RLHF
التعلُّم المعزز من التغذية الراجعة البشرية
تقنية تدريب يُجمع فيها آلاف التفضيلات البشرية بين أزواج الاستجابات لتدريب نموذج مكافأة، ثمّ يُستخدم هذا النموذج لتحسين سياسة الذكاء الاصطناعي عبر التعلُّم المعزز.
A training technique where thousands of human preferences between response pairs train a reward model, which then optimizes the AI policy via reinforcement learning.
Also translated asRLHF، التعزيز بالتقييم البشري، التعلم المعزز بالتغذية الراجعة البشرية، التعلم المعزز بالتقييم البشري، التعلم المعزز بتقييم بشري
First appears in this corpus in: Learning to Summarize from Human Feedback (2020)
Appears in these papers
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Constitutional AI: Harmlessness from AI Feedback2022in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Mixtral of Experts2024in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision2023in the sky ✦