Reward Model
نموذج المكافأة
شبكة عصبية مُدرَّبة على تفضيلات بشرية لتقدير جودة استجابة النموذج، تُستخدم كإشارة تدريب في التعلم المعزز بدلاً من التقييم البشري المباشر.
A neural network trained on human preferences to estimate the quality of a model's response, used as a training signal in reinforcement learning instead of direct human evaluation.
Also translated asمُقيِّم المكافأة، نموذج التفضيلات، نموذج التقييم، نموذج التقييم البشري، نموذج الثواب، نموذج الجزاء
First appears in this corpus in: Deep Reinforcement Learning from Human Preferences (2017)
Appears in these papers
- AI Safety via Debate2018in the sky ✦
- Constitutional AI: Harmlessness from AI Feedback2022in the sky ✦
- Deep Reinforcement Learning from Human Preferences2017in the sky ✦
- Deep Reinforcement Learning from Human Preferences2017in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- LIMA: Less Is More for Alignment2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Mixtral of Experts2024in the sky ✦
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model2020in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection2023in the sky ✦
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection2023in the sky ✦
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision2023in the sky ✦
- WebGPT: Browser-Assisted Question-Answering with Human Feedback2021in the sky ✦
- WebGPT: Browser-Assisted Question-Answering with Human Feedback2021in the sky ✦