Reward Hacking
اختراق المكافأة
حين تتعلم السياسة استغلال ثغرات نموذج المكافأة بدلاً من التحسّن الحقيقي، فتُنتج مخرجات عالية الدرجة لكنها منخفضة الجودة فعلياً.
When a policy learns to exploit flaws in the reward model rather than genuinely improving, producing high-scoring but actually low-quality outputs.
Also translated asاستغلال المكافأة، استغلال دالة المكافأة، الاستغلال السلبي الثغرات دالة التحفيز، التحايل على أهداف البيئة الحوسبية، التحايل على المكافأة، التلاعب بالمكافأة، التلاعب بنموذج المكافأة، خداع نموذج المكافأة
First appears in this corpus in: Concrete Problems in AI Safety (2016)
Appears in these papers
- Concrete Problems in AI Safety2016in the sky ✦
- Concrete Problems in AI Safety2016in the sky ✦
- Deep Reinforcement Learning from Human Preferences2017in the sky ✦
- Deep Reinforcement Learning from Human Preferences2017in the sky ✦
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- Hindsight Experience Replay2017in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Training Language Models to Follow Instructions with Human Feedback2022in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Learning to Summarize from Human Feedback2020in the sky ✦
- Learning to Reason with LLMs2024in the sky ✦
- Learning to Reason with LLMs2024in the sky ✦