Direct Preference Optimization
التحسين المباشر للتفضيلات
أسلوب مواءمة يُحسِّن النموذج اللغوي مباشرة على أزواج التفضيل بدالة خسارة تصنيفية مُوجَّهة، متخطياً نمذجة المكافأة والتعلم المعزز بالكامل.
An alignment method that optimizes a language model directly on preference pairs using a supervised classification loss, bypassing reward modeling and reinforcement learning entirely.
Also translated asDPO، التحسين المباشر بالتفضيل، المواءمة الفورية للاختيارات اللغوية، ضبط الترجيح المباشر للحلول
First appears in this corpus in: Direct Preference Optimization: Your Language Model Is Secretly a Reward Model (2023)
Appears in these papers
- Direct Preference Optimization: Your Language Model Is Secretly a Reward Model2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Mixtral of Experts2024in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦
- RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2023in the sky ✦