Implicit Reward
المكافأة الضمنية
في DPO، المكافأة المُعرَّفة بنسبة لوغاريتم السياسة إلى المرجع: β·log(π_θ/π_ref)، دون الحاجة لنموذج مكافأة منفصل.
In DPO, the reward defined by the policy-to-reference log-ratio: β·log(π_θ/π_ref), requiring no separate reward model.
Also translated asالمكافأة المُتضمَّنة، الجزاء الضمني
First appears in this corpus in: Direct Preference Optimization: Your Language Model Is Secretly a Reward Model (2023)
Appears in these papers