Glossary

Implicit Reward

المكافأة الضمنية

في DPO، المكافأة المُعرَّفة بنسبة لوغاريتم السياسة إلى المرجع: β·log(π_θ/π_ref)، دون الحاجة لنموذج مكافأة منفصل.

In DPO, the reward defined by the policy-to-reference log-ratio: β·log(π_θ/π_ref), requiring no separate reward model.

Also translated asالمكافأة المُتضمَّنة، الجزاء الضمني

First appears in this corpus in: Direct Preference Optimization: Your Language Model Is Secretly a Reward Model (2023)

Appears in these papers