Log-Ratio
نسبة اللوغاريتم
log(π_θ(y|x) / π_ref(y|x)): تقيس مقدار انحراف السياسة عن المرجع لإجابة معينة، وهي اللبنة الأساسية للمكافأة الضمنية في DPO.
log(π_θ(y|x) / π_ref(y|x)): measures how much the policy has drifted from the reference for a given response, the building block of DPO's implicit reward.
Also translated asاللوغاريتم النسبي، نسبة الأرجحية اللوغاريتمية
First appears in this corpus in: Direct Preference Optimization: Your Language Model Is Secretly a Reward Model (2023)
Appears in these papers