Glossary

Log-Ratio

نسبة اللوغاريتم

log(π_θ(y|x) / π_ref(y|x)): تقيس مقدار انحراف السياسة عن المرجع لإجابة معينة، وهي اللبنة الأساسية للمكافأة الضمنية في DPO.

log(π_θ(y|x) / π_ref(y|x)): measures how much the policy has drifted from the reference for a given response, the building block of DPO's implicit reward.

Also translated asاللوغاريتم النسبي، نسبة الأرجحية اللوغاريتمية

First appears in this corpus in: Direct Preference Optimization: Your Language Model Is Secretly a Reward Model (2023)

Appears in these papers