Language Models2023intermediate11 min read

Direct Preference Optimization: Your Language Model Is Secretly a Reward Model

التحسين المباشر للتفضيلات: نموذجك اللغوي هو سرًّا نموذج مكافأة

Rafailov, R. · Sharma, A. · Mitchell, E. · Ermon, S. · Manning, C. D. · Finn, C. — NeurIPS

The problem

Aligning language models with human preferences via RLHF requires a complex, unstable pipeline: first train a on human comparisons, then run PPO to optimize the language model against that reward — while keeping it close to a via a KL penalty. This pipeline holds four models in GPU memory simultaneously, is sensitive to hyperparameters, and is notoriously difficult to stabilize. The question: can we achieve the same quality without the reward model and without reinforcement learning?

The contribution

DPO: a closed-form reparameterization of the RLHF objective. The key insight is that the optimal under the KL-constrained reward maximization has an analytical solution that lets us express the reward as a function of the policy itself. Substituting this into the Bradley-Terry preference model yields a simple loss over preference pairs — no reward model, no PPO, no value network. Two models in memory (policy + frozen reference), one supervised optimizer, and stable gradients.

The impact

DPO became the default alignment method for open-source LLMs within a year of publication. Llama 2, Zephyr, Mixtral, and dozens of other models adopted it as a simpler, more stable alternative to PPO-based RLHF. It spawned a family of variants — IPO, KTO, ORPO, SimPO, CPO — and fundamentally shifted how the field thinks about alignment: not as a reinforcement learning problem, but as a supervised classification problem on preference pairs.

Imagine you're training a chef. The traditional way (RLHF) hires a food critic first: you show them hundreds of dish pairs, they learn to score dishes, and then the chef spends weeks cooking and adjusting recipes to maximize the critic's scores. The critic and the chef might disagree, the critic might be gamed, and you're paying two salaries.

DPO takes a shortcut: skip the critic. Show the chef the same dish pairs directly — "diners preferred this plate over that one" — and let the chef adjust recipes to match those preferences immediately. Mathematically, this produces the exact same chef as the critic-then-adjust pipeline, but with half the staff and none of the drama.

The standard pipeline: RLHF in three stages

Before DPO, aligning a language model with human preferences followed a three-stage recipe that the field called RLHF:

Stage 1 — (SFT). Start from a pre-trained language model and fine-tune it on high-quality demonstration data. This produces a policy πSFT\pi_{\text{SFT}} that can follow instructions but doesn't yet reflect nuanced human preferences.

Stage 2 — Reward Modeling. Collect preference data: for each prompt xx, two candidate responses are sampled and a human labels which one is better (yw≻yly_w \succ y_l). A reward model rϕ(x,y)r_\phi(x, y) is trained to predict these preferences using the — the probability that response y1y_1 is preferred over y2y_2 is σ(r(x,y1)−r(x,y2))\sigma(r(x, y_1) - r(x, y_2)).

Stage 3 — RL Optimization. Use PPO to maximize the learned reward while staying close to the SFT policy via a penalty. This requires holding four models in memory: the policy being trained, the frozen reference, the reward model, and a value network for variance reduction.

Open in Lab
Compare the RLHF and DPO pipelines side by side. Notice how DPO collapses three stages into one.
The demo wakes as you arrive…

The key insight: the reward hides inside the policy

The RLHF objective asks: find the policy π\pi that maximizes expected reward while staying close to the reference policy πref\pi_{\text{ref}}:

max⁡π  Ex∼D, y∼π[r(x,y)]−β DKL[π(y∣x)∥πref(y∣x)]\max_\pi \; \mathbb{E}_{x \sim \mathcal{D}, \, y \sim \pi}[r(x, y)] - \beta \, D_{\text{KL}}[\pi(y|x) \| \pi_{\text{ref}}(y|x)]

This looks like it needs RL to solve. But Rafailov et al. noticed something remarkable: this optimization has a . The optimal policy is:

π∗(y∣x)=1Z(x)πref(y∣x)exp⁡ ⁣(1βr(x,y))\pi^*(y|x) = \frac{1}{Z(x)} \pi_{\text{ref}}(y|x) \exp\!\left(\frac{1}{\beta} r(x,y)\right)

where Z(x)Z(x) is a normalizing constant. Now comes the crucial trick: rearrange this equation to solve for the reward:

r(x,y)=βlog⁡π∗(y∣x)πref(y∣x)+βlog⁡Z(x)r(x, y) = \beta \log \frac{\pi^*(y|x)}{\pi_{\text{ref}}(y|x)} + \beta \log Z(x)

The reward is fully determined by the ratio of the optimal policy to the reference policy. The Z(x)Z(x) cancels when we compare two responses to the same prompt — and that's exactly what the Bradley-Terry model does.

Open in Lab
See how adjusting the policy's probability directly changes the implicit reward — no separate reward model needed.
The demo wakes as you arrive…

The DPO loss: alignment as classification

Substituting the implicit reward into the Bradley-Terry preference model gives us the DPO loss — a binary cross-entropy objective that directly optimizes the policy:

LDPO(πθ;πref)=−E(x,yw,yl)∼D[log⁡σ ⁣(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{\text{DPO}}(\pi_\theta; \pi_{\text{ref}}) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma\!\left( \beta \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right) \right]
The DPO loss — the entire alignment algorithm in one line — π_θ = policy being trained · π_ref = frozen reference (SFT model) · y_w = preferred response · y_l = rejected response · β = temperature controlling KL penalty strength · σ = sigmoid function

Read it from the inside out. The term log⁡πθ(y∣x)πref(y∣x)\log \frac{\pi_\theta(y|x)}{\pi_{\text{ref}}(y|x)} is the of the policy to the reference — how much the model has drifted from the starting point for a given response. The difference of two such log-ratios is the implicit reward margin: how much more the model rewards the preferred response versus the rejected one. The sigmoid wraps this margin into a probability, and the log turns it into a cross-entropy loss.

The has an elegant interpretation. It increases the probability of ywy_w and decreases the probability of yly_l, but weighted by how wrong the model currently is. If the model already strongly prefers ywy_w, the gradient is small — no need to push further. If it incorrectly prefers yly_l, the gradient is large. This automatic curriculum means DPO focuses its updates where they matter most.

Open in Lab
Drag the policy probabilities and watch how the DPO loss and gradient respond.
The demo wakes as you arrive…

The preference model: Bradley-Terry

The entire derivation rests on the Bradley-Terry model of preferences, a classic from 1952. It says: the probability that item A is preferred over item B depends on the difference in their "strengths" (here, rewards):

P(y1≻y2∣x)=σ(r(x,y1)−r(x,y2))P(y_1 \succ y_2 | x) = \sigma(r(x, y_1) - r(x, y_2))

where σ\sigma is the . This is the same logistic model used in Elo ratings for chess — one player's probability of winning depends on the rating gap. When the reward gap is large, the preference is nearly certain; when it's small, it's a coin flip.

DPO replaces the explicit reward rr with the implicit reward defined by the policy's log-ratio against the reference. The model that was "just" generating text is secretly assigning reward scores to every response it could produce.

Open in Lab
Adjust the implicit rewards and see how the Bradley-Terry model predicts preferences.
The demo wakes as you arrive…

The role of β: how far can the model drift?

The parameter β\beta controls the trade-off between reward maximization and staying close to the reference policy. Think of it as a leash length:

  • High β (long leash): the model is free to diverge far from the reference. It can chase the reward aggressively, but risks degenerating — producing text that scores well on the preference data but sounds unnatural.
  • Low β (short leash): the model must stay close to the reference. It's more conservative, producing safer but potentially less aligned outputs.

In the DPO loss, β\beta appears as a scaling factor on the log-ratios. A larger β\beta amplifies the reward margin, making the sigmoid more confident and the gradients sharper. A smaller β\beta compresses the margin, making updates gentler. In practice, β≈0.1\beta \approx 0.1 to 0.50.5 works well for most alignment tasks.

Open in Lab
Drag β to see how it controls the trade-off between alignment strength and reference fidelity.
The demo wakes as you arrive…

How DPO processes a preference pair

Let's trace what happens for a single training example (x,yw,yl)(x, y_w, y_l):

Step 1. Compute the log-probabilities of both responses under the current policy πθ\pi_\theta and the frozen reference πref\pi_{\text{ref}}.

Step 2. Calculate the log-ratios: r^θ(x,y)=βlog⁡πθ(y∣x)πref(y∣x)\hat{r}_\theta(x, y) = \beta \log \frac{\pi_\theta(y|x)}{\pi_{\text{ref}}(y|x)} for both ywy_w and yly_l. These are the implicit rewards.

Step 3. Compute the reward margin: r^θ(x,yw)−r^θ(x,yl)\hat{r}_\theta(x, y_w) - \hat{r}_\theta(x, y_l). This is how much the model currently "rewards" the preferred response over the rejected one.

Step 4. Pass the margin through the sigmoid σ\sigma to get the model's predicted preference probability.

Step 5. Compute the binary cross-entropy loss: −log⁡σ(margin)-\log \sigma(\text{margin}). If the model correctly ranks ywy_w above yly_l by a large margin, the loss is near zero. If it ranks them incorrectly, the loss is high.

Step 6. Backpropagate. The gradient increases πθ(yw∣x)\pi_\theta(y_w|x) and decreases πθ(yl∣x)\pi_\theta(y_l|x), weighted by the model's current error.

Open in Lab
Step through the 6 stages of DPO processing a single preference pair.
The demo wakes as you arrive…

The same idea in code

DPO loss, completepython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn.functional as F

def dpo_loss(
    policy_logps_w,    # log π_θ(y_w | x)
    policy_logps_l,    # log π_θ(y_l | x)
    ref_logps_w,       # log π_ref(y_w | x)
    ref_logps_l,       # log π_ref(y_l | x)
    beta=0.1,          # KL penalty strength
):
    """
    Direct Preference Optimization loss.
    All inputs: (batch_size,) tensors of per-example log-probabilities.
    """
    # Step 1: log-ratios = implicit rewards (up to β scaling)
    log_ratio_w = policy_logps_w - ref_logps_w   # how much policy drifted for y_w
    log_ratio_l = policy_logps_l - ref_logps_l   # how much policy drifted for y_l

    # Step 2: reward margin — positive means model prefers y_w
    margin = beta * (log_ratio_w - log_ratio_l)

    # Step 3: binary cross-entropy on the preference
    loss = -F.logsigmoid(margin).mean()

    # Bonus: extract the implicit reward for monitoring
    with torch.no_grad():
        implicit_reward_w = beta * log_ratio_w
        implicit_reward_l = beta * log_ratio_l
        accuracy = (margin > 0).float().mean()

    return loss, implicit_reward_w, implicit_reward_l, accuracy

# That's it. No reward model, no PPO, no value network.
# The policy IS the reward model — you just didn't know it yet.

Why DPO is equivalent to RLHF

DPO is not an approximation — it is mathematically equivalent to the RLHF objective under the Bradley-Terry preference model. The derivation has three steps:

  1. The KL-constrained reward maximization has a closed-form optimal policy (the result from statistical mechanics).
  2. This optimal policy can be rearranged to express the reward in terms of the policy and the reference.
  3. Substituting this "implicit reward" into the Bradley-Terry log-likelihood for preference data yields the DPO loss.

At no step was an approximation made. If the Bradley-Terry model is the correct preference model, the global minimum of the DPO loss is the same policy that the full RLHF pipeline would converge to — with infinite data and perfect optimization.

The practical difference is optimization trajectory: PPO explores by sampling from the current policy, while DPO optimizes on a fixed offline dataset. This makes DPO simpler but means it can't discover reward signals not already captured in the preference data.

What the experiments showed

Rafailov et al. tested DPO on three tasks:

Controlled sentiment generation — fine-tuning GPT-2 to generate positive movie reviews. DPO matched or exceeded PPO while being simpler to tune.

Summarization — aligning a model on the TL;DR summarization dataset. DPO achieved comparable reward scores to PPO-based approaches with significantly less computational cost.

Single-turn dialogue — training on the Anthropic Helpful and Harmless dataset. DPO produced responses with higher win rates against the SFT baseline than PPO, while requiring no reward model and converging faster.

Across all tasks, DPO showed a consistent pattern: similar or better alignment quality, with dramatically simpler training. The absence of a reward model eliminated , and the supervised optimization eliminated PPO's instability.

Why it mattered

  1. 2017

    RLHF introduced

    Christiano et al. introduced learning from human preferences via reward modeling + PPO. Effective but complex and computationally expensive.

  2. 2022

    InstructGPT / ChatGPT

    OpenAI scaled RLHF to production with InstructGPT, then launched ChatGPT. Proved alignment works at scale but exposed the pipeline's fragility.

  3. 2023

    DPO published

    Rafailov et al. showed the RLHF objective has a closed-form solution expressible as a supervised loss. Alignment became a classification problem.

  4. 2023

    Zephyr, Llama 2 Chat

    Open-source models rapidly adopted DPO. HuggingFace's Zephyr used DPO to align Mistral 7B, achieving ChatGPT-level performance on chat benchmarks.

  5. 2024

    The DPO family expands

    IPO (generalized beyond Bradley-Terry), KTO (unpaired preferences), SimPO (no reference model), ORPO (unified SFT + alignment) — each addressing a DPO limitation.

DPO's conceptual contribution may outlast its technical one. It revealed that reinforcement learning was not essential for alignment — the "RL" in "RLHF" was a choice, not a necessity. The same objective could be optimized with supervised learning, and the reward model was always implicit in the policy. This reframing opened a new design space for alignment algorithms that the field is still exploring.

CitationRafailov, Sharma, Mitchell, Ermon, Manning, Finn. Direct Preference Optimization: Your Language Model Is Secretly a Reward Model. NeurIPS, 2023.

Terms in this paper