Reinforcement Learning2020intermediate13 min read

Learning to Summarize from Human Feedback

تعلُّم التلخيص من التغذية الراجعة البشرية

Stiennon, N. · Ouyang, L. · Wu, J. · Ziegler, D. M. · Lowe, R. · Voss, C. · Radford, A. · Amodei, D. · Christiano, P. — NeurIPS

The problem

By 2020, language models were typically fine-tuned to maximize the likelihood of human-written text and evaluated using automatic metrics like . But maximizing likelihood treats all errors equally — inventing a fact is penalized the same as choosing a synonym. ROUGE, in turn, correlates poorly with actual summary quality as judged by humans. The core problem: the training objective and the evaluation metric are both rough proxies for what we actually care about — producing summaries that humans find genuinely good.

The contribution

A three-step pipeline that replaces proxy metrics with direct human judgment. First, collect 64,000+ human comparisons between pairs of summaries. Second, train a to predict which summary humans prefer. Third, use that reward model as a to fine-tune a GPT-style with , regularized by a KL penalty against the supervised baseline. The result: a 1.3B parameter model that produces summaries preferred over those from a 10× larger supervised model, and a 6.7B model whose summaries are preferred over the original human reference summaries. The models also transfer zero-shot to CNN/DM news .

The impact

This paper established the recipe — → reward model → PPO — that became the backbone of InstructGPT, ChatGPT, and virtually every aligned since. It proved that human preferences can be distilled into a learnable reward and used to steer generation beyond what or automatic metrics can achieve. The methodology generalized far beyond summarization: dialogue, instruction following, code generation, and safety alignment all adopted the same paradigm.

Imagine a cooking competition where contestants are judged not by a recipe book, but by diners tasting pairs of dishes and picking the one they prefer. After thousands of tastings, a food critic learns exactly what the diners value — balance, freshness, depth of flavor.

Now a new chef enters. Instead of following a recipe (supervised learning) or optimizing a calorie counter (ROUGE), this chef cooks dish after dish and asks the critic for a score. The critic's palate is the reward signal. Over many rounds, the chef's dishes become so refined that diners prefer them even over the original award-winning dishes written by human chefs.

This paper turned language models into that chef, and human preference into that critic.

The misalignment problem: training loss ≠ quality

Before this paper, the standard recipe for training a summarization model was straightforward: take a large language model, fine-tune it on pairs of (article, human-written summary), and measure success with ROUGE — a metric that counts overlapping n-grams between the generated summary and the reference.

This approach has a deep structural flaw. Maximum likelihood training treats every prediction equally: confusing "Monday" with "Tuesday" (a factual error) is penalized the same as choosing "big" instead of "large" (a stylistic preference). The model has no way to learn that some errors matter far more than others.

ROUGE makes the problem worse. Because it only measures surface overlap, a summary can score well on ROUGE while missing the point entirely, or score poorly while capturing the essence perfectly in different words. The authors found that ROUGE's agreement with human judgments drops to chance level (~50%) when comparing their best models — precisely where evaluation matters most.

Think of it like grading essays by counting how many words match the teacher's answer key. A student who copies random sentences from the source might score higher than a student who wrote a thoughtful, accurate paraphrase.

Open in Lab
Compare a supervised model summary (optimizing likelihood) with a human feedback model summary (optimizing preferences). Notice how the RLHF summary captures the key points more faithfully.
The demo wakes as you arrive…

The three-step RLHF pipeline

The paper's central contribution is a pipeline that replaces proxy metrics with direct human judgment, organized in three clear steps. Each step feeds into the next, building a bridge from raw human preferences to a policy that generates high-quality text.

Think of it as building a self-improving writing assistant in three stages: first, learn what good writing looks like by watching an expert compare drafts; second, build an internal quality detector based on those expert opinions; third, practice writing, using that internal detector as a coach.

Open in Lab
The three steps of the RLHF pipeline — click each step to explore its role.
The demo wakes as you arrive…

Step 1: collecting human comparisons

The foundation of the pipeline is a of human preferences. For each Reddit post from the TL;DR dataset, the authors sampled summaries from multiple sources — the current policy, the supervised baseline, the original human reference, and various other models. Pairs of summaries were then shown to human labelers, who selected the better one.

This binary comparison format is crucial. Rather than asking labelers to rate quality on a numerical scale (where "7 out of 10" means different things to different people), comparisons ask a simpler question: "Which summary better captures the post?" This reduces noise and produces more consistent signal.

The authors invested heavily in labeler quality. They on-boarded labelers with detailed instructions, maintained an active chat channel for questions, and continuously monitored agreement rates. The result: labelers agreed with researchers 77% of the time — nearly matching the 73% rate at which researchers agreed with each other.

The final dataset contains over 64,000 summary comparisons — an unusually large and carefully curated preference dataset that the authors publicly released.

Step 2: training the reward model

The reward model converts pairwise human preferences into a scalar score. Given a post and a candidate summary, it outputs a number representing how good that summary is — higher means more likely to be preferred by a human.

Architecturally, the reward model starts as a copy of the supervised summarization model (a GPT-style ), with one modification: a randomly initialized linear head replaces the language modeling head, outputting a single scalar instead of token probabilities. This is key — the reward model inherits all the language understanding from and , and only needs to learn the relatively simpler task of quality comparison.

Think of it as taking a skilled writer and teaching them to be a judge. They already understand language deeply; they just need to learn what makes one summary better than another.

LRM(θ)=− E(x, yw, yl) ∼ D ⁣[log⁡ σ ⁣(rθ(x, yw)−rθ(x, yl))]\mathcal{L}_{\text{RM}}(\theta) = -\,\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\,\mathcal{D}}\!\Big[\log\,\sigma\!\big(r_\theta(x,\,y_w) - r_\theta(x,\,y_l)\big)\Big]
Reward Model Loss (Bradley-Terry) — For a post x, the model assigns scalar rewards to the preferred summary y_w and the rejected summary y_l. The loss pushes the reward of the preferred summary higher than the rejected one. σ is the sigmoid function — the same function used in logistic regression, converting the reward gap into a probability.
Open in Lab
See how the reward model scores different summaries. Adjust quality dimensions to understand what the model learned.
The demo wakes as you arrive…

Step 3: optimizing the policy with PPO

With a trained reward model in hand, the final step is to use it as the reward signal for . The policy — a GPT model initialized from the supervised fine-tuned — generates summaries token by token. When it finishes a summary, the reward model scores it, and the PPO algorithm updates the policy to increase the probability of actions that led to higher rewards.

But raw reward optimization is dangerous. Without constraints, the policy would quickly learn to exploit quirks in the reward model — producing text that scores highly but reads as gibberish to humans. This is , and it's a fundamental challenge in RL.

The key defense is a penalty that anchors the policy to the original supervised model. This penalty grows as the RL policy's outputs diverge from what the supervised model would produce, acting like a leash that prevents the policy from wandering too far into unexplored territory where the reward model's predictions are unreliable.

R(x,y)=rθ(x,y)  −  β log⁡ ⁣πϕRL(y∣x)πSFT(y∣x)R(x, y) = r_\theta(x, y) \;-\; \beta \,\log\!\frac{\pi_\phi^{\text{RL}}(y \mid x)}{\pi^{\text{SFT}}(y \mid x)}
Full Reward with KL Penalty — The total reward combines the reward model score r_θ with a KL penalty. β controls how tightly the RL policy π^RL is leashed to the supervised policy π^SFT. Higher β means the policy stays closer to supervised behavior; lower β allows more exploration but risks reward hacking.
Open in Lab
Explore the tension between reward optimization and KL penalty. As KL divergence grows, the policy drifts from supervised behavior and eventually over-optimizes.
The demo wakes as you arrive…

Results: human feedback beats scale

The results are striking. The 1.3B model produces summaries that humans prefer over those from a supervised model 10 times its size (6.7B supervised). Even more remarkably, the 6.7B human feedback model produces summaries that humans prefer over the original human-written reference summaries in the dataset.

On a detailed quality analysis across four axes — coverage, accuracy, coherence, and overall quality — the human feedback models outperformed supervised baselines on every dimension, with the largest gains in coverage. The 6.7B PPO model achieved a perfect 7/7 overall score 45% of the time, compared to 20% for the supervised baseline and 23% for human references.

Perhaps most impressive is the transfer result. The models, trained only on Reddit posts, produced high-quality summaries of CNN/DM news articles without any news-specific fine-tuning. The 6.7B human feedback model nearly matched the quality of a model specifically fine-tuned on CNN/DM data — despite never seeing a single news article during training.

Open in Lab
Fraction of time human evaluators preferred each model's summaries over the reference summaries.
The demo wakes as you arrive…

The reward hacking frontier

One of the paper's most important analyses concerns what happens when you optimize the reward model too much. The authors created a range of policies with varying optimization strength by adjusting the KL penalty coefficient β.

Under light optimization, quality improves — the reward model successfully guides the policy toward better summaries. But as optimization intensifies, actual human preferences plateau and then decline, even as the reward model's predicted score keeps rising. The reward model becomes "fooled" by outputs that exploit its imperfections.

This is reward hacking (also called Goodhart's Law in this context): when a measure becomes a target, it ceases to be a good measure. The KL penalty is the primary defense: it limits how far the policy can drift from the the reward model was trained to evaluate.

The authors showed that the same overoptimization problem occurs with ROUGE — optimizing ROUGE directly peaks sooner and at a lower quality level than optimizing the learned reward model. This suggests that learned reward models, while imperfect, capture more of what humans care about than hand-crafted metrics.

Open in Lab
Drag the optimization slider to see how summary quality first improves, then degrades as the reward model is over-optimized.
The demo wakes as you arrive…

Why learned rewards beat ROUGE

The paper provides compelling evidence that learned reward models capture quality dimensions that ROUGE simply cannot. When the authors optimized ROUGE directly using best-of-N , quality peaked much sooner and at a lower level than when optimizing the learned reward model.

The reason is structural. ROUGE counts n-gram overlaps — it measures surface similarity, not semantic faithfulness. A summary that copies random sentences from the source can score well on ROUGE. A summary that captures the meaning perfectly in completely different words scores poorly. The learned reward model, by contrast, has been trained on genuine human judgments of quality, so it can reward semantic accuracy, appropriate coverage, and coherent writing — none of which ROUGE measures.

This result carries a broader lesson for machine learning: when your metric is a rough proxy, optimizing it aggressively makes things worse. Learning the metric from human feedback, while harder, can yield genuine improvements.

Key design decisions

Several technical choices in this paper became standard practice for future RLHF systems.

The supervised fine-tuned (SFT) model serves triple duty: it initializes the reward model, the RL policy, and the value function. Starting all components from the same pre-trained checkpoint ensures they share the same understanding of language, which stabilizes training.

Temperature 0 sampling was used for all final evaluations. The authors found that (T=0) consistently outperformed higher temperatures and for their fine-tuned models, suggesting that once the model distribution is well-shaped by RLHF, you don't need stochastic diversity to get quality.

The TL;DR dataset was chosen over CNN/DM specifically because extractive baselines perform well on CNN/DM (a lead-3 baseline — just taking the first three sentences — was preferred over CNN/DM reference summaries by labelers). TL;DR requires genuinely abstractive summarization, making it a harder and more informative testbed.

The output length was constrained to 24-48 tokens for all summaries, minimizing the confounding effect of length on quality judgments. Even after controlling for length, the human feedback models maintained a clear quality advantage.

Pseudocode: PPO training step for RLHFpython

Simplified to show the idea — not the real implementation.

# Pseudocode for one PPO training step
def rlhf_training_step(post, policy, reward_model, sft_model, beta):
    # 1. Generate summary from current policy
    summary = policy.generate(post, max_tokens=48)

    # 2. Score with reward model
    rm_score = reward_model(post, summary)

    # 3. Compute KL penalty against supervised model
    log_ratio = policy.log_prob(summary, post) - sft_model.log_prob(summary, post)
    kl_penalty = beta * log_ratio

    # 4. Full reward = RM score minus KL penalty
    reward = rm_score - kl_penalty

    # 5. Update policy using PPO
    policy.ppo_update(reward, summary, post)

Legacy: the recipe that aligned AI

  1. 2017

    Deep RL from Human Preferences

    Christiano et al. demonstrated RLHF for Atari and MuJoCo environments — training agents from 900 binary human comparisons instead of hand-crafted reward functions.

  2. 2019

    Fine-Tuning LMs from Human Preferences

    Ziegler et al. applied RLHF to language models for the first time, including summarization. However, they trained in an online manner with lower labeler agreement.

  3. 2020

    Learning to Summarize (this paper)

    Scaled RLHF to 6.7B parameters with batch collection, high labeler quality, and separated policy/value networks. First model to surpass human reference summaries.

  4. 2021

    WebGPT

    Applied the RLHF recipe to question answering with web browsing. The model learned to search, cite sources, and produce factual answers preferred by humans.

  5. 2022

    InstructGPT

    Applied the same three-step recipe (SFT → RM → PPO) to general instruction following. A 1.3B InstructGPT was preferred over a 175B GPT-3, demonstrating that alignment techniques can substitute for raw scale.

  6. 2022

    ChatGPT

    A sibling model of InstructGPT fine-tuned for dialogue. The RLHF pipeline from this summarization paper became the core technology behind the most widely adopted AI system in history.

The deepest lesson of this paper is not about summarization — it's about the gap between what we optimize and what we want. Maximum likelihood, ROUGE, and other proxy metrics are convenient but fundamentally misaligned with human preferences. This paper showed that learning directly from human judgment, despite being harder and more expensive, produces qualitatively better results.

Every time you interact with a modern aligned language model — from ChatGPT to Claude — the conversation traces back to the ideas in this paper. The specific details have evolved, but the core recipe remains: collect human preferences, train a reward model, optimize with RL, and use KL to stay grounded. It is one of the most consequential papers in the alignment era of AI.

CitationStiennon, Ouyang, Wu, Ziegler, Lowe, Voss, Radford, Amodei, Christiano. Learning to Summarize from Human Feedback. NeurIPS, 2020.

Terms in this paper