Language Models2022intermediate12 min read

Training Language Models to Follow Instructions with Human Feedback

تدريب النماذج اللغوية على اتّباع التعليمات باستخدام التغذية الراجعة البشرية

Ouyang, L. · Wu, J. · Jiang, X. · Almeida, D. · Wainwright, C. L. · Mishkin, P. · Zhang, C. · Agarwal, S. · Slama, K. · Ray, A. · Schulman, J. · Hilton, J. · Kelton, F. · Miller, L. · Simens, M. · Askell, A. · Welinder, P. · Christiano, P. · Leike, J. · Lowe, R. — NeurIPS

The problem

Making language models bigger does not make them better at following instructions. GPT-3 with 175 billion parameters could write fluent text but often hallucinated facts, produced toxic content, or simply ignored what the user asked for. The language modeling objective — predict the next token — is fundamentally misaligned with the user's actual goal: get a helpful, honest, and harmless answer. Public NLP benchmarks did not capture this gap because they measured capabilities, not .

The contribution

A three-step method to align language models with human intent: (1) (SFT) on human-written demonstrations of ideal behavior, (2) a (RM) on human comparisons of model outputs, and (3) optimizing the SFT model against the model using (PPO) with a KL penalty to stay close to the original . The resulting 1.3B- model was preferred by human labelers over the 175B GPT-3 despite being 100× smaller, and showed improvements in truthfulness and reduced with minimal regression on standard NLP benchmarks.

The impact

InstructGPT established RLHF as the standard post-training recipe for large language models. ChatGPT, Claude, Gemini, and Llama 2 all use variants of this pipeline. The paper proved that alignment is not just an ethics concern but a practical engineering technique: a small aligned model outperforms a much larger unaligned one. It shifted the field's focus from "make it bigger" to "make it follow instructions" and launched the era of conversational AI products.

GPT-3 is a brilliant student who aced every exam but has zero social skills. Ask it to summarize a paragraph and it might write a poem instead, or repeat your question back, or add something offensive — not because it can't do better, but because nobody ever told it what "better" means.

InstructGPT is the same student after an apprenticeship: first, a mentor shows them a few ideal answers (SFT). Then, a panel of judges ranks their attempts (Reward Model). Finally, the student practices under the judges' scoring system until their instincts align with what people actually want (PPO).

The surprise: the apprenticed student with a small notebook outperforms the unapprenticied genius with a library.

The alignment problem: bigger ≠ better

Language models are trained to predict the next token. That objective teaches fluency, grammar, and world knowledge — but it does not teach the model to be helpful. A model that predicts well might complete "Tell me a joke" with another question, a Wikipedia excerpt, or a toxic rant, because all of these are plausible next tokens in some internet context.

The core insight of InstructGPT is that the training objective is misaligned with the use case. Users want helpful, truthful, and harmless answers; the model was trained to mimic internet text. Scaling the model to 175 billion parameters made it more capable but did not fix this fundamental mismatch.

Three failure modes made the gap concrete:

  • — confidently stating false facts.
  • Toxicity — generating offensive or harmful content.
  • Instruction-ignoring — producing irrelevant responses that ignore what was asked.
Open in Lab
Same prompt, different behavior. See how the raw language model drifts while the aligned model follows the instruction.
The demo wakes as you arrive…

The solution: three steps from raw model to aligned assistant

InstructGPT transforms GPT-3 through a three-stage pipeline. Each stage addresses a different aspect of alignment, and each builds on the output of the previous one. Think of it as teaching by example → teaching by comparison → teaching by practice.

Step 1 — Supervised Fine-Tuning (SFT): Human labelers write ideal responses to a set of prompts. The model is fine-tuned on these demonstrations using standard — the same next-token prediction , but on high-quality instruction–response pairs instead of raw internet text. This gives the model an initial sense of what a good answer looks like.

Step 2 — Reward Model (RM): The SFT model generates multiple responses to each prompt. Human labelers rank these responses from best to worst. A separate model — the reward model — is trained to predict which response a human would prefer. It learns to output a scalar score: higher for better responses.

Step 3 — with PPO: The SFT model is further trained using Proximal Policy Optimization. For each prompt, the model generates a response, the reward model scores it, and the model's parameters are updated to maximize this score — with a KL penalty that prevents it from drifting too far from the SFT model and producing nonsense that tricks the reward model.

Open in Lab
Click each stage to explore the three-step RLHF pipeline that transforms GPT-3 into InstructGPT.
The demo wakes as you arrive…

Step 1: Supervised Fine-Tuning — learning from demonstrations

The simplest step, yet the foundation everything else rests on. OpenAI hired a team of 40 labelers and asked them to write ideal responses to a diverse set of prompts — some written by the labelers themselves, some sampled from the OpenAI API. The prompts covered generation, open QA, brainstorming, chat, rewriting, classification, and extraction.

The model was fine-tuned on roughly 13,000 demonstration pairs using the standard language modeling loss. Think of it as showing the model a gallery of excellent answers and saying "write more like these." SFT alone produces a noticeably better model, but it has a ceiling: you can only show so many examples, and the model can only generalize so far from imitation alone.

Step 2: The Reward Model — teaching a judge

Here is where the real leverage comes from. Instead of writing more examples (expensive and slow), we teach a model to evaluate quality — a judge that scores any response on a continuous scale.

The process: for each prompt, the SFT model generates 4 to 9 candidate responses. Human labelers rank all candidates from best to worst. From a ranking of KK responses, we extract (K2)\binom{K}{2} pairwise comparisons — for example, ranking 4 responses yields 6 pairs. Each pair becomes a training signal: "response A is better than response B."

The reward model is initialized from the SFT model with its final language-modeling head replaced by a single linear layer that outputs a scalar score. It is trained with a cross-entropy loss over the pairwise comparisons.

LRM(θ)=−1(K2)E(x,yw,yl)∼D[log⁡σ(rθ(x,yw)−rθ(x,yl))]\mathcal{L}_{RM}(\theta) = -\frac{1}{\binom{K}{2}} \mathbb{E}_{(x, y_w, y_l) \sim D} \left[ \log \sigma\left( r_\theta(x, y_w) - r_\theta(x, y_l) \right) \right]
Reward model loss — the Bradley-Terry pairwise ranking objective — The reward model learns from human preference comparisons rather than from absolute ratings. Given two candidate responses to the same prompt, it is trained to assign a higher score to the response preferred by humans and a lower score to the less preferred one. Over many comparisons, the model learns a ranking function that captures human judgments of quality, helpfulness, and alignment.

Why pairwise comparisons instead of absolute scores? It turns out that asking humans "is A better than B?" is much more consistent than asking "rate this response 1–7." Relative judgments produce cleaner training signal. And from a single ranking of KK items, you get (K2)\binom{K}{2} training pairs for free — an efficient use of expensive human attention.

Open in Lab
Rank 4 model responses yourself and watch the K-choose-2 comparison pairs being extracted.
The demo wakes as you arrive…

Step 3: PPO — learning from the judge

Now we have a trained judge (the reward model). The final step uses reinforcement learning to optimize the model's outputs against this judge.

The framework maps directly to RL concepts: the policy is the language model, the action is generating the next token, the state is the prompt plus tokens generated so far, and the reward comes from the reward model's score of the complete response.

But there is a trap. If we optimize purely for the reward model's score, the policy quickly learns to exploit quirks in the judge — producing gibberish that happens to score highly. This is called . The defense is a penalty: the model is penalized for straying too far from the original SFT policy, keeping its outputs coherent while improving alignment.

J(πθ)=Ex∼D, y∼πθ(⋅∣x)[rϕ(x,y)−β KL(πθ(⋅∣x) ∥ πSFT(⋅∣x))]J(\pi_\theta) = \mathbb{E}_{x \sim D,\, y \sim \pi_\theta(\cdot|x)} \left[ r_\phi(x, y) - \beta \, \text{KL}\big(\pi_\theta(\cdot|x) \,\|\, \pi_{\text{SFT}}(\cdot|x)\big) \right]
PPO-RLHF objective — maximize reward while staying close to SFT — During RLHF training, the model is encouraged to produce responses that receive higher scores from the reward model. At the same time, it is discouraged from changing too drastically from the supervised fine-tuned model it started from. A balancing factor controls this trade-off between improvement and stability. Stronger constraints keep the model closer to its original behavior, while weaker constraints allow larger behavioral changes in pursuit of higher reward.

The paper also introduces PPO-ptx, a variant that mixes PPO updates with a small amount of the original pre-training loss. This acts as an additional anchor, preventing the model from forgetting its general language abilities while learning to follow instructions — a phenomenon known as the . In practice, PPO-ptx showed the best balance of instruction-following ability and general NLP performance.

Open in Lab
Drag β to explore the trade-off. Too low = reward hacking. Too high = the model barely changes.
The demo wakes as you arrive…

The RLHF pipeline in code

The three-step RLHF pipeline, simplifiedpython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(x):
    return 1.0 / (1.0 + np.exp(-x))

# ── Step 1: Supervised Fine-Tuning ────────────────────────────
def sft_loss(model, prompt, ideal_response):
    """Standard cross-entropy on human-written ideal responses."""
    logits = model(prompt + ideal_response)
    return cross_entropy(logits, ideal_response)

# ── Step 2: Reward Model Training ─────────────────────────────
def reward_model_loss(rm, prompt, y_win, y_lose):
    """Bradley-Terry: push winner score above loser score."""
    r_win  = rm.score(prompt, y_win)
    r_lose = rm.score(prompt, y_lose)
    return -np.log(sigmoid(r_win - r_lose))

# ── Step 3: PPO with KL Penalty ───────────────────────────────
def ppo_reward(rm, policy, sft_policy, prompt, response, beta=0.2):
    """Reward = RM score minus KL divergence from SFT."""
    rm_score = rm.score(prompt, response)
    kl = compute_kl(policy, sft_policy, prompt, response)
    return rm_score - beta * kl

# The entire RLHF loop:
# 1. SFT on ~13k demonstrations → produces π_SFT
# 2. Train RM on ~33k comparisons → produces r_ϕ
# 3. PPO on ~31k prompts using r_ϕ → produces π_RLHF
# That's it. ChatGPT, Claude, Gemini all refine this recipe.

The surprising result: small aligned beats large unaligned

The headline result: InstructGPT with 1.3 billion parameters was preferred by human labelers over GPT-3 with 175 billion parameters — a model 100× its size. Labelers preferred InstructGPT outputs 71% of the time.

Beyond preference, InstructGPT showed concrete improvements:

  • Truthfulness increased on the TruthfulQA benchmark — the model hallucinated less.
  • Toxicity decreased — the model generated fewer harmful outputs.
  • improved dramatically — the model actually answered what was asked.
  • NLP benchmark regression was minimal — the alignment tax was small, especially with PPO-ptx.

The result was so compelling that OpenAI replaced the default GPT-3 API endpoint with InstructGPT — effectively saying "alignment is not optional, it is the product."

Open in Lab
The 1.3B InstructGPT beats the 175B GPT-3. Alignment wins over scale.
The demo wakes as you arrive…

The human labelers: who teaches the model?

OpenAI hired 40 contractors through Upwork and Scale AI. Labelers were selected for sensitivity to the preferences of different demographics and for their ability to identify and flag potentially harmful outputs. The prompt dataset came from two sources: labeler-written prompts covering diverse tasks, and prompts submitted to the OpenAI API by real users (with personal information removed).

The data breakdown across the three steps:

  • SFT dataset: ~13,000 prompt–demonstration pairs.
  • RM dataset: ~33,000 prompts, each with 4–9 ranked outputs → ~270,000 pairwise comparisons.
  • PPO dataset: ~31,000 prompts (no human labels needed — the reward model scores on the fly).

A critical design choice: the labelers ranked batches of outputs for each prompt. From KK ranked outputs, (K2)\binom{K}{2} comparison pairs are extracted — making each unit of labeler effort produce many training signals.

Limitations and open questions

InstructGPT was a breakthrough, but the paper is refreshingly honest about its limitations:

  • Labeler bias: The model aligns with labelers' preferences, not all of humanity's. Different labelers, different alignment.
  • Reward hacking is not fully solved: Even with KL penalties, the model can find subtle exploits in the reward model, especially on distribution-shift prompts.
  • The alignment tax: RLHF slightly degrades performance on standard NLP benchmarks — the model trades some "capability" for "behavior."
  • Still hallucinates: InstructGPT reduced hallucination but did not eliminate it. The model still confidently states false claims, just less often.
  • Safety is not solved: The model can still be manipulated into producing harmful outputs through adversarial prompting.

Why it mattered

  1. 2017

    PPO Algorithm

    Schulman et al. introduce Proximal Policy Optimization, the RL algorithm that would later power RLHF. Stable, efficient, and simpler than TRPO.

  2. 2020

    Learning to Summarize from Human Feedback

    Stiennon et al. show the three-step SFT → RM → PPO pipeline works for summarization. The direct precursor to InstructGPT.

  3. 2022

    InstructGPT

    Ouyang et al. scale the RLHF pipeline to general instruction-following. A 1.3B model outperforms 175B GPT-3. OpenAI replaces the GPT-3 API with InstructGPT.

  4. 2022

    ChatGPT

    Building on InstructGPT's recipe, ChatGPT brings RLHF-aligned models to the public, reaching 100 million users in two months.

  5. 2022

    Constitutional AI (Anthropic)

    Bai et al. replace human labelers with a set of written principles, letting AI generate its own feedback. Scales alignment without scaling human labor.

  6. 2023

    DPO — Direct Preference Optimization

    Rafailov et al. show RLHF's objective can be optimized without a separate reward model or RL — a single supervised loss on preference pairs. Simpler and cheaper.

  7. 2023

    Llama 2 (Meta)

    Meta applies RLHF at scale to an open-source model family, democratizing alignment techniques beyond proprietary labs.

InstructGPT transformed the field's understanding of what makes a language model useful. It showed that alignment is not a philosophical luxury but a practical requirement — and that the gap between a capable model and a useful one can be bridged with surprisingly little human data: 13,000 demonstrations and 33,000 comparisons to transform the most powerful language model into something people actually preferred.

CitationOuyang, Wu, Jiang, Almeida, Wainwright, Mishkin, Zhang, Agarwal, Slama, Ray, Schulman, Hilton, Kelton, Miller, Simens, Askell, Welinder, Christiano, Leike, Lowe. Training language models to follow instructions with human feedback. NeurIPS, 2022.

Terms in this paper