AI Safety2023advanced12 min read

Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

التعميم من الضعيف إلى القوي: استخراج القدرات القوية بإشراف ضعيف

Burns, C. · Izmailov, P. · Kirchner, J. H. · Baker, B. · Gao, L. · Aschenbrenner, L. · Chen, Y. · Ecoffet, A. · Joglekar, M. · Leike, J. · Sutskever, I. · Wu, J. — OpenAI Research

The problem

Current techniques like rely on humans being able to evaluate model behavior reliably. But future superhuman models will produce outputs — million-line codebases, complex scientific reasoning — that humans cannot fully evaluate. Humans will become weak supervisors of superhuman models. If we naively finetune a superhuman model with human feedback, will it learn what we actually want, or will it learn to imitate our limitations? We have no empirical framework to study this question today.

The contribution

A concrete empirical framework for studying : use weak models to supervise strong models and measure how much of the strong model's capability can be recovered. Across NLP benchmarks, chess puzzles, and ChatGPT reward modeling, naively finetuning strong GPT-4-family models on weak labels consistently recovers partial performance gaps. Simple methods — an auxiliary confidence loss, through intermediate sizes, and unsupervised generative finetuning — substantially improve recovery. With GPT-2-level supervision and the confidence loss, GPT-4 recovers close to GPT-3.5-level NLP performance.

The impact

This paper operationalized the superalignment problem — turning a theoretical concern into a measurable research agenda. It demonstrated that weak-to-strong is a real and widespread phenomenon, not just a theoretical hope, and showed that improving it is tractable with simple methods. It launched an active research area on scalable alignment methods and directly influenced OpenAI's superalignment program and subsequent work on alignment tax, Constitutional AI improvements, and scaling.

Imagine a junior coach hired to train an Olympic athlete. The coach can't run faster or jump higher than the athlete — but the coach can still point them toward the finish line, correct obvious form mistakes, and set up the right drills.

The question is: will the Olympian merely perform at the coach's level, copying the coach's own limitations? Or will the athlete use their superior abilities to generalize beyond the coach's guidance, reaching performance the coach could never achieve alone?

This paper turns that question into a precise experiment — replacing the coach with a weak AI model and the athlete with a strong one — and discovers that strong models do generalize beyond their weak supervisors, but not perfectly. Simple techniques can close most of the remaining gap.

The core challenge: humans as weak supervisors

Today's alignment pipeline works because humans can judge model outputs. When ChatGPT writes a paragraph, a human can tell if it is helpful, accurate, and safe. RLHF works precisely because human evaluators are competent supervisors.

But what happens when models become superhuman? If a model writes a million lines of code, designs a novel protein, or produces a mathematical proof that spans hundreds of pages, humans can no longer provide reliable supervision. The human becomes a weak supervisor — they can catch surface-level errors but miss deeper failures.

This creates a fundamental tension: the very technique we use to align models (human feedback) breaks down exactly when alignment matters most (superhuman capabilities). The authors call this the superalignment problem: how can weak supervisors control models much smarter than them?

Open in Lab
Today humans supervise weaker models. Tomorrow humans will be weak supervisors of superhuman models. This paper studies the analogous problem using weak models to supervise strong models.
The demo wakes as you arrive…

The experiment: weak models supervise strong models

The authors propose a clean three-step experimental framework. First, finetune a small model on ground truth labels to create a weak supervisor. Second, use that weak supervisor's predictions as training labels to finetune a much larger strong student model. Third, finetune the same strong model on ground truth labels to establish a strong ceiling — the best the strong model can do with perfect supervision.

The key metric is the Performance Gap Recovered (PGR): the fraction of the gap between weak performance and strong ceiling that the weak-to-strong student manages to close. A PGR of 0 means the student learned nothing beyond what the weak supervisor knows. A PGR of 1 means perfect generalization — the student recovered everything the strong model is capable of, despite only seeing weak labels.

The beauty of this setup is that it can be studied today with any pair of models — no superhuman AI required. It removes a key bottleneck: we no longer need to wait for superhuman models to study the alignment challenges they will create.

Open in Lab
Adjust the weak, weak-to-strong, and strong ceiling accuracies to see how PGR changes. PGR = 0 means the student only matches its supervisor; PGR = 1 means perfect generalization.
The demo wakes as you arrive…
PGR=accweak-to-strong−accweakaccstrong ceiling−accweak\text{PGR} = \frac{\text{acc}_{\text{weak-to-strong}} - \text{acc}_{\text{weak}}}{\text{acc}_{\text{strong ceiling}} - \text{acc}_{\text{weak}}}
Performance Gap Recovered — measuring weak-to-strong generalization — PGR measures what fraction of the gap between weak and strong ceiling performance is recovered by the weakly-supervised strong student. Values range from 0 (no generalization beyond the weak supervisor) to 1 (perfect recovery of strong model capabilities).

Main results: strong models generalize beyond weak labels

The authors tested their framework across three domains using pretrained models from the GPT-4 family, spanning 7 orders of magnitude of compute.

NLP benchmarks (22 tasks): The strongest results. When GPT-4 is finetuned on labels from a GPT-2-level supervisor, the PGR often exceeds 50%. The strong model consistently outperforms its weak supervisor by a wide margin — it doesn't just copy the weak labels; it leverages its own superior representations.

Chess puzzles: More mixed. Small supervisor-student gaps yield decent PGR (above 40%), but for very large gaps the PGR drops and even shows signs of inverse scaling — larger students sometimes generalize worse.

ChatGPT reward modeling: The hardest setting. PGR is typically around 10%, rarely exceeding 20%. This suggests that naive RLHF will scale poorly to superhuman models without new methods.

A critical finding across all three domains: PGR is almost universally positive. In nearly every configuration, the strong student outperforms its weak supervisor. The strong model doesn't merely copy — it generalizes. This is the encouraging core result.

Open in Lab
Explore how PGR varies across NLP, chess, and reward modeling. Click each domain to see how generalization changes with model size gap.
The demo wakes as you arrive…

Three methods to improve weak-to-strong generalization

Naive finetuning gives positive PGR, but falls short of recovering the strong model's full capability. The authors tested three simple methods to close this gap.

Method 1: auxiliary confidence loss

The most successful method. The core problem with naive finetuning is that the student is trained to imitate the weak supervisor — including its errors. We want the student to learn the supervisor's intent, not its mistakes.

The solution: add an auxiliary loss term that encourages the strong model to be confident in its own predictions, even when they disagree with the weak labels. This is closely related to minimization in . The intuition is that the strong model's own confident predictions are more likely to be correct than the weak labels where they disagree, because the strong model has better internal representations.

The confidence loss adds a term to the training objective that rewards the strong student for making high-confidence predictions on unlabeled data. This regularizes the student away from blindly imitating weak labels and toward its own internal knowledge.

With this simple addition, PGR on NLP tasks jumps from about 25% to nearly 80% when using the smallest weak supervisor and largest strong student. On many individual tasks, the method recovers almost all of the strong model's performance.

Open in Lab
See how the confidence loss helps the strong student disagree with weak errors. Toggle between naive finetuning and confidence loss to compare how PGR changes.
The demo wakes as you arrive…
L=LCE(fθ(x), yweak)+α⋅Lconf(fθ(x))\mathcal{L} = \mathcal{L}_{\text{CE}}(f_\theta(x),\, y^{\text{weak}}) + \alpha \cdot \mathcal{L}_{\text{conf}}(f_\theta(x))
Training with auxiliary confidence loss — The first term is the standard cross-entropy against weak labels. The second term is the confidence loss — it encourages the model to make high-confidence (low-entropy) predictions, acting as a regularizer that steers the student toward its own internal knowledge rather than blindly imitating weak labels. The hyperparameter α\alpha controls the balance.

Method 2: bootstrapping through intermediate sizes

Instead of one giant leap from weak to strong, take many small steps. Train a slightly-stronger model on the weak labels, use it to generate better labels, and repeat — climbing through a sequence of models like ascending a staircase.

Formally: construct a chain M1→M2→⋯→Mn\mathcal{M}_1 \to \mathcal{M}_2 \to \cdots \to \mathcal{M}_n of increasing sizes. Use weak labels from M1\mathcal{M}_1 to finetune M2\mathcal{M}_2, use M2\mathcal{M}_2's predictions to finetune M3\mathcal{M}_3, and so on.

This works well for chess puzzles, where naive PGR decreases with larger gaps. Since each small step has high PGR, chaining many small steps preserves generalization. With bootstrapping, accuracy continues to improve monotonically with model size, fixing the flat scaling observed with naive transfer. On NLP tasks, where naive PGR already scales well, bootstrapping provides only modest additional gains.

Open in Lab
Compare one big leap vs. many small steps. Click to toggle bootstrapping and see how accuracy improves when chaining intermediate models.
The demo wakes as you arrive…

Method 3: unsupervised generative finetuning

If weak-to-strong generalization works better when the task is "salient" — when the strong model has strong internal representations of the relevant concepts — then we can improve generalization by increasing saliency before doing weak-to-strong transfer.

The idea is to first finetune the strong model with a standard language modeling loss on task-relevant data, without using any labels. For reward modeling, this means finetuning on all prefix-completion pairs from the ChatGPT comparison dataset, ignoring which completion humans preferred.

This unsupervised step makes the concepts relevant to the task (like "what makes a good assistant response") more linearly represented in the model's activations. When weak-to-strong training follows, the weak supervisor's signal finds a better-prepared landscape. This boosts PGR by 10–20% in the difficult reward modeling setting.

The imitation trap: when strong models copy weak mistakes

Why doesn't the strong model always reach its full potential? The main failure mode is imitation: the student learns to reproduce the weak supervisor's predictions — including its systematic errors — rather than learning the underlying task.

Three pieces of evidence reveal this failure mode:

to weak labels. Ground truth test accuracy initially rises during training, then drops — even within a single epoch. The strong model starts by leveraging its own knowledge, but gradually learns to imitate the weak supervisor's errors. This "weak label overfitting" is distinct from standard overfitting to training examples.

High student-supervisor agreement. When measured on test data, the strong student agrees with the weak supervisor more often than the weak supervisor is correct. This means the student is copying some of the supervisor's mistakes.

Inverse scaling for imitation. The surprising finding: larger strong models agree less with their supervisor's errors. Despite having more capacity to memorize weak labels, larger pretrained models seem to have a harder time fitting the error patterns of other models. This suggests that pretrained representations actively resist learning incorrect patterns — an encouraging sign for scaling.

Open in Lab
Watch the tension between imitation and generalization during training. Early in training, the student leverages its own knowledge. Later, it starts imitating weak errors. The confidence loss slows this collapse.
The demo wakes as you arrive…

Saliency: when is the task easy to elicit?

Weak-to-strong generalization works best when the desired task is internally "salient" to the strong model — meaning the relevant concepts are well-represented in the model's activations.

Prompting reveals saliency. For the largest models, and 5-shot prompting can nearly match finetuning on many NLP tasks, indicating that the task knowledge is easily accessible. Even prompting with weak labels shows weak-to-strong generalization similar to finetuning, suggesting the phenomenon is not an artifact of the training procedure.

Weak finetuning linearizes representations. A fascinating discovery: finetuning on weak labels makes the ground truth concept more linearly represented in the model's activations. A trained on ground truth labels after weak finetuning significantly outperforms one trained on the base model. This means reorganizes the model's internal representations in a way that actually helps, even though the labels themselves are imperfect.

This insight connects to findings in the spurious correlations literature: finetuning on biased data can produce biased outputs, but surprisingly strong representations of the correct concept.

Remaining disanalogies and open problems

The authors are careful to note important limitations of their framework:

Imitation saliency. Future superhuman models will be explicitly trained on human data and will likely have very salient representations of human- level behavior. This makes the "human simulator" failure mode — where the model learns to output what a human would say rather than its own best answer — much more dangerous than in the current setup, where strong models are not pretrained to imitate weak models.

Pretraining leakage. Many NLP tasks studied here may appear indirectly in the pretraining data. This could make eliciting capabilities artificially easy today. Superhuman capabilities may be more "latent" — learned through self-supervised learning rather than directly observed — and thus harder to elicit.

Despite these caveats, the authors conducted a proof-of-concept using self-supervised vision models (DINO) where capabilities are latent and prompting is impossible. Even there, weak-to-strong generalization succeeded, giving reason for cautious optimism.

Context: the road to superalignment

  1. 2017

    RLHF introduced (Christiano et al.)

    Deep reinforcement learning from human preferences established the paradigm of training reward models from human comparisons, then optimizing against them with RL.

  2. 2018

    AI safety via debate (Irving et al.)

    Proposed scalable oversight through adversarial debate between AI systems, with humans serving as judges. A precursor to thinking about weak human supervision of strong models.

  3. 2022

    InstructGPT and RLHF at scale (Ouyang et al.)

    Demonstrated that RLHF can align large language models to follow instructions. Made alignment practical but raised the question of what happens when models surpass human evaluators.

  4. 2023

    OpenAI Superalignment team formed

    Jan Leike and Ilya Sutskever announced a dedicated team to solve the core technical challenges of aligning superhuman AI systems, with this paper as a key early output.

  5. 2023

    This paper — weak-to-strong generalization

    Provided the first large-scale empirical framework for studying superalignment. Showed that weak-to-strong generalization is real and improvable, making the superalignment problem empirically tractable.

The paper concludes with a call to action: superhuman models may arrive this decade, and aligning them is among the most important unsolved technical problems. The weak-to-strong framework provides a way to make empirical progress now — before superhuman models exist — on the fundamental challenge of ensuring they do what we want.

CitationBurns, Izmailov, Kirchner, Baker, Gao, Aschenbrenner, Chen, Ecoffet, Joglekar, Leike, Sutskever, Wu. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. OpenAI Research, 2023.

Terms in this paper