AI Safety2023advanced12 min read
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
التعميم من الضعيف إلى القوي: استخراج القدرات القوية بإشراف ضعيف
Burns, C. · Izmailov, P. · Kirchner, J. H. · Baker, B. · Gao, L. · Aschenbrenner, L. · Chen, Y. · Ecoffet, A. · Joglekar, M. · Leike, J. · Sutskever, I. · Wu, J. — OpenAI Research
The problem
Current techniques like rely on humans being able to evaluate model behavior reliably. But future superhuman models will produce outputs — million-line codebases, complex scientific reasoning — that humans cannot fully evaluate. Humans will become weak supervisors of superhuman models. If we naively finetune a superhuman model with human feedback, will it learn what we actually want, or will it learn to imitate our limitations? We have no empirical framework to study this question today.
The contribution
A concrete empirical framework for studying : use weak models to supervise strong models and measure how much of the strong model's capability can be recovered. Across NLP benchmarks, chess puzzles, and ChatGPT reward modeling, naively finetuning strong GPT-4-family models on weak labels consistently recovers partial performance gaps. Simple methods — an auxiliary confidence loss, through intermediate sizes, and unsupervised generative finetuning — substantially improve recovery. With GPT-2-level supervision and the confidence loss, GPT-4 recovers close to GPT-3.5-level NLP performance.
The impact
This paper operationalized the superalignment problem — turning a theoretical concern into a measurable research agenda. It demonstrated that weak-to-strong is a real and widespread phenomenon, not just a theoretical hope, and showed that improving it is tractable with simple methods. It launched an active research area on scalable alignment methods and directly influenced OpenAI's superalignment program and subsequent work on alignment tax, Constitutional AI improvements, and scaling.
Imagine a junior coach hired to train an Olympic athlete. The coach can't run faster or jump higher than the athlete — but the coach can still point them toward the finish line, correct obvious form mistakes, and set up the right drills.
The question is: will the Olympian merely perform at the coach's level, copying the coach's own limitations? Or will the athlete use their superior abilities to generalize beyond the coach's guidance, reaching performance the coach could never achieve alone?
This paper turns that question into a precise experiment — replacing the coach with a weak AI model and the athlete with a strong one — and discovers that strong models do generalize beyond their weak supervisors, but not perfectly. Simple techniques can close most of the remaining gap.
The core challenge: humans as weak supervisors
Today's alignment pipeline works because humans can judge model outputs. When ChatGPT writes a paragraph, a human can tell if it is helpful, accurate, and safe. RLHF works precisely because human evaluators are competent supervisors.
But what happens when models become superhuman? If a model writes a million lines of code, designs a novel protein, or produces a mathematical proof that spans hundreds of pages, humans can no longer provide reliable supervision. The human becomes a weak supervisor — they can catch surface-level errors but miss deeper failures.
This creates a fundamental tension: the very technique we use to align models (human feedback) breaks down exactly when alignment matters most (superhuman capabilities). The authors call this the superalignment problem: how can weak supervisors control models much smarter than them?
The experiment: weak models supervise strong models
The authors propose a clean three-step experimental framework. First, finetune a small model on ground truth labels to create a weak supervisor. Second, use that weak supervisor's predictions as training labels to finetune a much larger strong student model. Third, finetune the same strong model on ground truth labels to establish a strong ceiling — the best the strong model can do with perfect supervision.
The key metric is the Performance Gap Recovered (PGR): the fraction of the gap between weak performance and strong ceiling that the weak-to-strong student manages to close. A PGR of 0 means the student learned nothing beyond what the weak supervisor knows. A PGR of 1 means perfect generalization — the student recovered everything the strong model is capable of, despite only seeing weak labels.
The beauty of this setup is that it can be studied today with any pair of models — no superhuman AI required. It removes a key bottleneck: we no longer need to wait for superhuman models to study the alignment challenges they will create.
Main results: strong models generalize beyond weak labels
The authors tested their framework across three domains using pretrained models from the GPT-4 family, spanning 7 orders of magnitude of compute.
NLP benchmarks (22 tasks): The strongest results. When GPT-4 is finetuned on labels from a GPT-2-level supervisor, the PGR often exceeds 50%. The strong model consistently outperforms its weak supervisor by a wide margin — it doesn't just copy the weak labels; it leverages its own superior representations.
Chess puzzles: More mixed. Small supervisor-student gaps yield decent PGR (above 40%), but for very large gaps the PGR drops and even shows signs of inverse scaling — larger students sometimes generalize worse.
ChatGPT reward modeling: The hardest setting. PGR is typically around 10%, rarely exceeding 20%. This suggests that naive RLHF will scale poorly to superhuman models without new methods.
A critical finding across all three domains: PGR is almost universally positive. In nearly every configuration, the strong student outperforms its weak supervisor. The strong model doesn't merely copy — it generalizes. This is the encouraging core result.
Three methods to improve weak-to-strong generalization
Naive finetuning gives positive PGR, but falls short of recovering the strong model's full capability. The authors tested three simple methods to close this gap.
Method 1: auxiliary confidence loss
The most successful method. The core problem with naive finetuning is that the student is trained to imitate the weak supervisor — including its errors. We want the student to learn the supervisor's intent, not its mistakes.
The solution: add an auxiliary loss term that encourages the strong model to be confident in its own predictions, even when they disagree with the weak labels. This is closely related to minimization in . The intuition is that the strong model's own confident predictions are more likely to be correct than the weak labels where they disagree, because the strong model has better internal representations.
The confidence loss adds a term to the training objective that rewards the strong student for making high-confidence predictions on unlabeled data. This regularizes the student away from blindly imitating weak labels and toward its own internal knowledge.
With this simple addition, PGR on NLP tasks jumps from about 25% to nearly 80% when using the smallest weak supervisor and largest strong student. On many individual tasks, the method recovers almost all of the strong model's performance.
Method 2: bootstrapping through intermediate sizes
Instead of one giant leap from weak to strong, take many small steps. Train a slightly-stronger model on the weak labels, use it to generate better labels, and repeat — climbing through a sequence of models like ascending a staircase.
Formally: construct a chain of increasing sizes. Use weak labels from to finetune , use 's predictions to finetune , and so on.
This works well for chess puzzles, where naive PGR decreases with larger gaps. Since each small step has high PGR, chaining many small steps preserves generalization. With bootstrapping, accuracy continues to improve monotonically with model size, fixing the flat scaling observed with naive transfer. On NLP tasks, where naive PGR already scales well, bootstrapping provides only modest additional gains.
Method 3: unsupervised generative finetuning
If weak-to-strong generalization works better when the task is "salient" — when the strong model has strong internal representations of the relevant concepts — then we can improve generalization by increasing saliency before doing weak-to-strong transfer.
The idea is to first finetune the strong model with a standard language modeling loss on task-relevant data, without using any labels. For reward modeling, this means finetuning on all prefix-completion pairs from the ChatGPT comparison dataset, ignoring which completion humans preferred.
This unsupervised step makes the concepts relevant to the task (like "what makes a good assistant response") more linearly represented in the model's activations. When weak-to-strong training follows, the weak supervisor's signal finds a better-prepared landscape. This boosts PGR by 10–20% in the difficult reward modeling setting.
The imitation trap: when strong models copy weak mistakes
Why doesn't the strong model always reach its full potential? The main failure mode is imitation: the student learns to reproduce the weak supervisor's predictions — including its systematic errors — rather than learning the underlying task.
Three pieces of evidence reveal this failure mode:
to weak labels. Ground truth test accuracy initially rises during training, then drops — even within a single epoch. The strong model starts by leveraging its own knowledge, but gradually learns to imitate the weak supervisor's errors. This "weak label overfitting" is distinct from standard overfitting to training examples.
High student-supervisor agreement. When measured on test data, the strong student agrees with the weak supervisor more often than the weak supervisor is correct. This means the student is copying some of the supervisor's mistakes.
Inverse scaling for imitation. The surprising finding: larger strong models agree less with their supervisor's errors. Despite having more capacity to memorize weak labels, larger pretrained models seem to have a harder time fitting the error patterns of other models. This suggests that pretrained representations actively resist learning incorrect patterns — an encouraging sign for scaling.
Saliency: when is the task easy to elicit?
Weak-to-strong generalization works best when the desired task is internally "salient" to the strong model — meaning the relevant concepts are well-represented in the model's activations.
Prompting reveals saliency. For the largest models, and 5-shot prompting can nearly match finetuning on many NLP tasks, indicating that the task knowledge is easily accessible. Even prompting with weak labels shows weak-to-strong generalization similar to finetuning, suggesting the phenomenon is not an artifact of the training procedure.
Weak finetuning linearizes representations. A fascinating discovery: finetuning on weak labels makes the ground truth concept more linearly represented in the model's activations. A trained on ground truth labels after weak finetuning significantly outperforms one trained on the base model. This means reorganizes the model's internal representations in a way that actually helps, even though the labels themselves are imperfect.
This insight connects to findings in the spurious correlations literature: finetuning on biased data can produce biased outputs, but surprisingly strong representations of the correct concept.
Remaining disanalogies and open problems
The authors are careful to note important limitations of their framework:
Imitation saliency. Future superhuman models will be explicitly trained on human data and will likely have very salient representations of human- level behavior. This makes the "human simulator" failure mode — where the model learns to output what a human would say rather than its own best answer — much more dangerous than in the current setup, where strong models are not pretrained to imitate weak models.
Pretraining leakage. Many NLP tasks studied here may appear indirectly in the pretraining data. This could make eliciting capabilities artificially easy today. Superhuman capabilities may be more "latent" — learned through self-supervised learning rather than directly observed — and thus harder to elicit.
Despite these caveats, the authors conducted a proof-of-concept using self-supervised vision models (DINO) where capabilities are latent and prompting is impossible. Even there, weak-to-strong generalization succeeded, giving reason for cautious optimism.
Context: the road to superalignment
2017
RLHF introduced (Christiano et al.)
Deep reinforcement learning from human preferences established the paradigm of training reward models from human comparisons, then optimizing against them with RL.
2018
AI safety via debate (Irving et al.)
Proposed scalable oversight through adversarial debate between AI systems, with humans serving as judges. A precursor to thinking about weak human supervision of strong models.
2022
InstructGPT and RLHF at scale (Ouyang et al.)
Demonstrated that RLHF can align large language models to follow instructions. Made alignment practical but raised the question of what happens when models surpass human evaluators.
2023
OpenAI Superalignment team formed
Jan Leike and Ilya Sutskever announced a dedicated team to solve the core technical challenges of aligning superhuman AI systems, with this paper as a key early output.
2023
This paper — weak-to-strong generalization
Provided the first large-scale empirical framework for studying superalignment. Showed that weak-to-strong generalization is real and improvable, making the superalignment problem empirically tractable.
The paper concludes with a call to action: superhuman models may arrive this decade, and aligning them is among the most important unsolved technical problems. The weak-to-strong framework provides a way to make empirical progress now — before superhuman models exist — on the fundamental challenge of ensuring they do what we want.
CitationBurns, Izmailov, Kirchner, Baker, Gao, Aschenbrenner, Chen, Ecoffet, Joglekar, Leike, Sutskever, Wu. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. OpenAI Research, 2023.
Terms in this paper
- Weak Supervisionالإشراف الضعيف
- Superalignmentالمحاذاة الفائقة
- Scalable Oversightالإشراف القابل للتوسُّع
- Generalizationالتعميم
- Fine-Tuningالضبط الدقيق
- Knowledge Distillationتقطير المعرفة
- Auxiliary Objectiveالهدف المساعد
- Confidence Scoreدرجة الثقة
- Student Modelنموذج الطالب
- Teacher Modelنموذج المعلم
- Bootstrappingالتمهيد الذاتي
- Alignmentالمحاذاة
- RLHFالتعلُّم المعزز من التغذية الراجعة البشرية
- Representationالتمثيل الرقمي
- Overfittingفرط التخصيص