Language Models2022intermediate11 min read
Emergent Abilities of Large Language Models
القدرات الناشئة في النماذج اللغوية الكبيرة
Wei, J. · Tay, Y. · Bommasani, R. · Raffel, C. · Zoph, B. · Borgeaud, S. · Yogatama, D. · Bosma, M. · Zhou, D. · Metzler, D. · Chi, E. H. · Hashimoto, T. · Vinyals, O. · Liang, P. · Dean, J. · Fedus, W. — TMLR
The problem
By 2022, had established that increasing model size, data, and compute predictably reduces language model loss. But practitioners kept noticing something the smooth loss curves could not explain: certain capabilities — multi-digit arithmetic, multi-step reasoning, transliteration — appeared to be completely absent below a threshold scale and then suddenly present above it. There was no systematic study of these abrupt transitions, no shared vocabulary for discussing them, and no catalog of which tasks exhibited this pattern.
The contribution
A formal definition of "" — a capability absent in smaller models that appears in larger ones, making it unpredictable by extrapolating small-scale performance. The paper catalogs dozens of such abilities across five model families (GPT-3, LaMDA, PaLM, Chinchilla, Gopher), organized into few-shot prompted tasks (, , TruthfulQA) and augmented prompting strategies (chain-of-thought, , ). It draws a parallel to phase transitions in physics and argues that further scaling may unlock still-unknown capabilities.
The impact
This paper gave the AI community a shared framework — the concept of "emergence" — to discuss why scaling produces qualitative, not just quantitative, improvements. It directly motivated the race to build ever-larger models and focused safety research on the unpredictability of future capabilities. The paper also sparked an important counter-debate: Schaeffer et al. (2023) argued that many are measurement artifacts caused by discontinuous metrics, not genuine phase transitions — a debate that continues to shape how the field evaluates and interprets model scaling.
Imagine teaching a child to swim. At first, they flounder no matter how much practice they get — splashing, sinking, going nowhere. You might conclude they simply cannot swim. Then one day, something clicks: arms coordinate with legs, breathing finds a rhythm, and they glide across the pool. The skill did not improve gradually; it was absent and then present.
Large language models do the same thing. A model with 1 billion parameters cannot add three-digit numbers — it guesses randomly. A model with 10 billion parameters still cannot. But at 100 billion parameters, it suddenly can. This paper documents dozens of these "click moments" and asks whether they are genuine phase transitions or just the natural consequence of passing a hidden threshold.
What is an emergent ability?
The concept of emergence comes from physics. When water molecules are heated, each molecule moves a little faster — a smooth, predictable change. But at 100°C, the system undergoes a : liquid becomes gas. The collective behavior is qualitatively different, even though the underlying change (more kinetic energy) was gradual.
Wei et al. borrow this idea for language models. They define an emergent ability as one that is "not present in smaller models but is present in larger models." Crucially, this means the ability cannot be predicted by extrapolating the performance curve of smaller models. If you plotted performance against scale for a non-emergent ability, you would see a smooth upward curve. For an emergent ability, you see a flat line at random-chance performance, then a sharp jump.
The authors measure scale along two axes: the number of model parameters and the total compute in (floating-point operations). Using FLOPs is important because it captures both model size and how long the model was trained — a smaller model trained longer can match a larger model trained less.
Few-shot prompting: abilities that appear from nowhere
The first category of emergence the paper examines is . In this paradigm, popularized by GPT-3, you give the model a natural language instruction plus a handful of input-output examples, then ask it to perform the task on a new input — with no gradient updates or .
The key finding: for many tasks, small models perform at or near random chance no matter how good the prompt is. Then, above a critical scale, performance jumps sharply. The paper documents this pattern across diverse tasks from two major benchmarks.
BIG-Bench — a crowd-sourced suite of over 200 tasks — yielded striking examples. Three-digit addition and subtraction shows near-zero for GPT-3 below 13B parameters ( FLOPs), then jumps to high accuracy. LaMDA shows the same pattern but at 68B parameters ( FLOPs). Other emergent BIG-Bench tasks include transliterating from the International Phonetic Alphabet, recovering scrambled words, and Persian question-answering.
MMLU (Massive Multitask Language Understanding) — a 57-subject — showed that different subject areas become emergent at different scales. Humanities and Social Sciences were the most emergent categories, appearing above random only at the largest model sizes. The pattern was consistent across GPT-3, Chinchilla, and Gopher.
A third benchmark, TruthfulQA, revealed an initially counterintuitive pattern: larger Gopher models actually performed worse on truthfulness, generating more confident but incorrect answers. Only at the very largest scale did performance recover. This suggests that emergence is not always a simple story of "bigger is better" — some abilities may temporarily degrade during the scaling process before ultimately emerging.
Augmented prompting: strategies that only work at scale
The second category goes beyond individual tasks. Some prompting strategies themselves are emergent — they hurt small models but help large ones.
asks the model to show its reasoning step by step before giving a final answer. For the LaMDA model family, chain-of-thought prompting only surpasses standard prompting above FLOPs (roughly 100B parameters). Below that threshold, asking the model to reason step-by-step actually hurts performance compared to direct answering. The model is too small to maintain coherent multi-step reasoning, so the extra steps introduce errors rather than helping.
Instruction following shows a similar pattern. Finetuning on a mixture of tasks phrased as instructions (the FLAN approach) enables generalization to unseen tasks — but only above about 68B parameters. Below that scale, instruction finetuning actually degrades performance relative to the base model.
Scratchpad training, where the model writes out intermediate computation steps (like carrying digits in addition), only helps larger models. For smaller LaMDA models, the scratchpad adds noise rather than structure.
Why does emergence happen?
The paper is deliberately open about this: we do not know. The authors present the empirical pattern but do not provide a mechanistic explanation. They offer several hypotheses worth considering.
One possibility is that emergent abilities require a constellation of sub-skills. A model needs to simultaneously master , syntax, semantics, world knowledge, and task format — and only when all of these cross individual quality thresholds does the combined ability appear. Think of it like a combination lock: each dial moving closer to the right number does nothing visible until all dials click into place at once.
Another hypothesis involves the relationship between pretraining loss and task performance. While loss decreases smoothly, the mapping from loss to performance on a specific task may be highly nonlinear. A small improvement in loss near a critical region could correspond to a large jump in task accuracy — especially for tasks with sharp success criteria like exact-match accuracy.
The paper does not resolve which explanation is correct. This open question became one of the most active research areas following publication.
The counter-argument: emergence as a measurement artifact
In 2023, Schaeffer, Miranda, and Koyejo published a provocative response: Are Emergent Abilities of Large Language Models a Mirage? Their argument is elegant and important.
They observed that almost all "emergent" tasks were measured with discontinuous metrics — metrics like exact-match accuracy that give zero credit for partial answers. Under such metrics, a model that gets 4 out of 5 digits right in a math problem scores identically to a model that guesses randomly. The result: apparent flat-line performance even though the model is improving steadily.
When Schaeffer et al. re-evaluated the same tasks with continuous metrics (like token-level edit distance or per-digit accuracy), the sharp transitions vanished. Performance improved smoothly and predictably with scale. Over 92% of emergent abilities on BIG-Bench disappeared when evaluated with linear metrics.
Their conclusion: emergence may not be a property of the model but a property of the measurement. The model's underlying capability improves smoothly, but a harsh metric creates the illusion of a sudden jump — like measuring a high-jumper's ability as pass/fail on a fixed bar height rather than tracking how high they actually jump.
Implications for AI safety and the future
Whether emergence is fundamental or partly a measurement artifact, the practical implications are profound.
For model developers: if truly emergent abilities exist, they cannot be predicted by evaluating smaller models. This means you cannot fully characterize a model's capabilities until it is built and tested — a deeply uncomfortable position for safety planning. The only way to know what a 10-trillion-parameter model can do is to build it.
For AI safety: unpredictable capability acquisition is a core concern. If models can gain abilities that were not anticipated, they could also gain dangerous capabilities — like generating convincing misinformation or exploiting security vulnerabilities — without warning. The paper explicitly notes that emergence could include harmful abilities.
For benchmarks: the debate about metrics reveals that how we measure model capabilities matters as much as what we measure. Coarse binary metrics may hide gradual improvement, while fine-grained metrics may miss qualitative transitions. The field needs evaluation frameworks that capture both.
The paper closes with an observation that doubles as a prediction: "Further scaling of language models could potentially expand the range of capabilities even further." As of 2024 and beyond, that prediction has been borne out repeatedly.
Timeline: emergence in context
2020
GPT-3: the first glimpse of emergence
Brown et al. showed that a 175B-parameter model could perform few-shot tasks that smaller GPT models could not — arithmetic, translation, code generation. The term "emergent" was not yet used, but the phenomenon was visible.
2022
Chinchilla: compute-optimal scaling
Hoffmann et al. showed that many large models were undertrained. Chinchilla (70B parameters, trained on more data) matched Gopher (280B). This reframed emergence in terms of training compute, not just parameter count.
2022
This paper — Emergent Abilities (Wei et al.)
First systematic catalog of emergent abilities across five model families and dozens of tasks. Introduced the formal definition and the physics analogy to phase transitions. Published in TMLR.
2022
PaLM: emergence at unprecedented scale
Google's 540B-parameter model showed new emergent abilities on BIG-Bench that even GPT-3 175B could not achieve, including joke explanation and logical deduction.
2023
The Mirage paper (Schaeffer et al.)
Argued that emergent abilities are measurement artifacts from discontinuous metrics. Won an Outstanding Paper award at NeurIPS 2023. The debate between "real emergence" and "metric mirage" remains active.
The emergence debate illuminated a deeper truth about AI research: what we observe depends on how we measure. Whether emergent abilities are phase transitions or metric artifacts, the practical reality is the same — scaling language models produces capabilities that were not anticipated, cannot be easily predicted, and demand careful evaluation. The paper's lasting contribution is not a definitive answer about emergence but a framework for asking the right questions about what scaling does and does not tell us.
CitationWei, Tay, Bommasani, Raffel, Zoph, Borgeaud, Yogatama, Bosma, Zhou, Metzler, Chi, Hashimoto, Vinyals, Liang, Dean, Fedus. Emergent Abilities of Large Language Models. Transactions on Machine Learning Research (TMLR), 2022.
Terms in this paper
- Emergent Abilitiesالقدرات المعرفية الناشئة فجأة
- Scaling Lawsقوانين التوسعة
- Few-Shot Promptingالتحفيز بأمثلة قليلة
- In-Context Learningالتعلم في السياق
- Chain-of-Thought Promptingتلقين سلسلة التفكير
- BIG-BenchBIG-Bench
- MMLUMMLU
- FLOPsالعمليات الحسابية العائمة
- Model Scaleحجم النموذج
- Prompt Engineeringهندسة التحفيز
- Instruction Followingاتباع التعليمات
- Phase Transitionالتحوّل الطوري
- Benchmarkالمعيار المرجعي
- Perplexityمعيار الحيرة الاحتمالية
- Few-Shot Learningالتعلّم بأمثلة قليلة