AI Safety2023intermediate11 min read
Are Emergent Abilities of Large Language Models a Mirage?
هل القدرات الناشئة في النماذج اللغوية الكبيرة مجرد سراب؟
Schaeffer, R. · Miranda, B. · Koyejo, S. — NeurIPS
The problem
By 2022, researchers had observed that large language models seem to acquire new abilities abruptly — arithmetic, multi-step reasoning, translation — at unpredictable scales. These "" raised urgent questions for : if capabilities appear without warning, how can we anticipate dangerous ones? But the evidence for emergence rested almost entirely on two metrics — and — both of which impose sharp thresholds on model outputs. Nobody had systematically asked whether the appearance of emergence was a property of the models or a property of the measurement.
The contribution
The authors demonstrate that emergent abilities are a measurement artifact, not a fundamental property of scaling. They show mathematically that nonlinear metrics (like Accuracy) and discontinuous metrics (like Multiple Choice Grade) can transform smooth, predictable improvements in per- error rates into sharp, seemingly unpredictable jumps. They confirm this with three lines of evidence: (1) switching metrics on GPT-3 arithmetic tasks makes emergence disappear, (2) a meta-analysis of reveals that over 92% of claimed emergent abilities use just two metrics, and (3) they deliberately induce fake "emergence" in vision models by choosing sharp metrics.
The impact
This paper fundamentally changed how the AI community interprets scaling results. It showed that the dramatic "emergent abilities" narrative — central to both AI hype and AI safety arguments — may be an artifact of poor metric choice. The work prompted researchers to report continuous metrics alongside discrete ones, increased scrutiny of design, and tempered claims about unpredictable capability jumps. It remains one of the most cited critiques of scaling-era LLM evaluation.
Imagine a high-jump bar. An athlete improving from 1.50 m to 2.10 m by 5 cm each month shows steady progress — but if you only record "cleared the bar" or "didn't" at a fixed height of 2.00 m, the athlete appears to have zero ability for months, then "suddenly" succeeds. The jump didn't change — the scorecard did.
This paper shows that the celebrated "emergent abilities" of large language models work exactly like that bar. The models improve smoothly with scale. The metric — accuracy, exact match — is the bar, and it turns gradual progress into a dramatic-looking leap.
The claim: abilities appear from nowhere
In 2022, Wei et al. coined the term "emergent abilities of large language models" to describe abilities that are absent in smaller models but present in larger ones. Two properties made this idea exciting and alarming:
Sharpness — the transition from "can't do it" to "can do it" seems instantaneous, not gradual. A model with 10 billion parameters scores near zero on arithmetic, but a model with 100 billion parameters scores near perfect.
Unpredictability — there is no way to foresee at which scale an ability will appear. You cannot extrapolate from smaller models to predict when the jump will happen.
These two properties together created an urgent narrative: if models can acquire dangerous capabilities without warning, how do we ensure safety? But Schaeffer et al. noticed something suspicious: the evidence for emergence rested almost entirely on metrics that impose hard thresholds on model outputs.
The alternative: it is the metric, not the model
The paper's core argument is elegant. Suppose that within a model family, the per-token cross-entropy loss falls smoothly as a power law with model size — a well-documented phenomenon known as neural scaling laws. Then the per-token probability of selecting the correct token also improves smoothly. So far, no emergence.
Now suppose the researcher chooses Accuracy as the metric: a model's output is scored 1 only if every token in the output is correct, and 0 otherwise. If the output is tokens long, the probability of getting a perfect score is approximately the per-token probability raised to the power . This exponential relationship is what creates the illusion of a sharp jump — the metric nonlinearly amplifies small improvements into an apparent phase transition.
Similarly, Multiple Choice Grade acts as a step function: 1 if the model's highest probability lands on the correct option, 0 otherwise. This discontinuous metric creates a cliff where a smooth probability shift crosses a decision boundary.
The fix is straightforward: use a linear or continuous metric instead. counts how many tokens differ between the model's output and the target — it scales approximately linearly with the per-token error rate. measures the mean squared error between predicted probabilities and outcomes — it is continuous and strictly proper. Under either of these metrics, the same model outputs that looked "emergent" under Accuracy reveal smooth, predictable improvement.
Evidence 1: GPT-3 arithmetic — same outputs, different story
The GPT-3 family was prominently cited as displaying emergent arithmetic abilities. On 2-digit multiplication and 4-digit addition, smaller models score near zero on Accuracy while the largest model (175B parameters) scores dramatically higher — a classic "emergent" pattern.
The authors made three predictions and confirmed all three:
Prediction 1: Switching from Accuracy to Token Edit Distance should reveal smooth improvement. It did — the "emergence" vanished entirely. Under Token Edit Distance, all models improved predictably with scale.
Prediction 2: Even under Accuracy, increasing the test set size should reveal that small models perform above chance. It did — with more test data, even the 350M-parameter model showed non-zero accuracy, consistent with the mathematical prediction of geometrically decaying accuracy.
Prediction 3: Increasing the target length should predictably reduce Accuracy (geometrically) and Token Edit Distance (quasi-linearly). Both patterns matched.
Evidence 2: the BIG-Bench meta-analysis
If emergent abilities are real, they should appear regardless of which reasonable metric you use. If they are a metric artifact, they should cluster around specific metrics. The authors analyzed all task-metric-model-family triplets in BIG-Bench and found a striking pattern:
Of the 39 preferred metrics in BIG-Bench, at most 5 produce any emergent abilities. And two metrics alone — Multiple Choice Grade and — account for over 92% of all claimed emergent abilities. Multiple Choice Grade is discontinuous (a step function). Exact String Match is nonlinear (requires every token correct).
When the authors re-evaluated the LaMDA model family — which showed emergence under Multiple Choice Grade — using the continuous Brier Score instead, the emergent abilities disappeared completely. The model outputs did not change. Only the metric changed.
Evidence 3: manufacturing emergence on demand
The strongest evidence is constructive: the authors deliberately induced "emergent" abilities in vision models — architectures where no one had ever observed emergence before.
Autoencoders on CIFAR-100: Shallow autoencoders display smoothly decreasing reconstruction error as model size grows. But define a new discontinuous metric — "Reconstruction ability: 1 if reconstruction error is below threshold , 0 otherwise" — and a sharp, unpredictable-looking transition appears.
Transformers on Omniglot: Transformers classifying handwritten characters show smoothly increasing single-image accuracy. But redefine accuracy as "correctly classifying all images in a sequence" and an "emergent" classification ability appears for larger .
LeNet on MNIST: Similarly, convolutional networks show smooth per-image accuracy on MNIST, but redefining accuracy as "correctly classifying all independent test images" produces an apparent emergence.
In every case, the model's actual behavior improved smoothly. The researcher manufactured the emergence by choosing the metric.
The formal picture: how metrics distort scaling curves
The paper provides a clean mathematical model connecting per-token cross-entropy loss to observed metrics. Start with the assumption that per-token cross-entropy falls as a power law with model size :
The four key metrics behave as follows when applied to the same underlying improvement:
Accuracy (nonlinear): — the per-token probability raised to the target length . Geometric decay creates sharp transitions.
Token Edit Distance (linear): — counts individual token errors. Smooth and predictable.
Multiple Choice Grade (discontinuous): a step function at the decision boundary. Tiny probability shifts near the boundary cause 0→1 jumps.
Brier Score (continuous): mean squared error of predicted probabilities. Smooth by construction.
Why this matters: benchmarks, safety, and scientific practice
The implications ripple across the AI landscape:
For benchmarking: Tasks and metrics are distinct choices. A task measures an ability; a metric measures how you score that ability. Conflating them — concluding "the model can't do arithmetic" because Accuracy is zero — is a category error. Researchers should always report continuous metrics alongside discrete ones.
For AI safety: Much of the urgency around is driven by the fear that dangerous capabilities will appear "without warning." If emergence is a metric artifact, then capabilities are actually predictable — they grow smoothly with scale. This doesn't make safety less important, but it changes the threat model from "unpredictable jumps" to "predictable growth that we need to monitor with the right measurements."
For scientific rigor: When models and their outputs are not made publicly available, independent verification is impossible. The authors emphasize that scientific progress is hampered when the community cannot reanalyze published claims with different metrics.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def per_token_prob(N, c=1e10, alpha=-0.5):
"""Per-token probability of correct token, from scaling law."""
cross_entropy = (N / c) ** alpha
return np.exp(-cross_entropy)
def accuracy(p_token, L):
"""Nonlinear metric: all L tokens must be correct."""
return p_token ** L # Geometric — creates sharp transitions
def token_edit_distance(p_token, L):
"""Linear metric: count of incorrect tokens."""
return L * (1 - p_token) # Linear — smooth improvement
# Compare: same model outputs, two metrics
model_sizes = np.logspace(8, 11.5, 50) # 100M to 300B parameters
for L in [1, 3, 5]:
p = per_token_prob(model_sizes)
acc = accuracy(p, L) # Looks "emergent" for L >= 3
ted = token_edit_distance(p, L) # Always looks smooth
# The lesson: the model's actual improvement (p) is always smooth.
# Accuracy raises p to the power L, creating a sharp cliff.
# Token Edit Distance scales linearly with (1-p), staying smooth.Context: the emergence debate
2020
Neural scaling laws (Kaplan et al.)
Showed that loss decreases as a smooth power law with model size, dataset size, and compute — establishing predictability as the default expectation.
2022
"Emergent Abilities of LLMs" (Wei et al.)
Coined the term and documented sharp, unpredictable jumps in performance across multiple model families and tasks, sparking intense interest and concern.
2022
BIG-Bench (Srivastava et al.)
A massive collaborative benchmark of 200+ tasks. Noted that while accuracy shows sharp transitions, cross-entropy does not — planting the seed of the metric hypothesis.
2023
This paper (Schaeffer, Miranda, Koyejo)
Provided the mathematical model, empirical evidence, and constructive demonstrations that emergent abilities are a metric artifact, not a model property.
2023
Post-mirage evaluation reforms
Growing adoption of continuous metrics in LLM evaluation. Researchers increasingly report per-token metrics, calibration scores, and log-probability alongside accuracy.
CitationSchaeffer, Miranda, Koyejo. Are Emergent Abilities of Large Language Models a Mirage?. NeurIPS, 2023.
Terms in this paper
- Emergent Abilitiesالقدرات المعرفية الناشئة فجأة
- Evaluation Metricمعيار قياس الأداء
- Scaling Lawقانون التحجيم
- Accuracyنسبة الدقة الإجمالية
- Cross Entropyالعشوائية المتقاطعة
- Large Language Modelالنموذج اللغوي الكبير
- Benchmarkالمعيار المرجعي
- Few-Shotالنمط القليل العيّنات
- Brier Scoreدرجة براير
- Token Edit Distanceمسافة تعديل الرموز
- Multiple Choice Gradeدرجة الاختيار المتعدد
- Exact String Matchالمطابقة التامة للسلسلة
- BIG-BenchBIG-Bench
- Model Scaleحجم النموذج
- Perplexityمعيار الحيرة الاحتمالية