Language Models2020intermediate11 min read

Scaling Laws for Neural Language Models

قوانين التوسُّع في النماذج اللغوية العصبية

Kaplan, J. · McCandlish, S. · Henighan, T. · Brown, T. B. · Chess, B. · Child, R. · Gray, S. · Radford, A. · Wu, J. · Amodei, D. — arXiv

The problem

By 2020, deep learning had produced increasingly capable language models — GPT-2 had just shown that larger models produce qualitatively better text. But scaling up was a gamble: nobody knew how much better a model would get if you doubled its size, or whether to spend your on a bigger model or more data. runs cost millions of dollars, and there was no principled way to predict the outcome before committing those resources.

The contribution

Empirical scaling laws — precise power-law equations relating language model to three variables: model size N, dataset size D, and compute budget C. The loss follows L ∝ N^(−0.076), L ∝ D^(−0.095), and L ∝ C^(−0.050) over seven orders of magnitude. These laws reveal that performance depends overwhelmingly on scale (not architecture details like depth vs. width), that larger models are more sample-efficient, and that training means training very large models on relatively little data, stopping well before .

The impact

Transformed AI scaling from guesswork into engineering. These equations directly guided the training of GPT-3 (175B parameters) and shaped every subsequent scaling decision in the field. Chinchilla later refined the compute-optimal tradeoff, and GPT-4 used scaling laws to predict performance from small proxies. The paper established that "bigger is predictably better" — the intellectual foundation for the large language model era.

Imagine you're a city planner deciding how tall to build a skyscraper. You don't just guess — you know that adding 10 floors costs a predictable amount and adds a predictable number of offices. You can compute the optimal height before laying a single brick.

Before this paper, training an AI model was like building a skyscraper blindfolded: you poured concrete and hoped for the best. Kaplan et al. discovered the architect's equations — simple power laws that predict exactly how much "smarter" a language model will get as you make it bigger, feed it more data, or train it longer. Suddenly, scaling wasn't a gamble — it was engineering.

The three knobs: parameters, data, and compute

When you train a language model, three resources determine how good it gets:

Model size (N) — the number of learnable parameters (weights) in the network, excluding embeddings. A 100M- model is a small brain; a 175B-parameter model is a vastly larger one.

Dataset size (D) — how many tokens the model sees during training. More tokens means more patterns to learn from.

Compute (C) — the total floating-point operations spent on training, roughly C≈6NBSC \approx 6NBS where BB is the and SS is the number of training steps.

The breakthrough insight is that these three variables are all you need. Architectural details like depth vs. width, number of heads, or feed-forward dimension barely matter once you hold the total parameter count fixed.

Open in Lab
Each panel shows loss vs. one scaling variable on a log-log plot. The straight lines are power laws — notice how they hold over many orders of magnitude.
The demo wakes as you arrive…

Power laws: straight lines on log-log paper

The paper's central discovery is that when you plot loss against any of the three variables on a log-log scale, you get a straight line. A straight line on a log-log plot means a — the relationship L∝X−αL \propto X^{-\alpha}, where α\alpha is the slope.

Think of it this way: every time you multiply the model size by 10, the loss drops by a fixed percentage. Not a fixed amount — a fixed fraction. This pattern holds whether you go from 1,000 parameters to 10,000 or from 100 million to 1 billion — the same proportional improvement each time.

Why does this matter? Because power laws are predictable. If you've measured the trend on small models, you can extrapolate to large ones with confidence. Training a small model is cheap; predicting how a large model will perform before you build it is invaluable.

The three fundamental equations

The paper fits three core power laws. Each describes what happens when one variable is the bottleneck and the other two are abundant. Before seeing the formulas, here is the intuition: imagine filling a bathtub. The water level (performance) is limited by whichever pipe (parameter count, data, compute) carries the least flow. The power law tells you how much the water rises when you widen that pipe.

L(N)=(NcN)αN,αN≈0.076L(N) = \left(\frac{N_c}{N}\right)^{\alpha_N}, \quad \alpha_N \approx 0.076
Scaling with model size — When data and compute are unlimited, loss falls as a power law with the number of non-embedding parameters. Doubling N reduces loss by about 5%.
L(D)=(DcD)αD,αD≈0.095L(D) = \left(\frac{D_c}{D}\right)^{\alpha_D}, \quad \alpha_D \approx 0.095
Scaling with dataset size — When the model is large enough and well trained, loss falls as a power law with the number of training tokens. The exponent is slightly larger than for model size, meaning data is slightly more "efficient" at reducing loss.
L(Cmin⁡)=(Ccmin⁡Cmin⁡)αCmin⁡,αCmin⁡≈0.050L(C_{\min}) = \left(\frac{C_c^{\min}}{C_{\min}}\right)^{\alpha_C^{\min}}, \quad \alpha_C^{\min} \approx 0.050
Scaling with compute budget — When compute is optimally allocated between model size and training duration, loss falls as a power law with the total FLOPs. This is the most practically useful equation: it lets you predict performance from a dollar amount.
Open in Lab
Adjust the exponent α to see how steeper power laws mean faster improvement. The paper's measured exponents are shown as presets.
The demo wakes as you arrive…

Architecture doesn't matter (much)

One of the most surprising findings: when you fix the total parameter count, the model's shape barely affects performance. The ratio of depth to width can vary by a factor of 40 — from a very deep, narrow model to a very wide, shallow one — with only a 3% change in loss.

This was a radical claim. Before this paper, architecture search was a major research area. The paper suggested that all that effort was optimizing the wrong thing: instead of finding the perfect shape, you should simply be making models bigger.

The one exception: excluding parameters from the count N gives much cleaner trends. This suggests embeddings are somewhat "free" capacity that doesn't contribute to the core scaling behavior.

Open in Lab
Vary depth and width while keeping total parameters fixed. Notice how loss barely changes across a wide range of shapes.
The demo wakes as you arrive…

Overfitting: when the model outgrows its data

What happens when you make the model larger but keep the dataset fixed? Eventually the model memorizes the data instead of learning patterns — this is .

The paper discovered a remarkably clean relationship: the degree of overfitting depends on a single ratio N0.74/DN^{0.74}/D. This means that every time you make the model 8× bigger, you only need roughly 5× more data to avoid performance penalties. Data requirements grow sub-linearly with model size.

This is great news for scaling: you don't need to proportionally increase your dataset every time you build a bigger model. The model becomes increasingly sample-efficient — it extracts more knowledge from each .

L(N,D)=[(NcN)αN/αD+DcD]αDL(N, D) = \left[\left(\frac{N_c}{N}\right)^{\alpha_N/\alpha_D} + \frac{D_c}{D}\right]^{\alpha_D}
Joint scaling with model size and data — This single equation unifies both power laws. When D is huge, it reduces to L(N). When N is huge, it reduces to L(D). In between, it precisely predicts the overfitting penalty.
Open in Lab
Explore the L(N, D) surface. The diagonal shows the compute-efficient frontier where model size and data are balanced.
The demo wakes as you arrive…

The budget question: how to spend your compute

This is the paper's most practically important result. Given a fixed compute budget CC, how should you distribute it between a bigger model and more training time?

The answer is dramatic: most of the compute should go to model size. As compute increases by 1000×, the optimal model size grows by about 500× (exponent 0.73), while the number of training steps barely increases (exponent 0.03). The data requirement grows slowly too — only about 7× (exponent 0.27).

In practical terms, this means compute-efficient training involves training very large models for a short time, stopping well before convergence. This is counterintuitive: most researchers in 2020 were training smaller models to full convergence. The paper argued they were wasting compute.

Open in Lab
Slide to increase the compute budget and watch how it should be allocated. Notice that model size absorbs almost all the increase.
The demo wakes as you arrive…

Bigger models learn faster

A consistently striking finding: larger models are dramatically more sample-efficient. A 1-billion parameter model reaches the same loss as a 10-million parameter model while seeing roughly 100× fewer data examples.

This creates a virtuous cycle for scaling: not only do larger models converge to lower loss, they also get there faster in terms of data. The optimal strategy is not "collect as much data as possible and train a big model on all of it," but rather "build the biggest model you can afford and train it on a modest amount of data."

The learning curves themselves follow a power law: L(N,S)=(Nc/N)αN+(Sc/S)αSL(N, S) = (N_c/N)^{\alpha_N} + (S_c/S)^{\alpha_S} where SS is the number of training steps. This means you can predict the final loss from the first few percent of training — a powerful tool for deciding whether to continue an expensive run.

Open in Lab
Watch models of different sizes race to the same loss target. The largest model gets there first with far fewer data points.
The demo wakes as you arrive…

The critical batch size

There is an optimal batch size for training, and it too follows a power law — but remarkably, it depends only on the loss value, not on the model size. The roughly doubles for every 13% decrease in loss, following Bcrit≈B∗/L1/αBB_{crit} \approx B^*/L^{1/\alpha_B} with αB≈0.21\alpha_B \approx 0.21.

Training below the critical batch size wastes time (too many serial steps). Training above it wastes compute (each step gives diminishing returns). The sweet spot, training at the critical batch size, makes the best use of both.

For the largest models approaching convergence, the critical batch size is roughly 1–2 million tokens — which maps to the large-batch training configurations that became standard practice in subsequent work.

Transformers vs. LSTMs: the architecture that scales

The paper includes a crucial comparison between Transformers and LSTMs. Both follow power-law scaling with parameters, but Transformers asymptotically outperform LSTMs — and the gap widens as models grow.

The reason is context utilization. LSTMs plateau after about 100 tokens of context: adding more context doesn't help. Transformers, by contrast, continue to improve throughout their entire 1024-token . Each additional word of context makes the 's predictions meaningfully better.

This finding provided quantitative evidence for what the Transformer paper had argued qualitatively: 's direct word-to-word connections are not just architecturally elegant — they are the mechanism that enables scaling.

The ceiling: where the laws must break down

The authors noticed a deep tension in their own equations. Compute-efficient training grows data very slowly — D∝C0.27D \propto C^{0.27}. But avoiding overfitting requires D∝N0.74∝C0.54D \propto N^{0.74} \propto C^{0.54}. These two trends eventually contradict: at some point, compute-efficient training won't have enough data to avoid overfitting, even if you never reuse a single token.

The authors estimated this intersection at roughly 101210^{12} parameters and 101210^{12} tokens, with a loss of about 1.7 nats/token. They conjectured this might represent the inherent of natural language — an irreducible floor below which no model can go, no matter how large.

This prediction was remarkably prescient. Modern frontier models have indeed approached the trillion-parameter, trillion-token scale, and the quest for data has become one of the field's most pressing challenges.

Open in Lab
Watch the two data-growth curves diverge. The gap between what compute-efficient training uses and what overfitting prevention requires eventually becomes unsustainable.
The demo wakes as you arrive…

Generalization transfers predictably

When the authors tested their WebText2-trained models on completely different text distributions — Wikipedia, Books, Common Crawl — they found a consistent pattern: the loss on other distributions follows the same power law with model size, just shifted by a constant offset.

This means that improving on the training distribution translates directly to improvement on other distributions. A model doesn't just memorize WebText2 patterns — it learns general language understanding that transfers. Furthermore, this gap doesn't depend on training duration or model depth; it depends only on the validation loss achieved.

What these equations unlocked

  1. 2020

    Scaling Laws (this paper)

    Established power-law relationships between loss and scale across seven orders of magnitude, transforming AI scaling from guesswork into predictive engineering.

  2. 2020

    GPT-3 — scaling laws in action

    Brown et al. trained a 175B-parameter model, guided by scaling law predictions. The result exhibited emergent few-shot abilities that nobody had programmed.

  3. 2022

    Chinchilla — refining the optimal tradeoff

    Hoffmann et al. showed that Kaplan's allocation was undertrained: data should scale equally with parameters. The framework held; only the exponents shifted.

  4. 2022

    Emergent abilities observed at scale

    Wei et al. documented capabilities that appear suddenly above certain model sizes — arithmetic, multi-step reasoning, instruction following — raising questions about what smooth power laws can and cannot predict.

  5. 2023

    GPT-4 — predictable scaling validated

    OpenAI reported that scaling laws fitted on small models accurately predicted GPT-4's final performance before the full training run completed.

  6. 2023

    LLaMA — scaling for inference efficiency

    Touvron et al. extended Chinchilla-optimal training far beyond 20 tokens per parameter, trading training compute for inference efficiency — a new dimension of scaling.

Kaplan et al. didn't just discover power laws — they created a predictive science of AI scaling. Before this paper, training a frontier model was an act of faith. After it, teams could write down equations, plug in their compute budget, and predict within a few percent what performance they would achieve. The specific numbers have been refined — the Chinchilla paper corrected the data/model balance — but the paradigm of "fit power laws on small runs, extrapolate to large ones" is now standard practice at every AI lab.

CitationKaplan, McCandlish, Henighan, Brown, Chess, Child, Gray, Radford, Wu, Amodei. Scaling Laws for Neural Language Models. arXiv, 2020.

Terms in this paper