Language Models2020intermediate11 min read
Scaling Laws for Neural Language Models
قوانين التوسُّع في النماذج اللغوية العصبية
Kaplan, J. · McCandlish, S. · Henighan, T. · Brown, T. B. · Chess, B. · Child, R. · Gray, S. · Radford, A. · Wu, J. · Amodei, D. — arXiv
The problem
By 2020, deep learning had produced increasingly capable language models — GPT-2 had just shown that larger models produce qualitatively better text. But scaling up was a gamble: nobody knew how much better a model would get if you doubled its size, or whether to spend your on a bigger model or more data. runs cost millions of dollars, and there was no principled way to predict the outcome before committing those resources.
The contribution
Empirical scaling laws — precise power-law equations relating language model to three variables: model size N, dataset size D, and compute budget C. The loss follows L ∝ N^(−0.076), L ∝ D^(−0.095), and L ∝ C^(−0.050) over seven orders of magnitude. These laws reveal that performance depends overwhelmingly on scale (not architecture details like depth vs. width), that larger models are more sample-efficient, and that training means training very large models on relatively little data, stopping well before .
The impact
Transformed AI scaling from guesswork into engineering. These equations directly guided the training of GPT-3 (175B parameters) and shaped every subsequent scaling decision in the field. Chinchilla later refined the compute-optimal tradeoff, and GPT-4 used scaling laws to predict performance from small proxies. The paper established that "bigger is predictably better" — the intellectual foundation for the large language model era.
Imagine you're a city planner deciding how tall to build a skyscraper. You don't just guess — you know that adding 10 floors costs a predictable amount and adds a predictable number of offices. You can compute the optimal height before laying a single brick.
Before this paper, training an AI model was like building a skyscraper blindfolded: you poured concrete and hoped for the best. Kaplan et al. discovered the architect's equations — simple power laws that predict exactly how much "smarter" a language model will get as you make it bigger, feed it more data, or train it longer. Suddenly, scaling wasn't a gamble — it was engineering.
The three knobs: parameters, data, and compute
When you train a language model, three resources determine how good it gets:
Model size (N) — the number of learnable parameters (weights) in the network, excluding embeddings. A 100M- model is a small brain; a 175B-parameter model is a vastly larger one.
Dataset size (D) — how many tokens the model sees during training. More tokens means more patterns to learn from.
Compute (C) — the total floating-point operations spent on training, roughly where is the and is the number of training steps.
The breakthrough insight is that these three variables are all you need. Architectural details like depth vs. width, number of heads, or feed-forward dimension barely matter once you hold the total parameter count fixed.
Power laws: straight lines on log-log paper
The paper's central discovery is that when you plot loss against any of the three variables on a log-log scale, you get a straight line. A straight line on a log-log plot means a — the relationship , where is the slope.
Think of it this way: every time you multiply the model size by 10, the loss drops by a fixed percentage. Not a fixed amount — a fixed fraction. This pattern holds whether you go from 1,000 parameters to 10,000 or from 100 million to 1 billion — the same proportional improvement each time.
Why does this matter? Because power laws are predictable. If you've measured the trend on small models, you can extrapolate to large ones with confidence. Training a small model is cheap; predicting how a large model will perform before you build it is invaluable.
The three fundamental equations
The paper fits three core power laws. Each describes what happens when one variable is the bottleneck and the other two are abundant. Before seeing the formulas, here is the intuition: imagine filling a bathtub. The water level (performance) is limited by whichever pipe (parameter count, data, compute) carries the least flow. The power law tells you how much the water rises when you widen that pipe.
Architecture doesn't matter (much)
One of the most surprising findings: when you fix the total parameter count, the model's shape barely affects performance. The ratio of depth to width can vary by a factor of 40 — from a very deep, narrow model to a very wide, shallow one — with only a 3% change in loss.
This was a radical claim. Before this paper, architecture search was a major research area. The paper suggested that all that effort was optimizing the wrong thing: instead of finding the perfect shape, you should simply be making models bigger.
The one exception: excluding parameters from the count N gives much cleaner trends. This suggests embeddings are somewhat "free" capacity that doesn't contribute to the core scaling behavior.
Overfitting: when the model outgrows its data
What happens when you make the model larger but keep the dataset fixed? Eventually the model memorizes the data instead of learning patterns — this is .
The paper discovered a remarkably clean relationship: the degree of overfitting depends on a single ratio . This means that every time you make the model 8× bigger, you only need roughly 5× more data to avoid performance penalties. Data requirements grow sub-linearly with model size.
This is great news for scaling: you don't need to proportionally increase your dataset every time you build a bigger model. The model becomes increasingly sample-efficient — it extracts more knowledge from each .
The budget question: how to spend your compute
This is the paper's most practically important result. Given a fixed compute budget , how should you distribute it between a bigger model and more training time?
The answer is dramatic: most of the compute should go to model size. As compute increases by 1000×, the optimal model size grows by about 500× (exponent 0.73), while the number of training steps barely increases (exponent 0.03). The data requirement grows slowly too — only about 7× (exponent 0.27).
In practical terms, this means compute-efficient training involves training very large models for a short time, stopping well before convergence. This is counterintuitive: most researchers in 2020 were training smaller models to full convergence. The paper argued they were wasting compute.
Bigger models learn faster
A consistently striking finding: larger models are dramatically more sample-efficient. A 1-billion parameter model reaches the same loss as a 10-million parameter model while seeing roughly 100× fewer data examples.
This creates a virtuous cycle for scaling: not only do larger models converge to lower loss, they also get there faster in terms of data. The optimal strategy is not "collect as much data as possible and train a big model on all of it," but rather "build the biggest model you can afford and train it on a modest amount of data."
The learning curves themselves follow a power law: where is the number of training steps. This means you can predict the final loss from the first few percent of training — a powerful tool for deciding whether to continue an expensive run.
The critical batch size
There is an optimal batch size for training, and it too follows a power law — but remarkably, it depends only on the loss value, not on the model size. The roughly doubles for every 13% decrease in loss, following with .
Training below the critical batch size wastes time (too many serial steps). Training above it wastes compute (each step gives diminishing returns). The sweet spot, training at the critical batch size, makes the best use of both.
For the largest models approaching convergence, the critical batch size is roughly 1–2 million tokens — which maps to the large-batch training configurations that became standard practice in subsequent work.
Transformers vs. LSTMs: the architecture that scales
The paper includes a crucial comparison between Transformers and LSTMs. Both follow power-law scaling with parameters, but Transformers asymptotically outperform LSTMs — and the gap widens as models grow.
The reason is context utilization. LSTMs plateau after about 100 tokens of context: adding more context doesn't help. Transformers, by contrast, continue to improve throughout their entire 1024-token . Each additional word of context makes the 's predictions meaningfully better.
This finding provided quantitative evidence for what the Transformer paper had argued qualitatively: 's direct word-to-word connections are not just architecturally elegant — they are the mechanism that enables scaling.
The ceiling: where the laws must break down
The authors noticed a deep tension in their own equations. Compute-efficient training grows data very slowly — . But avoiding overfitting requires . These two trends eventually contradict: at some point, compute-efficient training won't have enough data to avoid overfitting, even if you never reuse a single token.
The authors estimated this intersection at roughly parameters and tokens, with a loss of about 1.7 nats/token. They conjectured this might represent the inherent of natural language — an irreducible floor below which no model can go, no matter how large.
This prediction was remarkably prescient. Modern frontier models have indeed approached the trillion-parameter, trillion-token scale, and the quest for data has become one of the field's most pressing challenges.
Generalization transfers predictably
When the authors tested their WebText2-trained models on completely different text distributions — Wikipedia, Books, Common Crawl — they found a consistent pattern: the loss on other distributions follows the same power law with model size, just shifted by a constant offset.
This means that improving on the training distribution translates directly to improvement on other distributions. A model doesn't just memorize WebText2 patterns — it learns general language understanding that transfers. Furthermore, this gap doesn't depend on training duration or model depth; it depends only on the validation loss achieved.
What these equations unlocked
2020
Scaling Laws (this paper)
Established power-law relationships between loss and scale across seven orders of magnitude, transforming AI scaling from guesswork into predictive engineering.
2020
GPT-3 — scaling laws in action
Brown et al. trained a 175B-parameter model, guided by scaling law predictions. The result exhibited emergent few-shot abilities that nobody had programmed.
2022
Chinchilla — refining the optimal tradeoff
Hoffmann et al. showed that Kaplan's allocation was undertrained: data should scale equally with parameters. The framework held; only the exponents shifted.
2022
Emergent abilities observed at scale
Wei et al. documented capabilities that appear suddenly above certain model sizes — arithmetic, multi-step reasoning, instruction following — raising questions about what smooth power laws can and cannot predict.
2023
GPT-4 — predictable scaling validated
OpenAI reported that scaling laws fitted on small models accurately predicted GPT-4's final performance before the full training run completed.
2023
LLaMA — scaling for inference efficiency
Touvron et al. extended Chinchilla-optimal training far beyond 20 tokens per parameter, trading training compute for inference efficiency — a new dimension of scaling.
Kaplan et al. didn't just discover power laws — they created a predictive science of AI scaling. Before this paper, training a frontier model was an act of faith. After it, teams could write down equations, plug in their compute budget, and predict within a few percent what performance they would achieve. The specific numbers have been refined — the Chinchilla paper corrected the data/model balance — but the paradigm of "fit power laws on small runs, extrapolate to large ones" is now standard practice at every AI lab.
CitationKaplan, McCandlish, Henighan, Brown, Chess, Child, Gray, Radford, Wu, Amodei. Scaling Laws for Neural Language Models. arXiv, 2020.
Terms in this paper
- Scaling Lawقانون التحجيم
- Power Lawقانون القدرة
- Compute-Optimalالأمثل حوسبياً
- Irreducible Lossالفقد غير القابل للاختزال
- Tokens Per Parameterرموز لكل معامل
- Sample Efficiencyكفاءة استخدام العيّنات
- Critical Batch Sizeحجم الدفعة الحرج
- Compute Budgetالميزانية الحوسبية