Language Models2022intermediate10 min read

Training Compute-Optimal Large Language Models

تدريب النماذج اللغوية الكبيرة بالكفاءة الحوسبية المُثلى

Hoffmann, J. · Borgeaud, S. · Mensch, A. · Buchatskaya, E. · Cai, T. · Rutherford, E. · de Las Casas, D. · Hendricks, L. A. · Welbl, J. · Clark, A. · Rae, J. W. — NeurIPS

The problem

By 2022, the dominant strategy for improving language models was to scale up parameters: GPT-3 had 175B, Gopher had 280B, and Megatron-Turing NLG reached 530B. Yet all of them were trained on roughly the same 300 billion tokens. Nobody had systematically asked: for a fixed , what is the optimal split between model size and data? The field was flying blind on one of its most expensive decisions.

The contribution

By training over 400 models (70M to 16B parameters, 5B to 500B tokens), the authors discovered that training requires scaling model size and training data equally: doubling the model means doubling the data. They proved this by training Chinchilla — a 70B- model on 1.4 trillion tokens using the same compute as the 280B Gopher. Chinchilla outperformed Gopher, GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on virtually every benchmark, including a 7% improvement on MMLU. The L(N,D) = A/N^α + B/D^β + E formalized the relationship between model size, data, and .

The impact

Chinchilla rewrote the playbook for training large language models. It proved that most existing models were severely , and that a smaller, data-rich model could outperform giants at a fraction of the . LLaMA, Mistral, and subsequent efficient models all followed the Chinchilla recipe: train smaller models on far more data. The paper also highlighted an emerging crisis — the field was running out of high-quality text data — shaping the research agenda for years to come.

Imagine you're cooking a feast with a fixed grocery budget. Previous chefs kept buying more and more expensive ovens (bigger models) but only buying the same small bag of ingredients (training data). They had magnificent ovens sitting half-empty.

Chinchilla asked: what if you bought a good oven and spent the rest on ingredients?

The answer was stunning: a medium oven fully stocked with ingredients (70B parameters, 1.4T tokens) produced a better meal than a luxury oven with a thin pantry (280B parameters, 300B tokens) — for the same total grocery bill.

The problem: bigger models, same data

Between 2020 and 2022, the race to build the biggest language model was in full swing. GPT-3 scaled to 175 billion parameters. DeepMind's Gopher reached 280 billion. Megatron-Turing NLG pushed to 530 billion. Each generation was bigger than the last.

But there was a blind spot: training data stayed roughly constant. Nearly all of these models trained on around 300 billion tokens. The ratio was wildly unbalanced — GPT-3 saw only about 2 , when the optimal ratio (as this paper would reveal) is closer to 20.

This raises a critical question: if you have a fixed compute budget CC, how should you split it between model size NN and training tokens DD? Previous work by Kaplan et al. suggested scaling models aggressively while keeping data modest. Chinchilla challenged that conclusion head-on.

Open in Lab
Compare how different models allocated their compute budget. Notice how most models had very few tokens per parameter compared to Chinchilla's ~20.
The demo wakes as you arrive…

Three approaches, one answer

The authors didn't rely on a single experiment. They attacked the question from three independent angles, training over 400 models in total. Each approach slices the data differently, but all three converge on the same conclusion: model size and data should scale equally.

Approach 1 — Fix the model, vary the data. For each model size (70M to 16B), train on different amounts of data and track where the loss bottoms out at each compute budget. The optimal model size for each budget traces a frontier.

Approach 2 — Fix the compute, vary the split. Choose 9 different compute budgets (IsoFLOP curves). For each budget, train models of different sizes (each getting a different number of tokens to stay within budget). Plot loss against model size — the minimum of each parabola is the compute-optimal point.

Approach 3 — Fit a parametric law. Fit all 400+ runs to a single equation that predicts loss from model size and data jointly. Then solve for the optimal allocation analytically.

Open in Lab
Explore the three approaches. Each tab shows how the authors attacked the problem from a different angle — all arriving at the same conclusion.
The demo wakes as you arrive…

The scaling law: predicting loss from size and data

The parametric (Approach 3) captures the full relationship between model size, training data, and loss in one elegant equation. Before we see the math, here is the intuition:

Think of loss as a sum of three "imperfections." The first comes from having a finite-sized model — it can't represent everything. The second comes from having finite training data — the model hasn't seen enough examples to learn well. The third is the — even a perfect model trained on infinite data can't predict language perfectly, because language itself is noisy and unpredictable.

Making the model bigger reduces the first imperfection. Adding more data reduces the second. But the third never goes away.

L(N,D)=ANα+BDβ+EL(N, D) = \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}} + E
Chinchilla scaling law — loss as a function of model size and data — A/N^α = penalty for finite model size · B/D^β = penalty for finite data · E = irreducible loss · N = number of parameters · D = number of training tokens
Open in Lab
Drag the sliders to change model size (N) and data (D). Watch how loss responds — and notice that improving only one axis hits diminishing returns.
The demo wakes as you arrive…

To find the optimal allocation, we minimize L(N,D)L(N, D) subject to the constraint C≈6NDC ≈ 6ND. Using Lagrange multipliers, the optimal model size and data scale as:

Nopt∝CaN_{opt} \propto C^a and Dopt∝CbD_{opt} \propto C^b

where a=βα+βa = \frac{\beta}{\alpha + \beta} and b=αα+βb = \frac{\alpha}{\alpha + \beta}. Since a≈b≈0.5a ≈ b ≈ 0.5, both NN and DD should grow as the square root of compute — confirming the equal-scaling rule.

The proof: Chinchilla vs. the giants

To validate their scaling law, the authors made a bold bet. They took the exact compute budget used to train Gopher (280B parameters, 300B tokens) and asked: what does our law predict as the optimal model for this budget?

The answer: 70B parameters trained on 1.4 trillion tokens — a model 4× smaller but trained on 4× more data. They named it Chinchilla.

The result was decisive. Chinchilla didn't just match the bigger models — it outperformed every single one of them on a vast range of benchmarks. A 70B model beat a 530B model. On the MMLU benchmark (a test of broad knowledge), Chinchilla scored 67.5% — a 7% jump over Gopher's 60.0%. On BIG-Bench, it improved average accuracy from 54.4% to 65.1%.

Open in Lab
Chinchilla (70B) vs. larger models on major benchmarks. Hover over each bar for details.
The demo wakes as you arrive…

The compute-optimal frontier

For any compute budget, there is exactly one optimal (N, D) pair. The paper plots these pairs across many budgets, forming the . Models above the frontier are too big for their data; models below are too small to exploit their data. Almost every major model of that era sat well above the frontier — oversized and undertrained.

The frontier follows a simple rule: as compute doubles, both optimal NN and optimal DD should grow by roughly 2\sqrt{2} ≈ 1.41×. This means doubling compute should make your model ~40% bigger and give it ~40% more data — not double the model size with the same data.

Open in Lab
Drag the compute budget slider to see the optimal (N, D) point move along the frontier. Existing models are plotted for comparison.
The demo wakes as you arrive…

The scaling law in code

Chinchilla scaling law: predict loss and find the optimal splitpython

Simplified to show the idea — not the real implementation.

import numpy as np
from scipy.optimize import minimize_scalar

# Fitted constants from the paper (Approach 3)
A     = 406.4       # model-size coefficient
B     = 410.7       # data coefficient
alpha = 0.34        # model-size exponent
beta  = 0.28        # data exponent
E     = 1.69        # irreducible loss

def chinchilla_loss(N, D):
    """Predict cross-entropy loss for N parameters and D tokens."""
    return A / N**alpha + B / D**beta + E

def optimal_allocation(C):
    """Given compute budget C (in FLOPs), find optimal N and D.
    Constraint: C ≈ 6 * N * D, so D = C / (6 * N).
    """
    def loss_for_N(log_N):
        N = np.exp(log_N)
        D = C / (6 * N)
        if D < 1:
            return 1e10
        return chinchilla_loss(N, D)

    # Search over model sizes from 10M to 1T
    result = minimize_scalar(loss_for_N, bounds=(np.log(1e7), np.log(1e12)),
                             method='bounded')
    N_opt = np.exp(result.x)
    D_opt = C / (6 * N_opt)
    return int(N_opt), int(D_opt), result.fun

# Example: Gopher's compute budget (≈ 5.76e23 FLOPs)
C_gopher = 6 * 280e9 * 300e9
N_opt, D_opt, loss = optimal_allocation(C_gopher)
print(f"Optimal: {N_opt/1e9:.1f}B params, {D_opt/1e9:.0f}B tokens")
print(f"Predicted loss: {loss:.3f}")
# → Optimal: ~67B params, ~1400B tokens (close to Chinchilla!)

Why it mattered: the aftershocks

Chinchilla's implications rippled through the entire field. Three consequences reshaped how language models are built:

1. Smaller, data-rich models replaced giants. Meta's LLaMA (2023) trained a 65B model on 1.4T tokens — directly following Chinchilla's recipe — and matched GPT-3's performance at a fraction of its size. Mistral, Gemma, and Phi followed the same philosophy.

2. Data became the bottleneck. If optimal training requires ~20 tokens per parameter, a 1-trillion-parameter model needs 20 trillion tokens of high-quality text. The internet has a finite supply. This sparked intense research into , data filtering, and data recycling.

3. Inference cost dropped. A 70B model serves queries 4× faster and 4× cheaper than a 280B model. Chinchilla-optimal training makes deployment practical for far more organizations.

  1. 2020

    Kaplan et al. — first scaling laws

    OpenAI published the first systematic scaling laws for language models, suggesting that model size should be scaled faster than data. This became the conventional wisdom that Chinchilla would later overturn.

  2. 2020

    GPT-3 — 175B parameters, ~300B tokens

    OpenAI's landmark model demonstrated emergent abilities but was severely undertrained by Chinchilla's later analysis — only ~2 tokens per parameter.

  3. 2021

    Gopher — 280B parameters, 300B tokens

    DeepMind's own large model. Strong performance, but Chinchilla would later show that the same compute could achieve better results with a 4× smaller model on 4× more data.

  4. 2022

    Chinchilla — 70B params, 1.4T tokens

    The compute-optimal model that beat all the giants. Same compute as Gopher, 4× smaller, 4× more data — and better on virtually every benchmark.

  5. 2023

    LLaMA — Chinchilla's recipe goes open

    Meta trained LLaMA following Chinchilla-optimal ratios and released it openly. A 65B model on 1.4T tokens matched GPT-3. Democratized access to strong language models.

  6. 2024

    Beyond Chinchilla — inference-aware scaling

    Researchers extended Chinchilla's framework to account for inference costs. When serving billions of queries, training a slightly smaller model on even more data can minimize total lifetime cost — pushing the optimal ratio beyond 20 tokens per parameter.

Chinchilla's legacy is a shift in mindset. Before it, "bigger is better" was the unquestioned mantra. After it, the question became: "bigger at what ratio?" The answer — scale both axes equally — seems obvious in hindsight, but it took training 400 models to prove it, and it changed how every major lab allocates its most expensive resource: compute.

CitationHoffmann, Borgeaud, Mensch, Buchatskaya, Cai, Rutherford, de Las Casas, Hendricks, Welbl, Clark, Rae et al.. Training Compute-Optimal Large Language Models. NeurIPS, 2022.

Terms in this paper