Language Models2022advanced13 min read

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

توسيع نطاق النماذج اللغوية: أساليب وتحليلات ورؤى من تدريب Gopher

Rae, J. W. · Borgeaud, S. · Cai, T. · Millican, K. · Hoffmann, J. · Song, F. · Aslanides, J. · Henderson, S. · Ring, R. · Young, S. · Rutherford, E. · Hennigan, T. · Menick, J. · Cassirer, A. · Powell, R. · van den Driessche, G. · Hendricks, L. A. · Rauh, M. · Huang, P.-S. · Glaese, A. · Welbl, J. · Dathathri, S. · Huang, S. · Uesato, J. · Mellor, J. · Higgins, I. · Creswell, A. · McAleese, N. · Wu, A. · Elsen, E. · Jayakumar, S. M. · Buchatskaya, E. · Budden, D. · Sutherland, E. · Simonyan, K. · Paganini, M. · Sifre, L. · Martens, L. · Li, X. L. · Kuncoro, A. · Nematzadeh, A. · Gribovskaya, E. · Donato, D. · Lazaridou, A. · Mensch, A. · Lespiau, J.-B. · Tsimpoukelli, M. · Grigorev, N. · Fritz, D. · Sottiaux, T. · Pajber, M. · Ossowski, T. · Piot, B. · Vinyals, O. · Ayoub, K. · Stanway, J. · Bennett, L. · Hassabis, D. · Kavukcuoglu, K. · Irving, G. — arXiv

The problem

By late 2021, scaling language models had become a dominant research strategy, but the relationship between scale and downstream capability was poorly understood. GPT-3 showed that 175 billion parameters could unlock few-shot abilities, yet no systematic study had isolated scale effects across a truly broad evaluation suite, nor examined how scaling interacts with , , and . Researchers needed a rigorous map of what scaling buys and what it does not.

The contribution

The Gopher family: six autoregressive models from 44M to 280B parameters, all trained on 300B tokens from MassiveText (10.5 TB of filtered English text). Evaluated on 152 diverse tasks, Gopher achieved state-of-the-art on 100 of 124 comparable benchmarks. The paper provides a per-domain scaling map showing that , fact-checking, and general knowledge benefit most from scale, while mathematical and logical reasoning gain far less. It includes a comprehensive toxicity and bias analysis, dialogue experiments, and the first detailed study of data quality's impact on downstream performance.

The impact

Gopher was the catalyst for Chinchilla, which used Gopher's scaling insights to show that the 280B model was significantly undertrained — the same yields better results with a smaller model trained on more data. This insight reshaped the entire field's approach to compute allocation. Gopher's MassiveText pipeline and comprehensive evaluation methodology influenced subsequent models from DeepMind and beyond.

Imagine you are building telescopes of increasing power — from a handheld spyglass to a giant observatory mirror. Each bigger telescope reveals more distant stars and galaxies. But there is a catch: some things, like reading a street sign across the road, do not need a bigger lens — a pair of binoculars would do. And some things, like seeing through fog, are not helped by a bigger mirror at all.

This paper builds six "telescopes" (language models from 44 million to 280 billion parameters) and systematically tests what each size can see across 152 different tasks. The biggest telescope — Gopher — sees further than any before it on most tasks. But the paper's real value is the map it draws: where bigger helps, where it barely matters, and where something fundamentally different is needed.

The Gopher family: architecture and training

Gopher is not a single model but a family of six autoregressive Transformer language models spanning four orders of magnitude: 44M, 117M, 417M, 1.4B, 7.1B, and 280B parameters. The core architecture follows GPT-style design with two key modifications.

First, the team replaced LayerNorm with — a simpler normalization that skips the mean-centering step and normalizes only by the root mean square. This improves at scale without sacrificing quality. Think of it as replacing a full car wash with a quick rinse that cleans just as well — fewer moving parts, less to go wrong.

Second, they replaced absolute positional encodings with relative positional encodings from Transformer-XL. Instead of stamping each token with a fixed position number, the model learns the distance between tokens. This means the model can be evaluated on sequences longer than it trained on — like learning to read paragraphs and then being able to read whole chapters.

All six models were trained on the same 300 billion tokens from MassiveText using a 2048-token and with a 32,000-token vocabulary. The only changes across sizes were the number of layers, model dimension, number of heads, , and — everything else was held constant to isolate the effect of scale.

Open in Lab
Explore the six models in the Gopher family. Click on any model to see its architecture details: layers, heads, model dimension, learning rate, and batch size.
The demo wakes as you arrive…

MassiveText: the data pipeline

The quality of data matters as much as the size of the model. The Gopher team built MassiveText, a 10.5 TB dataset of filtered English text from six sources: MassiveWeb (curated web pages), Books, News articles, C4 (Common Crawl), GitHub code, and Wikipedia. In total, MassiveText contains 2.35 billion documents and approximately 2.3 trillion tokens.

Crucially, the models trained on only 300 billion of those tokens — about 12.8% of the full dataset. This means Gopher never saw most of its training data even once. The sampling was not uniform: web pages were given 48% of the training diet, books 27%, news 10%, GitHub code 3%, C4 10%, and Wikipedia 2%. This deliberate weighting reflected the team's belief that high-quality sources like books and curated web pages contribute disproportionately to model capability.

The data pipeline included five successive filtering stages: content filtering, text quality filtering using a classifier, removal of repetitious text, deduplication of near-duplicate documents, and removal of documents with significant test-set overlap. Each stage measurably improved downstream task performance, making a strong case that data curation is not optional — it is a first-class engineering concern.

Open in Lab
Walk through the five stages of the MassiveText data pipeline. Each stage shows how data quality improves and how much data is filtered.
The demo wakes as you arrive…

Training at scale: parallelism and stability

Training a 280B model is an engineering challenge as much as a research one. The Gopher team used TPUv3 pods with a combination of and . A key finding was that these two strategies are low-overhead on TPUs — only about 10% slowdown — thanks to fast cross-chip communication. This eliminated the need for , which greatly simplified the training setup.

The team used the optimizer and trained with bfloat16 mixed precision. Models smaller than 7.1B used float32 parameters with bfloat16 activations, while 7.1B and 280B used bfloat16 for both parameters and activations with stochastic rounding for stability. The authors later noted that stochastic rounding did not fully recover performance — a candid admission that advanced training techniques at this scale involve unresolved tradeoffs.

Gopher's batch size increased during training from 3 million to 6 million tokens, using a learning rate of 4×10−54 \times 10^{-5} with cosine decay. The learning rate decreased and batch size increased with model size — a pattern now standard practice.

The scaling map: where scale helps and where it does not

This is the paper's central contribution. Gopher was evaluated on 152 tasks spanning language modeling, reading comprehension, fact-checking, question answering, common sense reasoning, mathematical reasoning, ethical identification, and more. Compared against GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B), Gopher outperformed the state of the art on 100 out of 124 comparable tasks.

But the gains were uneven. The paper reveals a clear pattern:

Scale helps most for knowledge-intensive tasks — reading comprehension (RACE), fact-checking (FEVER), general knowledge (MMLU), and identification of toxic language. These are tasks where a larger model can simply store and recall more facts. On the MMLU , Gopher achieved 60.0% average accuracy across 57 subjects, a dramatic improvement over prior models.

Scale helps least for mathematical reasoning (GSM, MATH) and logical reasoning. Going from 7.1B to 280B barely moved the needle on these tasks. This suggests that reasoning is not simply a function of memorizing more patterns — it requires something architecturally or algorithmically different.

This asymmetry became a foundational insight for the field. It told researchers: if you want better math, do not just make the model bigger.

Open in Lab
The scaling map: compare how different task categories respond to increasing model size. Drag the model size slider to see which domains benefit most and least from scale.
The demo wakes as you arrive…

MMLU: measuring knowledge across 57 subjects

The Massive Multitask Language Understanding (MMLU) benchmark tests a model across 57 academic subjects — from abstract algebra to world religions. It is essentially a comprehensive university exam. Gopher's 60.0% average accuracy represented a massive improvement over prior models, advancing significantly toward human expert performance.

But the pattern within MMLU is telling. Gopher excelled at humanities and social sciences — subjects that rely on encyclopedic knowledge. It struggled with STEM subjects that require symbolic computation or multi-step reasoning. For example, Gopher achieved strong results in Medical Genetics and Marketing but lagged on Abstract Algebra and Formal Logic.

The MMLU breakdown became one of the most cited results in the paper. It provided a clear diagnostic: language models are becoming encyclopedias, not calculators. They are learning what to know, not how to think.

Open in Lab
Explore Gopher's MMLU performance by subject category. Hover or click categories to see how knowledge-heavy vs reasoning-heavy subjects compare.
The demo wakes as you arrive…

Toxicity and bias: what scale does and does not fix

The paper includes one of the most thorough toxicity and bias analyses of its time. The findings are nuanced and sometimes counterintuitive.

On toxicity: when given non-toxic prompts, larger models generate less toxic text — they are better at staying on topic. But when given toxic prompts (from the RealToxicityPrompts dataset), larger models produce more toxic continuations. The bigger model is better at both good behavior and bad — it has learned to continue whatever style it is given more faithfully.

On bias: the paper measured gender and occupation bias using coreference templates. The relationship between scale and bias showed no consistent trend. Some bias metrics increased with scale, others decreased, and others oscillated. This finding is critical: it means bias cannot be fixed by simply scaling up. Dedicated debiasing techniques are necessary.

The paper also noted that Gopher's training data is over 99% English, which limits its usefulness as a multilingual model and introduces cultural biases inherent to English-language internet text.

Open in Lab
Explore the toxicity paradox: larger models are less toxic with clean prompts but more toxic with toxic prompts. Toggle between prompt types to see the divergence.
The demo wakes as you arrive…

Dialogue: from language model to conversational agent

The paper also explored using Gopher for dialogue — both through prompting and through . By conditioning the model on a dialogue (a description of a helpful, polite assistant followed by example conversations), Gopher could conduct multi-turn conversations.

Fine-tuning on curated dialogue data further improved conversational quality. The fine-tuned model showed better factual accuracy and was preferred by human evaluators in pairwise comparisons. Importantly, dialogue prompting also reduced toxicity — the model adopted the tone of the polite assistant it was prompted to emulate.

These dialogue experiments foreshadowed the chatbot era that would arrive with ChatGPT a year later. The insight that prompting alone can turn a into a passable conversational agent was a key stepping stone toward instruction-tuned and RLHF-aligned models.

The autoregressive objective: predict the next token

Before diving into the training objective, let us build an intuition. Think of a language model as someone completing a sentence. Given "The capital of France is", the model should assign high probability to "Paris" and low probability to "banana." The training objective is to get better at this prediction for every position in every document in the training set.

Formally, given a sequence of tokens x1,x2,…,xTx_1, x_2, \ldots, x_T, the model is trained to maximize the :

L(θ)=∑t=1Tlog⁡Pθ(xt∣x1,…,xt−1)\mathcal{L}(\theta) = \sum_{t=1}^{T} \log P_\theta(x_t \mid x_1, \ldots, x_{t-1})
Autoregressive training objective — next-token prediction — At each position tt, the model predicts the next token xtx_t given all previous tokens. The training signal is the cross-entropy loss between the model's predicted distribution and the actual next token. This simple objective, applied at massive scale, produces models that can answer questions, summarize documents, and write code — all as emergent behaviors of next-token prediction.

Key findings: a summary

The Gopher paper's lasting contribution is not just the model itself — Chinchilla later showed it was significantly undertrained — but the systematic framework for understanding scaling. The paper's findings can be distilled into five core insights:

1. Scale is not uniform. Different tasks respond differently to scale. Knowledge-intensive tasks benefit enormously; reasoning tasks barely move. This means compute allocation should be task-aware.

2. Data quality is a force multiplier. Each stage of the MassiveText pipeline independently improved downstream performance. Better data makes bigger models worth their cost.

3. Toxicity scales in both directions. Larger models amplify whatever they are given — clean input gets cleaner output, toxic input gets more toxic output.

4. Bias does not trend with scale. No consistent pattern emerged between model size and bias metrics, meaning bias must be addressed through targeted interventions, not through scaling.

5. Dialogue emerges from prompting. A pre-trained language model can be turned into a conversational agent through careful prompting, without any fine-tuning — a finding that foreshadowed the chatbot revolution.

Open in Lab
The five core insights from the Gopher paper. Click each card to see the evidence and implications.
The demo wakes as you arrive…

The road to Chinchilla: Gopher was undertrained

Perhaps the most important legacy of the Gopher paper is what it made possible next. Jordan Hoffmann and colleagues at DeepMind, many of whom also worked on Gopher, used its scaling data to ask a radical question: was the 280B-parameter model the best use of that compute budget?

The answer was no. Chinchilla showed that a 70B model trained on 1.4 trillion tokens — four times smaller but trained on four times more data — matched or exceeded Gopher on virtually every benchmark. The Gopher team had focused on scaling parameters while holding training tokens constant at 300B. Chinchilla revealed that the optimal split allocates compute more evenly between model size and data size.

This was a paradigm shift. Instead of the "bigger is better" mantra, the field adopted "" as the guiding principle. Every major model released after Chinchilla — LLaMA, Falcon, Mistral — was designed with this balance in mind.

In hindsight, Gopher was both a landmark and a lesson. The landmark was the evaluation methodology and the scaling map. The lesson was that scale alone is not enough — how you allocate your resources between model size and training data is what determines the outcome.

Timeline: from GPT-3 to compute-optimal scaling

  1. 2020

    GPT-3 (175B)

    OpenAI demonstrated that scaling to 175 billion parameters unlocks few-shot learning abilities. Became the reference point for all subsequent scaling work.

  2. 2021

    Jurassic-1 (178B) & MT-NLG (530B)

    AI21 Labs and the Nvidia-Microsoft partnership pushed parameter counts further. MT-NLG became the largest dense model at the time.

  3. 2022

    Gopher (280B) — this paper

    DeepMind introduced the Gopher family and provided the most comprehensive scaling analysis to date: 152 tasks, toxicity analysis, bias study, and dialogue experiments. Showed scale helps knowledge more than reasoning.

  4. 2022

    Chinchilla (70B)

    Using Gopher's scaling data, Hoffmann et al. proved that a 70B model trained on 4× more data matches 280B Gopher — establishing compute-optimal scaling laws.

  5. 2023

    The compute-optimal era begins

    LLaMA, Falcon, and subsequent models all adopted Chinchilla-style compute allocation. Gopher's evaluation methodology became the standard for how to measure what scaling buys.

Gopher occupies a pivotal position in the history of language models. It was the most thorough scaling study of its time: not just "make it bigger and report accuracy," but a careful dissection of where scale helps, where it stalls, and what dangers it amplifies. The evaluation methodology — 152 tasks across dozens of domains — set a new standard that subsequent papers adopted. And the scaling data it produced directly enabled the Chinchilla insight that reshaped the field.

CitationRae, Borgeaud, Cai, Millican, Hoffmann, Song, et al.. Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv, 2022.

Terms in this paper