Language Models2019intermediate11 min read

Language Models Are Unsupervised Multitask Learners

النماذج اللغوية: تعلُّم مهام متعددة بلا إشراف

Radford, A. · Wu, J. · Child, R. · Luan, D. · Amodei, D. · Sutskever, I. — OpenAI Technical Report

The problem

By 2019 NLP systems were narrow specialists: each task — translation, QA, — required its own labeled dataset and its own run. GPT-1 had shown that pre- helps, but you still needed task-specific supervision to get good performance. Nobody had demonstrated that a single , with no fine-tuning at all, could perform multiple tasks competitively just by reading a natural-language .

The contribution

GPT-2: a 1.5-billion parameter pre-trained on WebText (40 GB of quality-filtered web pages). Three linked insights: (1) a new, curated web dataset far larger and more diverse than BookCorpus; (2) at the byte level, giving a single that handles any language and symbol; (3) the demonstration that scale alone — more data, more parameters, same architecture — unlocks multitask ability. GPT-2 achieved state of the art on 7 of 8 language-modeling benchmarks with no fine-tuning, performed at 55 F1 (CoQA), and even demonstrated rudimentary translation (11.5 BLEU on WMT-14 Fr→En) despite seeing almost no French data.

The impact

GPT-2 was the first convincing proof that scale turns a into a general-purpose engine. It reframed the field: instead of designing task-specific architectures, just train a bigger model on more data and prompt it. This insight directly seeded GPT-3's in-context learning, the scaling-laws research program, and ultimately the prompt-based paradigm behind every modern LLM — Claude, Gemini, LLaMA, and beyond.

GPT-1 was a specialist doctor: it studied general medicine (pre-training), then passed a board exam in one specialty at a time (fine-tuning). You needed a separate exam for each skill.

GPT-2 is a polyglot librarian who has read every book in the building: you walk up and say — in plain language — "summarize this chapter," "translate this passage," or "answer this question," and the librarian complies, not because it was trained on any exam, but because it has simply read so much that it already knows how those tasks work.

The problem: narrow specialists in a world of diverse tasks

Before GPT-2, the recipe for every NLP task was the same: collect labeled data, build a task-specific head, and fine-tune. Want translation? Train a translation model. Want summarization? Train a summarization model. Each model was a single-use tool: brilliant at its task, useless at everything else.

This approach has three deep problems:

  • Data hunger. Labeled datasets are expensive to create. Every new task needs thousands of curated examples — a bottleneck for real-world deployment.
  • Brittleness. Models trained on narrow distributions break when the world shifts. A sentiment model trained on movie reviews fails on product reviews.
  • No emergence. A fine-tuned specialist never surprises you — it can only do what you explicitly trained it to do. There is no path from "good at QA" to "suddenly also translates," because the training objective forbids it.

GPT-1 took a first step: pre-train a language model, then fine-tune with a small labeled set. But the fine-tuning step was still mandatory. The question GPT-2 asked was radical: what if we skip fine-tuning entirely?

The insight: task = prompt + prediction

The key realization behind GPT-2 is deceptively simple: every NLP task is just predicting what comes next in a text. Translation? The text "translate English to French: cheese →" naturally calls for "fromage." Summarization? "TL;DR:" at the end of a long passage calls for a short summary. ? "Q: Who wrote Hamlet? A:" calls for "Shakespeare."

If a language model has seen enough examples of these patterns in its training data — naturally occurring question-answer pairs, translations, summaries — it can learn to perform the task just by recognizing the pattern in a new prompt. No labels. No task-specific head. No fine-tuning. Just next- prediction at scale.

The authors formalized this as: instead of modeling p(output∣input)p(\text{output} | \text{input}) for each task, model p(output∣input,task)p(\text{output} | \text{input}, \text{task}) — and encode the task description as text. The language itself becomes the task specification.

Open in Lab
Explore how GPT-2 frames translation, QA, and summarization as next-word prediction — no task-specific heads needed.
The demo wakes as you arrive…

WebText: a new kind of training data

The quality of a language model is bounded by the quality of its data. GPT-1 trained on BookCorpus — ~7,000 unpublished books, roughly 1 GB. Diverse in style but narrow in scope and small by web standards.

GPT-2 introduced WebText: 40 GB of text scraped from outbound links on Reddit that received at least 3 upvotes — a simple but effective human-quality filter. The reasoning: pages that humans find interesting enough to share and upvote are, on average, higher quality than a random web crawl.

The resulting dataset contained 8 million documents. Crucially, it was diverse: news articles, blog posts, fiction, code, forum discussions, how-to guides, academic writing. This diversity is what gives the model exposure to the natural demonstrations of translation, QA, and summarization that make zero-shot transfer possible.

Open in Lab
See how WebText filters Reddit's links into high-quality training data, compared with BookCorpus.
The demo wakes as you arrive…

Byte-level BPE: one tokenizer for all languages

GPT-1 used a word-level tokenizer with a fixed 40,000-word . Any word not in the vocabulary became an <UNK> token — a black hole of lost meaning. Non-English text, rare words, code, and emoji were all at risk.

GPT-2 solved this with byte-level Byte Pair Encoding (BPE). Here is the intuition: start with a vocabulary of 256 raw bytes — the building blocks of any text in any language. Then, iteratively merge the most frequent adjacent byte pairs into new tokens. After enough merges, common words become single tokens ("the" → one token), while rare words are assembled from smaller pieces ("pneumonoultramicroscopic" → several sub-tokens).

The result is a vocabulary of 50,257 tokens that can encode any UTF-8 string — English, Arabic, Chinese, emoji, code — with zero <UNK> tokens. This was critical: the model would never encounter a character it couldn't represent.

Open in Lab
Watch byte-level BPE build tokens from raw bytes — type any text and see how GPT-2's tokenizer handles it.
The demo wakes as you arrive…

Architecture: the same recipe, scaled up

GPT-2's architecture is intentionally minimal in novelty — the insights are in the data and scale, not in structural changes. It uses the same Transformer decoder as GPT-1, with four refinements:

  • Pre-activation . Layer normalization is moved before each sub-block ( and feed-forward) instead of after. This stabilizes training at scale by ensuring each block receives normalized inputs, like a chef who washes ingredients before cooking rather than after.
  • Expanded . From 512 to 1,024 tokens — the model can now see twice as far into the past, crucial for tasks like summarization that need long-range context.
  • Modified initialization. Residual layer weights are scaled by 1/N1/\sqrt{N} at initialization, where NN is the number of residual layers. This prevents the accumulation of as signals flow through many layers — imagine turning down the volume slightly at each step so the signal doesn't clip at the end.
  • Larger vocabulary. 50,257 tokens (from byte-level BPE) instead of 40,000.
Open in Lab
Compare GPT-1 and GPT-2 side by side — same bones, different scale.
The demo wakes as you arrive…

The paper trained four model sizes to study how performance scales:

Open in Lab
GPT-2 models ranged from 117M to 1.5B parameters — explore how perplexity drops with scale.
The demo wakes as you arrive…

Zero-shot results: scale unlocks ability

The most striking finding in the paper is what happens when you simply increase model size — without changing the training objective or adding any labeled data. Every task improves as the model gets bigger, and some capabilities that barely existed in the smallest model become competitive in the largest.

Language modeling: GPT-2 achieved state-of-the-art on 7 of 8 benchmarks — including Penn Treebank (35.76), WikiText-2 (18.34), and LAMBADA (8.63) — all in zero-shot. Remarkably, it still underfits WebText, meaning larger models would likely improve further.

Reading comprehension: On CoQA (conversational QA), GPT-2 scored 55 F1 without seeing a single example — matching or exceeding several supervised baselines that were specifically designed for the task.

Translation: Despite deliberately removing non-English pages during data collection, GPT-2 achieved 11.5 BLEU on WMT-14 Fr→En, outperforming several unsupervised baselines. The model had found the small amount of French text naturally mixed into English web pages and learned rudimentary translation from it.

Summarization: On CNN/Daily Mail, GPT-2 generated summaries comparable to early abstractive models when prompted with "TL;DR:", though this was its weakest task.

The clear pattern: bigger models perform better on every task, and capabilities that don't exist at small scale emerge at large scale.

Open in Lab
Explore GPT-2's zero-shot results across tasks — notice how every task improves with model size.
The demo wakes as you arrive…

Why does this work? The generalist hypothesis

The paper advances a bold hypothesis: a language model trained on enough diverse data implicitly learns to perform many tasks, because performing those tasks is necessary for optimal next-token prediction.

Think about it: to predict the next word in a passage that contains a question and then its answer, the model must learn what question-answering is. To predict the next word in a passage that summarizes a longer document, the model must learn what summarization is. The web is so diverse that practically every NLP task appears somewhere in natural text — and the language model absorbs all of them.

This is the generalist hypothesis: sufficiently large language models are implicit multitask learners, and the performance on each task is a function of model capacity and data diversity. The implication is transformative: you don't need to define tasks at all. You just need a big model and a big, diverse dataset. The tasks emerge.

Memorization or generalization?

A natural worry: is GPT-2 just memorizing its training data and regurgitating it? The authors investigated this with overlap analysis. They computed the fraction of 8-grams in each that also appeared in WebText.

The overlap ranged from 1% to 6% depending on the benchmark — nonzero, but far too small to explain the performance gains. On LAMBADA, a dataset specifically designed to require long-range understanding, GPT-2 achieved dramatic improvements over its smaller variants, suggesting genuine comprehension rather than simple lookup. The model performs well on unseen patterns, not just seen ones.

What GPT-2 changed

GPT-2 shifted the field's center of gravity from task-specific design to scaling and prompting. Its lasting contributions:

  • The prompting paradigm. Demonstrated that natural language can specify tasks, replacing custom architectures with text. This directly leads to the practices used with every modern LLM.
  • Scaling as a research variable. Showed that simply making a model bigger, on more data, is itself a research contribution. This directly inspired the scaling laws research that followed.
  • The data recipe matters. WebText proved that data quality and diversity are as important as architecture. This idea matured into the careful data curation behind GPT-3, Chinchilla, and LLaMA.
  • Byte-level BPE as standard. GPT-2's tokenizer became the default for the field — adopted (with variations) by GPT-3, RoBERTa, and many others.
  1. 2018

    GPT-1

    Pre-train a Transformer decoder, then fine-tune on each task. Proved that unsupervised pre-training transfers, but fine-tuning was still required.

  2. 2019

    GPT-2 — Zero-shot multitasking

    Scale the model 10×, train on WebText, and skip fine-tuning. Showed that zero-shot performance improves log-linearly with model size.

  3. 2020

    GPT-3 — In-context learning

    Scale 100× again (175B parameters). Demonstrated few-shot learning from examples in the prompt — no gradient updates at all. Confirmed GPT-2's scaling hypothesis.

  4. 2020

    Scaling Laws (Kaplan et al.)

    Formalized the relationship GPT-2 hinted at: loss decreases as a power law with model size, data size, and compute. Made scaling predictable.

  5. 2020

    T5 — Text-to-Text Transfer

    Unified every NLP task as text-to-text, validating GPT-2's intuition that tasks are just text patterns — but with encoder-decoder instead of decoder-only.

  6. 2021

    Decision Transformer

    Applied GPT-2's autoregressive framework to reinforcement learning — modeling trajectories as sequences of states, actions, and rewards.

GPT-2's deepest legacy is not a benchmark score — it is the shift in mindset. Before GPT-2, researchers asked "which architecture solves this task?" After GPT-2, the question became "how much data and compute does this task need?" That shift is the foundation of the entire LLM era.

CitationRadford, Wu, Child, Luan, Amodei, Sutskever. Language Models Are Unsupervised Multitask Learners. OpenAI Technical Report, 2019.

Terms in this paper