Language Models2020intermediate12 min read

Language Models Are Few-Shot Learners

النماذج اللغوية قادرة على التعلّم بأمثلة قليلة

Brown, T. · Mann, B. · Ryder, N. · Subbiah, M. · Kaplan, J. · Dhariwal, P. · Neelakantan, A. · Shyam, P. · Sastry, G. · Askell, A. · Agarwal, S. · Herbert-Voss, A. · Krueger, G. · Henighan, T. · Child, R. · Ramesh, A. · Ziegler, D. · Wu, J. · Winter, C. · Hesse, C. · Chen, M. · Siber, E. · Litwin, M. · Gray, S. · Chess, B. · Clark, J. · Berner, C. · McCandlish, S. · Radford, A. · Sutskever, I. · Amodei, D. — NeurIPS

The problem

By 2020, the dominant paradigm was pre-train a , then fine-tune it on each downstream task with a separate labeled dataset. This approach had three fundamental problems: (1) every new task required thousands of labeled examples and a separate run, (2) on a narrow dataset could exploit spurious correlations rather than learn the actual task, and (3) humans don't need thousands of examples to learn a new task — a few demonstrations suffice.

The contribution

GPT-3: a 175-billion- autoregressive that performs tasks via — no updates at . You write a natural-language description of the task plus a few input-output examples inside the , and the generates the correct continuation. On many NLP benchmarks, GPT-3's performance rivals or exceeds fine-tuned models, demonstrating that scale unlocks emergent task-learning abilities that smaller models cannot exhibit.

The impact

GPT-3 redefined what a language model could be. It introduced in-context learning as a practical alternative to fine-tuning, sparked the field of , and proved that scaling alone — more parameters, more data, more compute — yields qualitatively new capabilities. It was the direct ancestor of ChatGPT, InstructGPT, Codex, and DALL·E, and its API became the first commercially deployed LLM product. The paper's scaling curves became the empirical foundation for the scaling-laws research program.

Fine-tuning a language model is like hiring a specialist for every question: need translation? hire a translator. Need summarization? hire an editor. Need code? hire a programmer. Each specialist needs training time and labeled data.

GPT-3 is the polyglot generalist sitting at your desk. You slide them a memo that says "Translate English to French" with three example pairs, and they just... do it. No new training. No new weights. The same model, the same parameters — the only thing that changed is the instructions you wrote in the prompt.

The shocking discovery: the bigger the generalist, the fewer examples they need. At 175 billion parameters, two or three examples are often enough.

The problem: fine-tuning is powerful but expensive and fragile

Before GPT-3, the recipe for NLP was established by BERT and GPT-1: pre-train a big model on unlabeled text, then fine-tune on a task-specific labeled dataset. This worked brilliantly — but it introduced three friction points:

  • Data hunger. Every new task needs thousands of labeled examples. Labeling is slow, expensive, and sometimes impossible (how do you label "creative writing quality"?).

  • Spurious shortcuts. Models fine-tuned on narrow datasets learn superficial patterns that break on shift. A sentiment model trained on movie reviews might key on the word "cinematography" rather than actually understanding sentiment.

  • Mismatch with human learning. Humans read a few examples and generalize. They don't need 10,000 labeled pairs to learn what "translate" means. There should be a way to teach a model the same way — by demonstration, not by gradient descent.

Open in Lab
Compare the fine-tuning pipeline (left) with in-context learning (right). Notice: no gradient updates happen on the right side.
The demo wakes as you arrive…

The idea: teach through the prompt, not through gradients

GPT-3's core insight is that a language model trained to predict the next has already learned an implicit learning algorithm. When you write a prompt like:

"Translate English to French: sea otter → loutre de mer cheese → fromage plaid →"

The model doesn't see this as a "translation task." It sees a pattern in the text and predicts the most likely continuation — which happens to be the correct French word. The model is doing pattern completion, but at scale, pattern completion becomes task execution.

The paper defined three evaluation regimes — each providing different amounts of task information inside the prompt:

  • : Only a natural-language task description. No examples. The model must rely entirely on its pre-training knowledge.

  • One-shot: The task description plus a single input-output example. This mirrors how humans might receive a brief demo.

  • Few-shot: The task description plus 10–100 examples, limited only by the model's (2048 tokens). No gradient updates — the examples are simply part of the input text.

Open in Lab
Click through the three prompting regimes to see how zero-shot, one-shot, and few-shot prompts are structured.
The demo wakes as you arrive…

The model: scaling the decoder-only Transformer to 175B parameters

GPT-3 is architecturally the same as GPT-2: a decoder-only trained with the standard autoregressive language modeling objective — predict the next token given all preceding tokens. The innovation is entirely in scale. The paper trained eight models spanning three orders of magnitude to disentangle the effect of size from the effect of architecture:

  • GPT-3 Small — 125M parameters, 12 layers, dmodeld_{model} = 768
  • GPT-3 Medium — 350M parameters, 24 layers, dmodeld_{model} = 1024
  • GPT-3 Large — 760M parameters, 24 layers, dmodeld_{model} = 1536
  • GPT-3 XL — 1.3B parameters, 24 layers, dmodeld_{model} = 2048
  • GPT-3 2.7B — 2.7B parameters, 32 layers, dmodeld_{model} = 2560
  • GPT-3 6.7B — 6.7B parameters, 32 layers, dmodeld_{model} = 4096
  • GPT-3 13B — 13B parameters, 40 layers, dmodeld_{model} = 5140
  • GPT-3 175B — 175B parameters, 96 layers, dmodeld_{model} = 12288, 96 heads

The context window is 2048 tokens across all sizes, and the model uses the same learned positional embeddings and ترميز زوج البايت () with a vocabulary of ~50,257 tokens. The key architectural tweak versus GPT-2 is the use of alternating dense and locally-banded in the Transformer layers, following the Sparse Transformer.

Open in Lab
Explore the eight GPT-3 model sizes. Click each bar to see the architectural details.
The demo wakes as you arrive…

Training data: curating 300 billion tokens from the web

The quality of training data is just as important as model size. GPT-3 was trained on a carefully curated mixture of five datasets, totaling approximately 300 billion tokens. The largest source, Common Crawl, was heavily filtered: a binary classifier trained on high-quality reference corpora (Wikipedia, WebText) was used to score every page, and only high-scoring pages were retained.

The five sources and their approximate weights in training were:

  • Common Crawl (filtered) — 410B tokens available, 60% of training mix
  • WebText2 — 19B tokens, 22% of mix (an expanded version of GPT-2's training set)
  • Books1 — 12B tokens, 8% of mix
  • Books2 — 55B tokens, 8% of mix
  • Wikipedia — 3B tokens, 3% of mix

Critically, higher-quality datasets were upsampled: Wikipedia was seen ~3.4 times during training while Common Crawl was not even fully traversed once. The paper also documented concerns about — benchmark test sets accidentally appearing in the training data — and ran contamination analyses for each evaluation.

Open in Lab
Hover over each dataset to see its size, quality level, and sampling weight during training.
The demo wakes as you arrive…

The scaling discovery: bigger models learn faster with fewer examples

The paper's most striking finding is a smooth power-law relationship between model size and few-shot performance. Across virtually every benchmark, a consistent pattern emerged:

  • Performance scales as a smooth function of the number of parameters — no discontinuities, no lucky breaks. More parameters → better performance, predictably.

  • The gap between zero-shot and few-shot widens with scale. Small models barely benefit from in-context examples; the 175B model sometimes jumps 10–30 percentage points when given a few demonstrations. In other words, in-context learning is an emergent ability that strengthens with scale.

  • On many tasks, few-shot GPT-3 175B matches or exceeds the state-of-the-art set by fine-tuned models, without any gradient updates.

This finding connects directly to the scaling laws research: the same power-law curves that predict training also predict downstream task performance, making model scaling a reliable research strategy rather than a gamble.

Open in Lab
Watch performance scale with model size across different tasks. Toggle between zero-shot, one-shot, and few-shot.
The demo wakes as you arrive…
L(N)≈(NcN)αN\mathcal{L}(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}
Scaling Law for Cross-Entropy Loss — This scaling law shows that training loss decreases predictably as model size increases. The relationship follows a power-law trend, meaning that larger models tend to achieve lower loss in a systematic and measurable way. Because this pattern remains remarkably consistent across a wide range of model sizes, researchers can often estimate future performance before training is complete.

Benchmarks: where GPT-3 shines and where it stumbles

The paper evaluates GPT-3 on over two dozen benchmarks spanning language modeling, question answering, translation, arithmetic, and common sense reasoning. The results paint a nuanced picture:

Strong few-shot results appeared on tasks closely related to language modeling — cloze completion (LAMBADA: 86.4% zero-shot, a new SOTA), reading comprehension, trivia question answering, and translation between high-resource language pairs.

Competitive but mixed results appeared on SuperGLUE (71.8 few-shot vs. 89.0 fine-tuned SOTA), showing that few-shot can approach but not always match targeted fine-tuning.

Weaker results appeared on tasks requiring structured reasoning — such as natural language inference (ANLI), some reading comprehension variants, and comparison tasks — where fine-tuned models still dominated. Arithmetic was particularly revealing: GPT-3 could do 2-digit addition (100%) and 3-digit addition (~80%) but failed on 5-digit addition (~10%), suggesting pattern matching rather than true computation.

Open in Lab
Compare GPT-3's performance across benchmarks. Toggle between zero-shot, one-shot, and few-shot, and compare with fine-tuned SOTA.
The demo wakes as you arrive…

Why does in-context learning work?

The paper observes the phenomenon but does not provide a definitive mechanistic explanation. However, several intuitions help build a mental model:

The format-recognition hypothesis: During pre-training, the model encountered many structured text patterns — Q&A pairs in forums, translation pairs in multilingual documents, instruction-following in how-to guides. At inference time, the few-shot prompt activates a matching pre-trained "subroutine." The model is not learning a new task; it's recognizing which pre-learned capability to invoke.

The Bayesian inference hypothesis: The prompt provides evidence that narrows the posterior distribution over tasks. Zero-shot is a broad prior ("any text completion"); each example sharpens the posterior toward the specific task. Larger models have richer priors, so they need fewer examples to converge.

The practical consequence: The user becomes the "programmer" — not by writing code, but by writing demonstrations. The quality of results depends on prompt design: the ordering of examples, the format of the delimiter, and the phrasing of the task description all matter significantly. This realization gave birth to the field of prompt engineering.

Open in Lab
Visualize how the model's internal "task recognition" sharpens as more examples are added to the prompt.
The demo wakes as you arrive…

Limitations the paper openly acknowledged

The GPT-3 paper is unusually candid about its own weaknesses — an entire section is dedicated to limitations:

  • Text generation quality. GPT-3 can produce fluent text that drifts, repeats, contradicts itself, or includes nonsensical statements — especially in long generations. The model has no internal fact-checker.

  • Structural reasoning. Tasks requiring comparison, multi-step deduction, or tracking entities across long passages remain weak. The model often "pattern matches" rather than reasons.

  • Sample efficiency is still low vs. humans. GPT-3 needs ~300B tokens of pre-training data. Humans learn language from orders of magnitude less data. The model is data-efficient at inference (few-shot) but data-hungry at training.

  • Ambiguity of in-context learning. It remains unclear whether the model is truly learning new tasks from examples or merely recognizing tasks it already learned during pre-training and applying them. The paper explicitly flags this as an open question.

  • Cost and accessibility. Training a 175B-parameter model required an estimated $4.6M in compute (in 2020 cloud prices) and a cluster of high-end GPUs for weeks. This made the research inherently inaccessible to most of the academic community.

Broader impacts: the social implications GPT-3 surfaced

The paper includes an extended discussion of societal concerns — unusually thorough for its time:

  • Misuse potential. Fluent text generation at scale could enable spam, phishing, and disinformation campaigns. The paper noted that GPT-3's outputs were difficult for humans to distinguish from human-written text, even with only 500 words of context.

  • . The model inherits biases from its training data. The paper conducted targeted evaluations on gender, race, and religious bias, finding that GPT-3 associates certain occupations with specific genders and reflects stereotypes present in its web training data.

  • Energy and environment. Training GPT-3 required substantial energy. The paper argued that one-time pre-training costs are amortized over many inference-time uses, but the environmental footprint of scaling remained a concern.

What GPT-3 unlocked

  1. 2020

    GPT-3 — 175B parameters

    In-context learning proved viable at scale. The API became the first commercial LLM product. Prompt engineering was born.

  2. 2021

    Codex — Code generation from natural language

    Fine-tuned GPT-3 on code. Powered GitHub Copilot. Showed that the same architecture could master programming, not just prose.

  3. 2021

    DALL·E — Images from text prompts

    Applied GPT-3's autoregressive framework to image token generation. Text-to-image became real.

  4. 2022

    InstructGPT — Alignment through human feedback

    Fine-tuned GPT-3 with RLHF to follow instructions more faithfully. The direct precursor to ChatGPT.

  5. 2022

    Chain-of-Thought Prompting

    Showed that adding "let's think step by step" to the prompt unlocks multi-step reasoning. A direct exploitation of GPT-3's in-context learning.

  6. 2022

    ChatGPT — LLMs for everyone

    Built on GPT-3.5 (a GPT-3 descendant). RLHF-tuned for dialogue. 100 million users in two months.

  7. 2023

    GPT-4 — Multimodal and more capable

    The next generation. Accepts images and text. Passes professional exams. The scaling curve continues.

GPT-3's deepest legacy is not any single benchmark result — it's the paradigm shift. Before GPT-3, "using a language model" meant fine-tuning. After GPT-3, it meant prompting. That single conceptual change — from gradient-based specialization to text-based instruction — is the foundation on which every modern LLM application is built.

CitationBrown, Mann, Ryder, Subbiah, Kaplan, Dhariwal, Neelakantan, Shyam, Sastry, Askell, Agarwal, Herbert-Voss, Krueger, Henighan, Child, Ramesh, Ziegler, Wu, Winter, Hesse, Chen, Siber, Litwin, Gray, Chess, Clark, Berner, McCandlish, Radford, Sutskever, Amodei. Language Models Are Few-Shot Learners. NeurIPS, 2020.

Terms in this paper