Language Models2018intermediate9 min read

Improving Language Understanding by Generative Pre-Training

تحسين فهم اللغة عبر التدريب المسبق التوليدي

Radford, A. · Narasimhan, K. · Salimans, T. · Sutskever, I. — OpenAI Technical Report

The problem

By 2018 most NLP systems were trained from scratch on labeled data for each task. Labeled data is expensive and scarce; each task — sentiment analysis, question answering, textual entailment — needed its own architecture and its own dataset. Unlabeled text was abundant but nobody had shown a general recipe for converting raw reading into broad language ability that transfers across tasks.

The contribution

A two-stage semi-supervised approach: (1) of a 12-layer on (~7,000 books) using next- prediction, then (2) on each downstream task with minimal architecture changes — just a linear output layer. Task-specific input transformations convert classification, entailment, similarity, and QA into a unified sequence format. An auxiliary language modeling objective during improves . The result: state-of-the-art on 9 of 12 benchmarks studied.

The impact

GPT-1 proved that unsupervised generative pre-training transfers to diverse supervised tasks — the "pre-train then fine-tune" paradigm that now powers all of modern NLP. It established the Transformer as a foundation model chassis, spawning GPT-2, GPT-3, GPT-4, and inspiring BERT (which flipped it to an ). Every today — Claude, Gemini, LLaMA — descends from this recipe.

Imagine a student who reads thousands of novels before ever taking a class. She never sees an exam question, yet she develops an intuition for how language works — grammar, reasoning, story structure. When she finally sits for a test, she only needs a few minutes of review to ace it.

GPT-1 is that student. Stage one: read books and predict the next word, millions of times. Stage two: glance at a handful of labeled examples and learn to classify, answer questions, or judge textual relationships. The reading did the heavy lifting; the exam prep is almost trivial.

The problem: labeled data is scarce, tasks are many

Before GPT-1, building an NLP system meant: collect a labeled dataset for your specific task, design a task-specific architecture, and train from scratch. Sentiment analysis? One model. Question answering? A different model. Textual entailment? Yet another.

This approach had three crippling weaknesses:

  • Labeled data hunger. Annotating thousands of examples for every new task is slow and expensive.
  • No knowledge transfer. Each model starts from zero — unable to reuse linguistic knowledge learned elsewhere.
  • Architectural fragmentation. Every task demanded bespoke engineering, from mechanisms to output heads.

Meanwhile, the internet was overflowing with unlabeled text — books, articles, web pages — containing implicit knowledge about grammar, world facts, and reasoning patterns. The question was: how do you distill that raw text into a model that transfers to any task?

The idea: pre-train generatively, fine-tune discriminatively

GPT-1's recipe has two stages, and the elegance is in their simplicity:

Stage 1 — Unsupervised pre-training. Take a large corpus of unlabeled text (BooksCorpus: ~7,000 unpublished books). Train a Transformer decoder to predict the next word given all previous words. This is a standard language model objective — but applied to a Transformer decoder rather than an , it captures long-range dependencies far better.

Stage 2 — Supervised fine-tuning. Take the pre-trained model, add a single linear output layer, and train on labeled data for a specific task. The pre-trained weights provide a massive head start: the model already "understands" language, so it only needs to learn the task mapping.

The authors also add the language modeling as an auxiliary objective during fine-tuning. This acts as a regularizer — preventing the model from forgetting what it learned during pre-training, and improving generalization especially on larger datasets.

Open in Lab
Walk through GPT-1's two-stage recipe — click each stage to see what happens inside.
The demo wakes as you arrive…

Now let's look at the formal objectives. Stage 1 maximizes the of each token given the previous k tokens:

L1(U)=∑ilog⁡P(ui∣ui−k,…,ui−1;Θ)L_1(\mathcal{U}) = \sum_i \log P(u_i \mid u_{i-k}, \ldots, u_{i-1}; \Theta)
Stage 1: Language Modeling Objective — Maximize the likelihood of each token given the k previous tokens. The model learns to predict the next word — and in doing so, learns grammar, facts, and reasoning patterns from raw text.

During fine-tuning, the total loss combines the task-specific loss with the language modeling loss:

L3(C)=L2(C)+λ⋅L1(C)L_3(\mathcal{C}) = L_2(\mathcal{C}) + \lambda \cdot L_1(\mathcal{C})
Combined Fine-Tuning Loss — The fine-tuning loss combines the task-specific classification loss L₂ with the original language modeling loss L₁, weighted by λ. This auxiliary objective prevents catastrophic forgetting and improves generalization.

The architecture: a decoder-only Transformer

GPT-1 takes the Transformer decoder from the original Attention Is All You Need paper and removes the layer (since there is no separate encoder). What remains is a stack of masked layers — each token can only attend to tokens before it, never to future tokens. This ensures the model genuinely predicts the next token rather than copying it.

The specific configuration:

  • 12 layers of Transformer decoder blocks
  • 768-dimensional hidden states (the dimension)
  • 12 attention heads per layer (each head operates on 64 dimensions)
  • 3072-dimensional feed-forward intermediate layer (a 4× expansion)
  • activation instead of — smoother gradients
  • Learned positional embeddings (not sinusoidal)
  • 512-token context window
  • ~117 million total parameters
  • tokenizer with ~40,000 merges

This is exactly the same size as BERT-Base — 12 layers, 768 dimensions, 12 heads. The only architectural difference: GPT-1 uses a causal mask (left-to-right), while BERT removes it (bidirectional). This makes their comparison a clean test of direction rather than scale.

Open in Lab
Explore GPT-1's decoder-only architecture — hover over each component to learn its role.
The demo wakes as you arrive…
GPT-1 forward pass (simplified)python

Simplified to show the idea — not the real implementation.

import numpy as np

def gpt1_forward(tokens, W_emb, W_pos, blocks, W_out):
    """
    tokens: list of token IDs, length ≤ 512
    W_emb:  token embedding matrix  (vocab × 768)
    W_pos:  position embedding matrix (512 × 768)
    blocks: list of 12 transformer decoder blocks
    W_out:  output projection (768 × vocab)
    """
    seq_len = len(tokens)

    # Step 1: token embedding + learned positional embedding
    h = W_emb[tokens] + W_pos[:seq_len]      # (seq_len, 768)

    # Step 2: pass through 12 masked self-attention blocks
    for block in blocks:
        h = block(h)  # causal mask inside: token i sees only 0..i

    # Step 3: project to vocabulary for next-token prediction
    logits = h @ W_out.T                       # (seq_len, vocab)
    return logits

# During fine-tuning, we take h[-1] (the last token's hidden state)
# and feed it to a task-specific linear layer:
#   y = softmax(h[-1] @ W_task)
# That's it. One new layer. Everything else is pre-trained.

Fitting all tasks into one sequence

A key innovation in GPT-1 is how it adapts to different downstream tasks without changing the model architecture. Instead of designing a unique head for each task, the authors convert every task into a text sequence with special delimiter tokens:

  • Classification: [Start] text [Extract] — the final token's is fed to a linear classifier.
  • Entailment: [Start] premise [Delim] hypothesis [Extract] — the model learns the relationship between two sentences.
  • Similarity: The two sentences are fed in both orders (A∥B and B∥A), and their representations are added element-wise before classification.
  • Multiple Choice / QA: Each answer option is concatenated with the context, scored independently, and softmaxed across options.

This "traversal-style" input formatting is a precursor to the that would later define GPT-3 and modern large language models. The insight: you don't need a new architecture for each task — you just need to phrase the task as text.

Open in Lab
See how GPT-1 reformats each task type into a single sequence.
The demo wakes as you arrive…

What transfers — and why layers matter

One of the paper's most revealing experiments: transferring only the first n layers of the pre-trained model and randomly initializing the rest. The result is striking — each additional transferred layer improves performance on downstream tasks by roughly 9% on MultiNLI.

This tells us something profound: each layer captures a different level of linguistic abstraction. Lower layers learn syntax and local patterns; higher layers learn semantics, discourse, and task-relevant reasoning. Throwing away the top layers means losing the most abstract — and most transferable — representations.

The paper also shows that without any fine-tuning at all (), the pre-trained model already performs reasonably on some tasks. As pre-training progresses, zero-shot performance steadily improves — evidence that the language modeling objective is learning genuine linguistic knowledge, not just surface statistics.

Open in Lab
Drag the slider to see how performance changes as more pre-trained layers are transferred.
The demo wakes as you arrive…

Results: 9 out of 12 benchmarks beaten

GPT-1 achieved state-of-the-art results on 9 of 12 tasks studied, with particularly striking improvements on tasks requiring long-range reasoning:

  • 8.9% absolute improvement on Stories Cloze Test (commonsense reasoning)
  • 5.7% improvement on RACE (multi-choice reading comprehension)
  • 1.5% improvement on MultiNLI (textual entailment)
  • 5.5% improvement on SciTail (scientific entailment)

The ablation studies reveal what matters most:

  • Removing pre-training entirely drops performance by 14.8% — the single biggest factor.
  • Replacing the Transformer with an LSTM drops performance by 5.6% on average — confirming that the architecture choice matters, not just the pre-training idea.
  • Removing the auxiliary language modeling loss during fine-tuning hurts larger datasets but helps smaller ones.
Open in Lab
See how removing each component (pre-training, Transformer, auxiliary loss) impacts average performance.
The demo wakes as you arrive…

A glimpse of zero-shot: emergent abilities before they had a name

Perhaps the most forward-looking finding: GPT-1 showed that the pre-trained model, without any fine-tuning, could perform some tasks at a reasonable level. The authors tracked zero-shot performance across pre-training and found it steadily improved — implying the language model was learning transferable task-relevant knowledge.

This was a faint signal of what GPT-2 would call "unsupervised multitask learning" and what GPT-3 would turn into few-shot and zero-shot prompting. The seeds of the engineering revolution were planted here.

Open in Lab
Watch zero-shot accuracy rise as pre-training progresses — the model learns task ability without task data.
The demo wakes as you arrive…

The lineage: from GPT-1 to the modern era

  1. 2018

    GPT-1 — 117M parameters

    Proved unsupervised pre-training transfers to downstream tasks. 12-layer decoder, BooksCorpus, 9/12 benchmarks beaten.

  2. 2018

    BERT — 110M / 340M parameters

    Flipped GPT-1's decoder to an encoder with bidirectional masking. Same "pre-train + fine-tune" paradigm, different direction. Dominated NLU benchmarks.

  3. 2019

    GPT-2 — 1.5B parameters

    Scaled up GPT-1 by 10×, trained on WebText. Showed zero-shot task performance without any fine-tuning — "language models are unsupervised multitask learners."

  4. 2020

    GPT-3 — 175B parameters

    Scaled 100× further. In-context learning emerged: describe the task in the prompt with a few examples and GPT-3 solves it — no fine-tuning at all.

  5. 2022

    InstructGPT & ChatGPT

    Added instruction tuning and RLHF to GPT-3.5. Turned a language model into a conversational assistant used by hundreds of millions.

  6. 2023

    GPT-4 — Multimodal

    Extended the decoder-only recipe to images + text. The architecture GPT-1 pioneered now processes vision, code, and language in one model.

GPT-1's deepest contribution was not a number on a leaderboard — it was a recipe: take a Transformer decoder, feed it enormous amounts of text, and fine-tune. That recipe, unchanged in its essence, powers every large language model you interact with today — including the one generating this summary.

CitationRadford, Narasimhan, Salimans, Sutskever. Improving Language Understanding by Generative Pre-Training. OpenAI Technical Report, 2018.

Terms in this paper