Language Models2018intermediate9 min read
Improving Language Understanding by Generative Pre-Training
تحسين فهم اللغة عبر التدريب المسبق التوليدي
Radford, A. · Narasimhan, K. · Salimans, T. · Sutskever, I. — OpenAI Technical Report
The problem
By 2018 most NLP systems were trained from scratch on labeled data for each task. Labeled data is expensive and scarce; each task — sentiment analysis, question answering, textual entailment — needed its own architecture and its own dataset. Unlabeled text was abundant but nobody had shown a general recipe for converting raw reading into broad language ability that transfers across tasks.
The contribution
A two-stage semi-supervised approach: (1) of a 12-layer on (~7,000 books) using next- prediction, then (2) on each downstream task with minimal architecture changes — just a linear output layer. Task-specific input transformations convert classification, entailment, similarity, and QA into a unified sequence format. An auxiliary language modeling objective during improves . The result: state-of-the-art on 9 of 12 benchmarks studied.
The impact
GPT-1 proved that unsupervised generative pre-training transfers to diverse supervised tasks — the "pre-train then fine-tune" paradigm that now powers all of modern NLP. It established the Transformer as a foundation model chassis, spawning GPT-2, GPT-3, GPT-4, and inspiring BERT (which flipped it to an ). Every today — Claude, Gemini, LLaMA — descends from this recipe.
Imagine a student who reads thousands of novels before ever taking a class. She never sees an exam question, yet she develops an intuition for how language works — grammar, reasoning, story structure. When she finally sits for a test, she only needs a few minutes of review to ace it.
GPT-1 is that student. Stage one: read books and predict the next word, millions of times. Stage two: glance at a handful of labeled examples and learn to classify, answer questions, or judge textual relationships. The reading did the heavy lifting; the exam prep is almost trivial.
The problem: labeled data is scarce, tasks are many
Before GPT-1, building an NLP system meant: collect a labeled dataset for your specific task, design a task-specific architecture, and train from scratch. Sentiment analysis? One model. Question answering? A different model. Textual entailment? Yet another.
This approach had three crippling weaknesses:
- Labeled data hunger. Annotating thousands of examples for every new task is slow and expensive.
- No knowledge transfer. Each model starts from zero — unable to reuse linguistic knowledge learned elsewhere.
- Architectural fragmentation. Every task demanded bespoke engineering, from mechanisms to output heads.
Meanwhile, the internet was overflowing with unlabeled text — books, articles, web pages — containing implicit knowledge about grammar, world facts, and reasoning patterns. The question was: how do you distill that raw text into a model that transfers to any task?
The idea: pre-train generatively, fine-tune discriminatively
GPT-1's recipe has two stages, and the elegance is in their simplicity:
Stage 1 — Unsupervised pre-training. Take a large corpus of unlabeled text (BooksCorpus: ~7,000 unpublished books). Train a Transformer decoder to predict the next word given all previous words. This is a standard language model objective — but applied to a Transformer decoder rather than an , it captures long-range dependencies far better.
Stage 2 — Supervised fine-tuning. Take the pre-trained model, add a single linear output layer, and train on labeled data for a specific task. The pre-trained weights provide a massive head start: the model already "understands" language, so it only needs to learn the task mapping.
The authors also add the language modeling as an auxiliary objective during fine-tuning. This acts as a regularizer — preventing the model from forgetting what it learned during pre-training, and improving generalization especially on larger datasets.
Now let's look at the formal objectives. Stage 1 maximizes the of each token given the previous k tokens:
During fine-tuning, the total loss combines the task-specific loss with the language modeling loss:
The architecture: a decoder-only Transformer
GPT-1 takes the Transformer decoder from the original Attention Is All You Need paper and removes the layer (since there is no separate encoder). What remains is a stack of masked layers — each token can only attend to tokens before it, never to future tokens. This ensures the model genuinely predicts the next token rather than copying it.
The specific configuration:
- 12 layers of Transformer decoder blocks
- 768-dimensional hidden states (the dimension)
- 12 attention heads per layer (each head operates on 64 dimensions)
- 3072-dimensional feed-forward intermediate layer (a 4× expansion)
- activation instead of — smoother gradients
- Learned positional embeddings (not sinusoidal)
- 512-token context window
- ~117 million total parameters
- tokenizer with ~40,000 merges
This is exactly the same size as BERT-Base — 12 layers, 768 dimensions, 12 heads. The only architectural difference: GPT-1 uses a causal mask (left-to-right), while BERT removes it (bidirectional). This makes their comparison a clean test of direction rather than scale.
Simplified to show the idea — not the real implementation.
import numpy as np
def gpt1_forward(tokens, W_emb, W_pos, blocks, W_out):
"""
tokens: list of token IDs, length ≤ 512
W_emb: token embedding matrix (vocab × 768)
W_pos: position embedding matrix (512 × 768)
blocks: list of 12 transformer decoder blocks
W_out: output projection (768 × vocab)
"""
seq_len = len(tokens)
# Step 1: token embedding + learned positional embedding
h = W_emb[tokens] + W_pos[:seq_len] # (seq_len, 768)
# Step 2: pass through 12 masked self-attention blocks
for block in blocks:
h = block(h) # causal mask inside: token i sees only 0..i
# Step 3: project to vocabulary for next-token prediction
logits = h @ W_out.T # (seq_len, vocab)
return logits
# During fine-tuning, we take h[-1] (the last token's hidden state)
# and feed it to a task-specific linear layer:
# y = softmax(h[-1] @ W_task)
# That's it. One new layer. Everything else is pre-trained.
Fitting all tasks into one sequence
A key innovation in GPT-1 is how it adapts to different downstream tasks without changing the model architecture. Instead of designing a unique head for each task, the authors convert every task into a text sequence with special delimiter tokens:
- Classification:
[Start] text [Extract]— the final token's is fed to a linear classifier. - Entailment:
[Start] premise [Delim] hypothesis [Extract]— the model learns the relationship between two sentences. - Similarity: The two sentences are fed in both orders (A∥B and B∥A), and their representations are added element-wise before classification.
- Multiple Choice / QA: Each answer option is concatenated with the context, scored independently, and softmaxed across options.
This "traversal-style" input formatting is a precursor to the that would later define GPT-3 and modern large language models. The insight: you don't need a new architecture for each task — you just need to phrase the task as text.
What transfers — and why layers matter
One of the paper's most revealing experiments: transferring only the first n layers of the pre-trained model and randomly initializing the rest. The result is striking — each additional transferred layer improves performance on downstream tasks by roughly 9% on MultiNLI.
This tells us something profound: each layer captures a different level of linguistic abstraction. Lower layers learn syntax and local patterns; higher layers learn semantics, discourse, and task-relevant reasoning. Throwing away the top layers means losing the most abstract — and most transferable — representations.
The paper also shows that without any fine-tuning at all (), the pre-trained model already performs reasonably on some tasks. As pre-training progresses, zero-shot performance steadily improves — evidence that the language modeling objective is learning genuine linguistic knowledge, not just surface statistics.
Results: 9 out of 12 benchmarks beaten
GPT-1 achieved state-of-the-art results on 9 of 12 tasks studied, with particularly striking improvements on tasks requiring long-range reasoning:
- 8.9% absolute improvement on Stories Cloze Test (commonsense reasoning)
- 5.7% improvement on RACE (multi-choice reading comprehension)
- 1.5% improvement on MultiNLI (textual entailment)
- 5.5% improvement on SciTail (scientific entailment)
The ablation studies reveal what matters most:
- Removing pre-training entirely drops performance by 14.8% — the single biggest factor.
- Replacing the Transformer with an LSTM drops performance by 5.6% on average — confirming that the architecture choice matters, not just the pre-training idea.
- Removing the auxiliary language modeling loss during fine-tuning hurts larger datasets but helps smaller ones.
A glimpse of zero-shot: emergent abilities before they had a name
Perhaps the most forward-looking finding: GPT-1 showed that the pre-trained model, without any fine-tuning, could perform some tasks at a reasonable level. The authors tracked zero-shot performance across pre-training and found it steadily improved — implying the language model was learning transferable task-relevant knowledge.
This was a faint signal of what GPT-2 would call "unsupervised multitask learning" and what GPT-3 would turn into few-shot and zero-shot prompting. The seeds of the engineering revolution were planted here.
The lineage: from GPT-1 to the modern era
2018
GPT-1 — 117M parameters
Proved unsupervised pre-training transfers to downstream tasks. 12-layer decoder, BooksCorpus, 9/12 benchmarks beaten.
2018
BERT — 110M / 340M parameters
Flipped GPT-1's decoder to an encoder with bidirectional masking. Same "pre-train + fine-tune" paradigm, different direction. Dominated NLU benchmarks.
2019
GPT-2 — 1.5B parameters
Scaled up GPT-1 by 10×, trained on WebText. Showed zero-shot task performance without any fine-tuning — "language models are unsupervised multitask learners."
2020
GPT-3 — 175B parameters
Scaled 100× further. In-context learning emerged: describe the task in the prompt with a few examples and GPT-3 solves it — no fine-tuning at all.
2022
InstructGPT & ChatGPT
Added instruction tuning and RLHF to GPT-3.5. Turned a language model into a conversational assistant used by hundreds of millions.
2023
GPT-4 — Multimodal
Extended the decoder-only recipe to images + text. The architecture GPT-1 pioneered now processes vision, code, and language in one model.
GPT-1's deepest contribution was not a number on a leaderboard — it was a recipe: take a Transformer decoder, feed it enormous amounts of text, and fine-tune. That recipe, unchanged in its essence, powers every large language model you interact with today — including the one generating this summary.
CitationRadford, Narasimhan, Salimans, Sutskever. Improving Language Understanding by Generative Pre-Training. OpenAI Technical Report, 2018.
Terms in this paper
- Generative Pre-Trainingالتدريب المسبق التوليدي
- Discriminative Fine-Tuningالضبط الدقيق التمييزي
- Task-Specific Input Transformationتحويل المدخلات الخاص بالمهمة
- Decoder-Onlyوحدة تفكيك ترميز فقط
- BooksCorpusمدوّنة الكتب (BooksCorpus)