Language Models2019intermediate12 min read

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

استكشاف حدود نقل التعلُّم باستخدام محوِّل موحَّد يعمل بصيغة نصّ إلى نصّ

Raffel, C. · Shazeer, N. · Roberts, A. · Lee, K. · Narang, S. · Matena, M. · Zhou, Y. · Li, W. · Liu, P. J. — JMLR

The problem

By 2019, in NLP had become the dominant paradigm — pre-train on unlabeled text, fine-tune on your task. But the field was fragmented: BERT used a masked , GPT used a causal , and others mixed architectures, objectives, and datasets in incomparable ways. No one had done a systematic, controlled comparison. Researchers couldn't tell whether gains came from better architectures, better pre-training objectives, more data, or just more compute.

The contribution

T5 ( Transfer ): a unified framework that recasts every NLP task — translation, , , — as a text-to-text problem. The model takes a text input prefixed with a task description and produces a text output. Using this unification, the authors ran a massive systematic study comparing architectures ( vs. decoder-only vs. encoder-only), pre-training objectives (language model vs. BERT-style vs. ), unlabeled datasets (, a new 750 GB cleaned Common Crawl), strategies, and scale — all on the same footing. The best recipe: an encoder-decoder Transformer with span corruption pre-training on C4, scaled to 11 billion parameters, achieving state-of-the-art on GLUE, SuperGLUE, SQuAD, and WMT translation.

The impact

T5 became the reference architecture for the text-to-text paradigm. Its systematic study settled debates (encoder-decoder wins for generation, span corruption beats MLM, data quality trumps quantity). C4 became a community-standard dataset. T5's framework spawned a family: FLAN (instruction tuning), Switch Transformer (trillion- MoE), Prompt Tuning (parameter-efficient adaptation), FiD (retrieval-augmented generation), Imagen (text-to-image), and Chronos (Time Series Forecasting) — all built on the T5 text-to-text backbone.

Before T5, NLP was like a hospital where every department spoke a different language: the classification ward used checkboxes, the translation unit used bilingual dictionaries, and the summarization team used highlighters. Each department had its own tools, its own intake forms, and its own training programs.

T5 is the moment the hospital switches to a single universal medical record system: every department fills in the same form — plain text in, plain text out — and a single doctor (the model) rotates across all departments. The magic isn't the doctor; it's the standardized form that lets one doctor serve every ward.

The problem: a fragmented landscape of incomparable experiments

By late 2019 the transfer learning recipe — pre-train then fine-tune — had become standard. But every team cooked differently:

  • BERT used a masked encoder with 340M parameters trained on BooksCorpus + Wikipedia.
  • GPT-2 used a causal decoder with 1.5B parameters trained on WebText.
  • XLNet permuted the order. RoBERTa simply trained BERT longer.
  • ALBERT shared parameters; ERNIE injected entity knowledge.

Each paper changed multiple ingredients at once — architecture, objective, data, and scale — so it was impossible to isolate which changes actually mattered. The field was advancing fast but blindly: we had dozens of recipes and no controlled experiment to compare them.

Open in Lab
Each card shows a different pre-2019 approach. Notice how each changes multiple variables simultaneously — making fair comparison impossible.
The demo wakes as you arrive…

The core idea: every task is text-to-text

T5's key insight is simple but powerful: reframe every NLP task as generating text from text. The model receives an input string and produces an output string. What makes each task different is just a short prefix that tells the model which task to perform:

  • Translation: "translate English to German: That is good" → "Das ist gut"
  • Summarization: "summarize: The article discusses..." → "Scientists found..."
  • Classification: "sst2 sentence: This movie is great" → "positive"
  • Question answering: "question: When did... context: The event..." → "1969"

Even classification outputs are generated as literal text strings — "positive" or "negative" — rather than class indices. This means the same model, the same function, the same decoder, the same training loop handles every task. Nothing is task-specific except the prefix.

Open in Lab
Select a task to see how T5 reformats it as text-in, text-out.
The demo wakes as you arrive…

The architecture: encoder-decoder wins

T5 uses the original Transformer encoder-decoder architecture, closely following the design from the Is All You Need paper with a few modern modifications:

  • instead of sinusoidal positional encodings — each attention head learns a small table of bias values indexed by the distance between the and key positions. This lets the model generalize to sequences longer than those seen during training.
  • Pre-norm instead of post-norm: is applied before each sub-layer (attention or feed-forward), not after. This stabilizes training at scale.
  • No bias terms in the dense layers, and a simplified layer norm without the additive bias.

The encoder processes the input and builds a contextualized representation. The decoder attends to the encoder's output via and generates the output tokens one at a time. This two-part structure is crucial: the encoder can see all input tokens bidirectionally (no masking), while the decoder sees tokens only up to the current position (causal masking).

Open in Lab
Click any layer to understand its role. Notice the cross-attention that connects encoder to decoder.
The demo wakes as you arrive…

A critical finding in the paper: the authors compared three architecture variants head to head — encoder-decoder, decoder-only (like GPT), and prefix language model (decoder with bidirectional attention on the input). Each had roughly the same number of parameters and FLOPs. The encoder-decoder consistently outperformed the others, especially on generative tasks. Why? Because the encoder can use fully bidirectional attention to deeply understand the input, while the decoder specializes in generating the output — a division of labor that a decoder-only model can't replicate.

Open in Lab
Compare the three architectures. Notice how encoder-decoder separates understanding (bidirectional) from generation (causal).
The demo wakes as you arrive…

Pre-training objective: span corruption beats masking

Instead of masking individual tokens like BERT, T5 uses span corruption: randomly select and remove contiguous spans of tokens from the input, replace each span with a unique (like <extra_id_0>, <extra_id_1>, ...), and train the model to reconstruct only the dropped spans.

Think of it like a redacted document: imagine blacking out a few phrases from a paragraph and asking someone to fill in just the blanks — not rewrite the whole paragraph. The input is the document with black bars, and the target output is only the missing phrases, each tagged with which bar they came from.

Why is this better than BERT-style masking? Two reasons. First, corrupting spans rather than individual tokens creates harder, more informative predictions — guessing a 3-word phrase requires understanding context more deeply than guessing one word. Second, the target sequence is much shorter than the input — typically only ~15% of the original length — which means faster training because the expensive autoregressive decoder processes fewer tokens.

Open in Lab
Click "Corrupt" to see how T5 masks spans and builds its training target.
The demo wakes as you arrive…
Input: x1  ⟨s1⟩  x4  x5  ⟨s2⟩  x8Target: ⟨s1⟩  x2  x3  ⟨s2⟩  x6  x7\text{Input: } x_1 \; \langle s_1 \rangle \; x_4 \; x_5 \; \langle s_2 \rangle \; x_8 \qquad \text{Target: } \langle s_1 \rangle \; x_2 \; x_3 \; \langle s_2 \rangle \; x_6 \; x_7
Span corruption — sentinels mark where dropped spans belong — Original tokens x₂, x₃ and x₆, x₇ are removed and replaced with sentinel tokens ⟨s₁⟩ and ⟨s₂⟩ in the input. The target contains only the dropped tokens, each preceded by its sentinel. The model learns to reconstruct what's missing, not regurgitate the whole input.

The data: C4 — cleaning the internet at scale

The paper introduced the Colossal Clean Crawled Corpus (C4) — a 750 GB, ~800 billion token English dataset derived from Common Crawl. The cleaning pipeline was aggressive: remove duplicates, filter non-English pages, discard short or boilerplate text, remove pages with offensive content, and drop any page with too few sentences.

The systematic study showed that data quality matters more than data quantity. Training on a cleaner but smaller dataset consistently outperformed training on a larger but noisier one. C4 became a community standard — many subsequent models used it as their pre-training corpus.

The systematic study: a controlled experiment across everything

The heart of the T5 paper isn't a single model — it's the systematic comparison of every major design decision in transfer learning. The authors held all variables constant except the one under study, using the same compute budget, the same data, and the same evaluation. Here are the key findings:

Architecture: Encoder-decoder > decoder-only ≈ prefix LM, especially for generation. Pre-training objective: Span corruption > BERT-style masking > language modeling. Span corruption was most compute-efficient because the target sequence is short. Corruption rate: 15% of tokens corrupted works best (matching BERT's 15%). Span length: Average span length of 3 tokens outperformed single-token masking. Unlabeled data: C4 (cleaned web) > unfiltered web > domain-specific (Wikipedia alone). Training strategy: Multi-task pre-training followed by individual fine-tuning slightly trails single-task fine-tuning but enables a single checkpoint for all tasks. : Bigger models consistently outperform smaller ones, with 11B parameters (T5-XXL) achieving state-of-the-art. Ensembling multiple models further improves results.

Open in Lab
Toggle each variable to see which choices the T5 study found best.
The demo wakes as you arrive…

Scaling up: from 60M to 11B parameters

The authors trained five model sizes to study scaling behavior:

  • T5-Small: 60M parameters — a quick-to-train baseline.
  • T5-Base: 220M parameters — comparable to BERT-Base.
  • T5-Large: 770M parameters — the sweet spot for many tasks.
  • T5-3B: 3 billion parameters — strong performance across the board.
  • T5-11B (XXL): 11 billion parameters — state-of-the-art at the time.

The scaling trend was clear and monotonic: every increase in model size improved performance on every benchmark. But the paper also found diminishing returns: doubling parameters from 3B to 11B gave smaller gains than doubling from 220M to 770M. This observation presaged the Chinchilla scaling laws, which later showed that most large models were undertrained relative to their size.

Open in Lab
Compare T5 sizes across benchmarks. Notice the consistent but diminishing gains with scale.
The demo wakes as you arrive…

The idea in code

T5 span corruption — the pre-training data pipelinepython

Simplified to show the idea — not the real implementation.

import numpy as np

def span_corruption(tokens, corruption_rate=0.15, mean_span_length=3):
    """T5's span corruption: drop random spans, replace with sentinels.

    Returns (input_ids, target_ids) for the text-to-text format.
    """
    n = len(tokens)
    n_corrupt = max(1, int(n * corruption_rate))

    # Randomly decide which tokens belong to corrupted spans
    mask = np.zeros(n, dtype=bool)
    i = 0
    while mask.sum() < n_corrupt and i < n:
        if np.random.random() < corruption_rate:
            span_len = max(1, np.random.poisson(mean_span_length))
            span_len = min(span_len, n - i, n_corrupt - mask.sum())
            mask[i:i + span_len] = True
            i += span_len
        i += 1

    # Build input: original tokens with spans replaced by sentinels
    input_ids, target_ids = [], []
    sentinel_id = 0
    in_span = False

    for i, tok in enumerate(tokens):
        if mask[i]:
            if not in_span:             # start of a new span
                input_ids.append(f"<extra_id_{sentinel_id}>")
                target_ids.append(f"<extra_id_{sentinel_id}>")
                sentinel_id += 1
                in_span = True
            target_ids.append(tok)       # corrupted token goes to target
        else:
            input_ids.append(tok)         # clean token stays in input
            in_span = False

    return input_ids, target_ids

# Example:
tokens = ["The", "quick", "brown", "fox", "jumps", "over", "the", "lazy", "dog"]
inp, tgt = span_corruption(tokens)
# inp: ["The", "<extra_id_0>", "fox", "jumps", "<extra_id_1>", "lazy", "dog"]
# tgt: ["<extra_id_0>", "quick", "brown", "<extra_id_1>", "over", "the"]
# The model learns to fill in what's missing — short target = fast training.

Results: state-of-the-art across the board

T5-11B achieved state-of-the-art on virtually every benchmark tested:

  • GLUE: 89.7 average — surpassing RoBERTa and XLNet.
  • SuperGLUE: 88.9 — the first model to approach human performance (89.8).
  • SQuAD: Exact match 90.1 — competitive with the best specialized models.
  • CNN/DailyMail summarization: ROUGE-2 of 21.55 — a new record.
  • WMT English-to-German: 30.90 BLEU — state-of-the-art for an NLP generalist model.

Crucially, these results came from a single framework: one model architecture, one loss function, one training pipeline. No task-specific tricks, no architectural modifications, no special decoding strategies. The text-to-text approach proved that unification doesn't sacrifice performance.

What T5 unlocked

  1. 2019

    T5

    Unified text-to-text framework. Systematic comparison of architectures, objectives, data, and scale. Introduced C4 and span corruption.

  2. 2020

    mT5 — Multilingual T5

    Extended T5 to 101 languages using mC4. Same architecture, same text-to-text framework, proving the approach generalizes across languages.

  3. 2021

    FLAN — Instruction Tuning

    Fine-tuned T5 on 60+ NLP tasks phrased as instructions. Zero-shot performance soared — showing that the text-to-text format is a natural fit for instruction following.

  4. 2021

    Switch Transformer — Trillion Parameters

    Built on T5's architecture with Mixture of Experts layers. Achieved trillion-parameter scale with sublinear compute cost. Each token routes to one expert.

  5. 2021

    Prompt Tuning

    Instead of fine-tuning all T5 parameters, learn only a small set of "soft prompt" vectors prepended to the input. At scale, matches full fine-tuning with <0.1% of the parameters.

  6. 2021

    FiD — Fusion-in-Decoder

    T5's encoder processes retrieved passages independently; the decoder fuses them. A key architecture for retrieval-augmented generation.

  7. 2022

    Imagen — Text-to-Image Generation

    Used a frozen T5-XXL encoder as the text understanding backbone for image generation. Demonstrated that T5's text representations transfer even to visual domains.

  8. 2024

    Chronos — Time Series Forecasting

    Adapted T5 for time-series by tokenizing numerical values. Showed the text-to-text paradigm extends beyond language to sequential numerical data.

T5's deepest legacy is the text-to-text paradigm itself. Before T5, each NLP task had its own output format and training pipeline. After T5, the field converged on a unified interface: if you can phrase a task as "text in, text out," you can solve it with the same model. This idea is the direct ancestor of instruction-following models like ChatGPT and Claude — they simply expanded "text to text" from a training framework to a user interface.

CitationRaffel, Shazeer, Roberts, Lee, Narang, Matena, Zhou, Li, Liu. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. JMLR, 2020.

Terms in this paper