Language Models2019intermediate13 min read
BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension
BART: التدريب المسبق بإزالة التشويش لنماذج تسلسل إلى تسلسل في التوليد والترجمة والفهم اللغوي
Lewis, M. · Liu, Y. · Goyal, N. · Ghazvininejad, M. · Mohamed, A. · Levy, O. · Stoyanov, V. · Zettlemoyer, L. — ACL
The problem
By 2019, methods like BERT and GPT had shown great promise, but each was limited. BERT used a excellent for understanding tasks but awkward for generation, since it predicts tokens independently rather than autoregressively. GPT used an great for generation but blind to rightward context. Other approaches like XLNet and UniLM tried compromises, but no single model could match state-of-the-art on both comprehension and generation tasks. The field needed a unified framework that combined the strength of bidirectional encoding with autoregressive decoding.
The contribution
BART: a that pre-trains a full using arbitrary text corruption. The encoder reads corrupted text bidirectionally; the decoder reconstructs the original text left-to-right. This generalizes BERT (bidirectional encoder) and GPT (autoregressive decoder) into one model. The paper evaluates five noising strategies and finds that (replacing random spans with a single mask) combined with yields the best results. BART matches RoBERTa on GLUE and SQuAD, and sets new state-of-the-art on (up to +6 on XSum), dialogue, abstractive QA, and (+1.1 on WMT RO→EN).
The impact
BART proved that a single encoder-decoder model, pre-trained with flexible denoising, can unify understanding and generation. It became the backbone for RAG (Retrieval-Augmented Generation), mBART for multilingual translation, and numerous production summarization systems. Its text infilling objective influenced T5 and later models. BART demonstrated that the choice of corruption strategy matters as much as the architecture, opening a new axis for pre- research.
BERT reads text through a two-way mirror: every word sees every other word, but to prevent cheating it wears a blindfold over 15% of them. GPT reads through a one-way turnstile: each word can only see what came before it, perfect for writing the next word but blind to the future.
BART does something different. Imagine a document restoration workshop: someone takes a clean manuscript, vandalizes it in creative ways — scribbling over words, ripping out phrases, shuffling paragraphs — then hands it to an apprentice whose job is to produce a perfect copy of the original. The apprentice has two desks: at the first desk (the encoder), they study the damaged manuscript from every angle. At the second desk (the decoder), they write the restored version word by word, left to right. After thousands of restorations, the apprentice masters both understanding damaged text and generating clean text.
The gap: understanding OR generating, never both
Before BART, the pre-training landscape was split into two camps. On one side, BERT and its variants used bidirectional encoders — every sees every other token — ideal for , question answering, and any task where you need to understand a whole passage. But BERT predicts masked tokens independently (not one after the other), so it struggles to generate fluent text.
On the other side, GPT used an autoregressive decoder — each token only sees what came before it — perfect for text generation, but limited for understanding tasks because it cannot look ahead. It is like writing a story while covering the page below your pen: you can only build on what you have already written.
Models like XLNet, UniLM, and MASS tried to bridge this gap with clever masking schemes, but each made compromises. No single pre-training strategy could match state-of-the-art on both discriminative and generative benchmarks. The question was: can we design one model that inherits the encoder's comprehension and the decoder's generation ability simultaneously?
The idea: corrupt, then reconstruct
BART's core idea is elegant: pre-train a full encoder-decoder Transformer as a denoising autoencoder. The process has two steps. First, take a clean document and apply a noising function — any transformation that corrupts the text. Second, train the model to reconstruct the original document from the corrupted version. The encoder reads the corrupted text bidirectionally (like BERT), and the decoder produces the clean text autoregressively (like GPT).
The training objective is simply the between the decoder's output and the original document. This means BART learns to map any corrupted input back to the original text — and the flexibility is the key: unlike BERT which only masks individual tokens, BART can handle deletions, insertions, span replacements, and even sentence reordering.
Think of it as training a translator, but instead of translating between languages, BART translates between damaged text and clean text. The encoder must understand the damaged input deeply enough to pass its meaning forward, and the decoder must generate fluent output — exactly the two skills needed for downstream tasks.
Five ways to break text
A core contribution of BART is the systematic study of different text corruption strategies. The key insight is that how you corrupt the text determines what the model learns. Each strategy teaches the model a different skill:
Token Masking replaces random tokens with [MASK], identical to BERT. The model learns to predict individual words from context — good for local understanding, but the model always knows how many tokens are missing.
removes random tokens entirely. Now the model must figure out which positions are missing, not just what the words are. This is harder than masking because the input and output lengths differ.
Text Infilling samples random spans (with lengths drawn from a Poisson distribution, λ=3) and replaces each span with a single [MASK] token, regardless of span length. A span of five words becomes one mask. A span of zero words inserts a mask between existing words. The model must predict how many tokens are missing behind each mask — a skill no previous method taught.
Sentence Permutation shuffles the order of sentences in the document. This forces the model to learn about discourse structure and how sentences relate to each other — a skill crucial for summarization and long-form generation.
picks a random token and rotates the document so it starts with that token. The model must learn to identify the true beginning of the document.
The architecture: the full Transformer, reunited
BART's architecture is the standard sequence-to-sequence Transformer from the original Transformer paper — an encoder and a decoder, just like what was designed for machine translation. While BERT uses only the encoder half and GPT uses only the decoder half, BART uses the complete architecture.
The encoder is bidirectional: every token attends to every other token with no . The decoder is autoregressive: each token can only attend to previous tokens (via a causal mask) plus the encoder output (via ). This combination means BART can understand input from all directions and generate output sequentially.
Two model sizes were explored. BART-Base has 6 encoder layers and 6 decoder layers with 768 hidden dimensions. BART-Large has 12 layers each with 1024 hidden dimensions. Both use activations instead of and initialize parameters from a N(0, 0.02). BART-Large contains roughly 10% more parameters than a similarly-sized BERT model because of the cross-attention layers in the decoder.
Fine-tuning: one model, four modes
Because BART has both an encoder and a decoder, it naturally adapts to a wider range of tasks than encoder-only or decoder-only models:
Sequence classification (sentiment analysis, NLI): Feed the same input to both encoder and decoder. Use the final decoder token's hidden state — which has attended to the entire input — as the classification vector. This is analogous to BERT's [CLS] token, but placed at the end so it can attend to the full decoder output.
Token classification (NER, SQuAD): Feed the full document to encoder and decoder. Use each decoder token's final hidden state to classify that position.
Sequence generation (summarization, abstractive QA): The most natural fit. The encoder receives the input document and the decoder generates the output autoregressively. This directly mirrors the pre-training setup — corrupted text in, clean text out — which is why BART excels at generation tasks.
Machine translation: A creative approach. Replace BART's encoder layer with a new randomly initialized encoder trained on the source language. The new encoder learns to map foreign text into the same representation space that BART's decoder was pre-trained to denoise. This lets BART act as a pre-trained target-language model for translation, even though it was only pre-trained on English.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def text_infill(tokens, mask_id, mask_ratio=0.30, poisson_lambda=3):
"""BART's text infilling: replace random spans with a single [MASK].
Span lengths are drawn from Poisson(λ=3). 0-length spans insert a [MASK]."""
n = len(tokens)
num_masked = int(n * mask_ratio)
corrupted = list(tokens)
masked_count = 0
insert_positions = []
while masked_count < num_masked:
# Sample span length from Poisson distribution
span_len = np.random.poisson(poisson_lambda)
span_len = min(span_len, num_masked - masked_count)
# Pick a random start position
start = np.random.randint(0, max(1, len(corrupted) - span_len + 1))
# Replace the span with a single [MASK] token
corrupted[start:start + span_len] = [mask_id]
masked_count += span_len
return corrupted # Shorter than original! Decoder must reconstruct full text.
def bart_loss(encoder, decoder, original_tokens, corrupted_tokens):
"""One training step: encode corrupted text, decode to original."""
# Encoder: bidirectional (no causal mask), reads corrupted input
memory = encoder(corrupted_tokens) # full attention over corrupted text
# Decoder: autoregressive, reconstructs original text left-to-right
loss = 0
for t in range(len(original_tokens)):
# Cross-attention lets decoder see encoder output
logits = decoder(original_tokens[:t], memory)
loss += cross_entropy(logits, original_tokens[t])
return loss / len(original_tokens)
# Key difference from BERT: input and output lengths can differ!
# A span of 5 words → 1 [MASK]. The decoder must figure out how many tokens.Results: strong everywhere, dominant in generation
BART-Large was pre-trained with text infilling (masking 30% of tokens) and sentence permutation, using the same 160GB corpus as RoBERTa and training for 500,000 steps with a of 8000.
Discriminative tasks: BART matched RoBERTa on GLUE and SQuAD, despite having a decoder that is not needed for these tasks. This proved that adding a decoder does not hurt understanding performance — a concern many researchers had.
Summarization: BART achieved new state-of-the-art on both CNN/DailyMail (44.16 ROUGE-1) and XSum (45.14 ROUGE-1), a gain of roughly 6 ROUGE points on XSum over the previous best. XSum requires highly abstractive summaries, making it the perfect test for BART's generation abilities.
Dialogue: On ConvAI2, BART outperformed all previous methods in both F1 and .
Abstractive QA: On ELI5, BART set a new state-of-the-art with 30.6 ROUGE-1.
Machine translation: Using the stacked encoder approach, BART improved over a strong back-translation baseline by 1.1 BLEU on WMT Romanian→English, reaching 37.96 BLEU — using only English pre-training.
Ablation insights: what matters most
The controlled ablation study across six tasks revealed several critical lessons about pre-training design:
Token-level corruption is essential. Rotation and sentence shuffling alone performed poorly. The methods that worked all involved token masking or deletion. The model needs fine-grained word-level signal during pre-training.
Left-to-right generation matters. The (BERT-style) and permuted language model underperformed on generation tasks. These are the only approaches that lack autoregressive left-to-right decoding during pre-training. Generating coherent text requires practice at sequential generation.
Bidirectional encoding is crucial for understanding. A pure left-to-right language model collapsed on SQuAD (76.7 F1 vs 90.8 for BART), confirming that comprehension tasks need the encoder to see the full context.
The pre-training objective is not the only factor. Controlling for data and optimization, different objectives led to very different results. But even within the same architecture, hyperparameters like learning rate and placement mattered significantly.
What BART unlocked
2019
BART
Denoising encoder-decoder pre-training. Unified understanding and generation. State-of-the-art on summarization, dialogue, and abstractive QA.
2020
mBART — Multilingual BART
Extended BART's denoising to 25 languages. Showed that multilingual denoising pre-training dramatically improves machine translation, especially for low-resource language pairs.
2020
RAG — Retrieval-Augmented Generation
Used BART as the generator in a retrieve-then-generate pipeline. Combined a retriever (DPR) with BART's decoder to answer questions using retrieved documents. Became a foundation for modern RAG systems.
2020
T5 — Text-to-Text Transfer Transformer
Independently explored span corruption similar to BART's text infilling. Unified all NLP tasks as text-to-text, validating BART's insight that encoder-decoder pre-training is highly versatile.
2020
PEGASUS
Designed a pre-training objective specifically for summarization by masking entire sentences. Conceptually similar to BART's sentence-level corruption, but specialized for abstractive summarization.
BART's deepest insight is that pre-training is restoration, not prediction. BERT predicts masked tokens. GPT predicts next tokens. But BART frames pre-training as reconstructing corrupted text — a more general objective that naturally spans understanding and generation. This framing opened the door to creative corruption strategies and showed that the diversity of corruption matters as much as the architecture. Every modern encoder-decoder model, from T5 to mBART to PEGASUS, builds on this foundation.
CitationLewis, Liu, Goyal, Ghazvininejad, Mohamed, Levy, Stoyanov, Zettlemoyer. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. ACL, 2020.
Terms in this paper
- Denoising Autoencoderالمرمِّز الذاتي لإزالة الضوضاء
- Sequence-to-Sequenceتسلسل إلى تسلسل
- Text Infillingملء النص
- Span Corruptionتشويه النطاقات
- Sentence Permutationخلط ترتيب الجمل
- Token Deletionحذف الرموز
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Autoregressiveذاتي الانحدار
- Pre-trainingالتدريب المسبق
- Fine-Tuningالضبط الدقيق
- Summarizationالتلخيص الآلي
- Machine Translationالترجمة الآلية
- Beam Searchبحث الحزمة
- Cross Entropyالعشوائية المتقاطعة
- Byte Pair Encoding (BPE)ترميز زوج البايت