Language Models2018intermediate8 min read

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

BERT: التدريب المسبق لمحوِّلات عميقة ثنائية الاتجاه لفهم اللغة

Devlin, J. · Chang, M.-W. · Lee, K. · Toutanova, K. — NAACL

The problem

By 2018 language models were unidirectional: GPT read left-to-right, ELMo concatenated two separate directions as a shallow afterthought. Neither built truly deep representations. Worse, every NLP task — sentiment, QA, NER — needed its own bespoke architecture. There was no single pre-trained model you could fine-tune for everything.

The contribution

BERT: a pre-trained with two self-supervised tasks — , where 15% of tokens are masked and the model predicts them from full bidirectional context, and , which teaches sentence-pair understanding. The result is a single pre-trained model that can be fine-tuned with just one added output to achieve state-of-the-art on 11 NLP benchmarks, including GLUE (80.5%), SQuAD v1.1 (F1 93.2), and SQuAD v2.0 (F1 83.1).

The impact

BERT established the pre-train → fine-tune paradigm that dominates NLP today. It proved that deep bidirectional context is crucial and that a single architecture can master many tasks. BERT spawned a family — RoBERTa, ALBERT, DistilBERT, SpanBERT, XLM — and its encoder lives on inside every retrieval system, search engine, and model. It was the bridge between the Transformer paper and the GPT scaling era.

GPT reads a sentence like peeking through a keyhole that only slides left-to-right: at each word you can see everything before it, but the future is dark. ELMo trains two separate keyholes — one left-to-right, one right-to-left — then tapes their views together at the end, a shallow compromise.

BERT removes the wall entirely. It sits inside a glass room: every word sees every other word simultaneously. To prevent the model from simply copying the answer, BERT blindfolds 15% of the words and says: "Use everyone else to guess who I am."

The problem: unidirectional models see half the picture

Consider the sentence: "The bank of the river was covered in wildflowers."

A left-to-right model reaching "river" has already read "The bank of the" — which could mean a financial institution. Only the future words ("was covered in wildflowers") disambiguate "bank" as a riverbank. A unidirectional model must guess the meaning before seeing the evidence.

ELMo partly addresses this by running two LSTMs — one forward, one backward — and concatenating their hidden states. But this is shallow: the two directions never interact during encoding. Each direction builds its in isolation, then they're pasted together. The left context and right context never inform each other at the attention level.

Open in Lab
Watch how GPT, ELMo, and BERT build context for the word "bank" differently.
The demo wakes as you arrive…

The idea: mask, then predict from all directions

BERT's breakthrough is simple: use the Transformer encoder (not the decoder — no ) so every can attend to every other token. But if every token sees every token, what stops the model from cheating — just copying the target word from itself?

The answer: Masked Language Modeling (MLM). Before feeding a sentence to BERT, randomly pick 15% of tokens and replace them. Of those 15%:

  • 80% are replaced with [MASK]
  • 10% are replaced with a random token
  • 10% are kept unchanged

The model must predict the original token at each masked position. Because the mask could be anywhere, every token must build a representation that's useful for any prediction — forcing deep bidirectional understanding.

Open in Lab
Click a word to mask it, then see BERT's prediction using full bidirectional context.
The demo wakes as you arrive…

The second task: next sentence prediction

Many tasks (, ) require understanding the relationship between two sentences. MLM alone doesn't train for this. So BERT adds a second pre-training task:

Next Sentence Prediction (NSP): Given a pair [Sentence A, Sentence B], predict whether B actually follows A in the original text (label: IsNext) or was randomly sampled from elsewhere (label: NotNext). 50/50 split.

The input format uses special tokens: [CLS] Sentence A [SEP] Sentence B [SEP]. The [CLS] token's final representation is used for the binary . This format carries over to fine-tuning: [CLS] becomes the classification vector for any task.

Open in Lab
Can you guess which sentence pairs are consecutive? Click to reveal BERT's answer.
The demo wakes as you arrive…

The architecture: the Transformer encoder, twice

BERT's architecture is the Transformer encoder from the original paper — no decoder, no causal mask. The innovation is purely in the pre-training tasks, not the architecture.

Two sizes:

  • BERT-Base: 12 layers, 768 hidden dim, 12 attention heads — 110M parameters
  • BERT-Large: 24 layers, 1024 hidden dim, 16 attention heads — 340M parameters

BERT-Base was designed to match GPT-1's size for a fair comparison. The input representation is the sum of three embeddings: token embeddings ( vocabulary of 30,000), segment embeddings (marking Sentence A vs B), and position embeddings (learned, up to 512 tokens).

Open in Lab
Click any layer to learn its role inside BERT.
The demo wakes as you arrive…
Input=Token_Emb+Segment_Emb+Position_Emb\text{Input} = \text{Token\_Emb} + \text{Segment\_Emb} + \text{Position\_Emb}
BERT input representation — three embeddings summed — Token embedding maps each WordPiece to a vector. Segment embedding marks which sentence a token belongs to (A or B). Position embedding encodes order. Their sum is the input to Layer 1.

Fine-tuning: one model, many tasks

BERT's most impactful idea may be how simple fine-tuning is. The entire pre-trained model is reused unchanged — you only add one thin output layer on top:

  • Classification (sentiment, NLI): take the [CLS] token → linear layer → .
  • Token-level tasks (NER, POS tagging): take each token's output → linear layer → label per token.
  • Question answering (SQuAD): take each token's output → predict start and end positions of the answer span.
  • Sentence-pair tasks: encode both sentences with [SEP] between them, classify from [CLS].

All parameters — including the pre-trained ones — are fine-tuned jointly. This typically takes 2–4 epochs with a around 2e-5 to 5e-5. No task-specific architecture needed.

Open in Lab
Click a task to see how the same BERT model adapts with one output layer.
The demo wakes as you arrive…

The idea in code

BERT masked language modeling — the core training looppython

Simplified to show the idea — not the real implementation.

import numpy as np

def mask_tokens(tokens, vocab_size, mask_id, mask_prob=0.15):
    """Apply BERT's 80/10/10 masking strategy."""
    masked = tokens.copy()
    labels = np.full_like(tokens, -1)     # -1 = don't predict here

    for i in range(len(tokens)):
        if np.random.random() < mask_prob:
            labels[i] = tokens[i]          # remember the real token
            roll = np.random.random()
            if roll < 0.8:
                masked[i] = mask_id        # 80% → [MASK]
            elif roll < 0.9:
                masked[i] = np.random.randint(vocab_size)  # 10% → random
            # else: 10% → keep original (masked[i] unchanged)
    return masked, labels

def bert_mlm_loss(model, tokens, mask_id, vocab_size):
    """One training step for BERT's MLM objective."""
    masked_input, labels = mask_tokens(tokens, vocab_size, mask_id)

    # Forward: full bidirectional attention — no causal mask!
    hidden = model.encode(masked_input)    # (seq_len, d_model)

    loss = 0
    count = 0
    for i in range(len(tokens)):
        if labels[i] != -1:                # only compute loss at masked positions
            logits = hidden[i] @ model.vocab_proj.T  # predict original token
            loss += cross_entropy(logits, labels[i])
            count += 1
    return loss / count

# The ENCODER is the same Transformer encoder from "Attention Is All You Need."
# The only difference: no causal mask. Every token sees every other token.
# After pre-training, fine-tuning = same model + one linear layer on top.

Results: 11 benchmarks, all at once

BERT-Large set state-of-the-art records across the board:

  • : 80.5% average — a 7.7-point absolute improvement over the previous best. BERT dominated all 8 tasks from sentiment analysis to textual entailment.
  • SQuAD v1.1: F1 93.2, surpassing human performance (91.2) for the first time on this reading comprehension benchmark.
  • SQuAD v2.0: F1 83.1 — a 5.1-point improvement. This harder version includes questions with no answer in the passage.
  • SWAG (commonsense reasoning): 86.3% accuracy — a massive jump that showed BERT captures world knowledge, not just syntax.

Critically, all these results came from the same pre-trained model with different one-layer heads. No task-specific architecture engineering needed.

What BERT unlocked

  1. 2018

    BERT

    Masked language modeling + bidirectional encoder. Pre-train once, fine-tune everywhere. 11 benchmarks broken at once.

  2. 2019

    RoBERTa — Robustly Optimized BERT

    Dropped NSP, trained longer on more data, dynamic masking. Same architecture, better recipe. Showed BERT was undertrained.

  3. 2019

    ALBERT — A Lite BERT

    Parameter sharing across layers + factorized embedding. 18× fewer parameters with competitive performance.

  4. 2019

    DistilBERT

    Knowledge distillation from BERT-Base into a 6-layer model. 60% the size, 97% the performance, 60% faster.

  5. 2019

    XLNet

    Replaced masking with permutation language modeling — captures bidirectional context without the [MASK] token mismatch issue.

  6. 2020

    ELECTRA

    Replaced MLM with replaced-token detection. Every token gets a training signal (not just 15%), so it trains 4× more efficiently.

  7. 2020

    Sentence-BERT

    Fine-tuned BERT with siamese networks for efficient sentence similarity. Made BERT embeddings practical for search and retrieval.

BERT's deepest legacy is the paradigm, not the parameters. Before BERT, every NLP task needed a custom architecture. After BERT, the recipe became universal: (1) pre-train a large encoder on unlabeled text, (2) fine-tune on your specific task with a thin head. This pattern flows directly into the GPT line — different pre-training objective (next-token instead of masking), but the same insight that scale + pre-training + fine-tuning is all you need.

CitationDevlin, Chang, Lee, Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL, 2019.

Terms in this paper