Model Compression2019intermediate10 min read

DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter

DistilBERT: نسخة مُقطَّرة من BERT — أصغر وأسرع وأرخص وأخفّ

Sanh, V. · Debut, L. · Chaumond, J. · Wolf, T. — EMC @ NeurIPS 2019

The problem

By 2019, BERT and similar large pre-trained Transformers had become the standard for NLP, but their size — 110M+ parameters — made them expensive to run at time. Deploying BERT on edge devices, mobile phones, or under real-time constraints was impractical. Prior work on focused on task-specific : you compress a model only after it for one task. There was no general-purpose compressed model you could fine-tune on any downstream task.

The contribution

DistilBERT: a 6-layer pre-trained via task-agnostic from BERT-base. The student is trained with a triple loss — soft-target cross-entropy (), , and cosine-distance alignment of hidden states. The result: 40% smaller (66M vs 110M parameters), 60% faster at inference, yet retaining 97% of BERT's performance on the GLUE . DistilBERT is a general-purpose model — fine-tune it on any downstream task, just like BERT.

The impact

DistilBERT proved that knowledge distillation during — not just fine-tuning — produces compact general-purpose models. It became the go-to lightweight NLP backbone at Hugging Face and beyond, running in production on mobile and edge devices. The approach inspired a wave of distillation work: TinyBERT, MiniLM, DistilGPT-2, and the broader trend of compressing foundation models for deployment.

BERT is like building a massive encyclopedia — exhaustive, detailed, and expensive to carry around. What if you could train a student to read that encyclopedia and write a pocket guide that captures 97% of the knowledge?

That's knowledge distillation. The original BERT is the teacher: it has read millions of sentences and formed rich opinions about language. DistilBERT is the student: it has half the layers, but instead of learning from raw text alone, it watches the teacher's soft predictions — the subtle probability distributions over every word — and absorbs the teacher's "" about which wrong answers are almost right.

The problem: BERT is too big for the real world

BERT-base has 110 million parameters spread across 12 Transformer encoder layers. On a modern GPU it runs well, but in production — where you need to classify thousands of sentences per second, or run on a phone, or serve users with tight latency budgets — BERT is simply too large and too slow. Its model file weighs over 400 MB and a single inference pass on CPU can take hundreds of milliseconds.

The conventional approach by 2019 was task-specific distillation: first fine-tune BERT on your task, then compress that fine-tuned model into a smaller one. But this means you need a separate compressed model for every task — sentiment analysis, question answering, named entity recognition, each with its own distillation pipeline. This does not scale.

The question DistilBERT asks is: can we distill BERT during pre-training itself, to get a single compact model that works as a universal starting point for any task?

Open in Lab
Compare BERT-base and DistilBERT side-by-side in size, speed, and performance.
The demo wakes as you arrive…

Knowledge distillation: learning from a teachers soft opinions

Knowledge distillation, introduced by Hinton et al. (2015), is based on a key insight: a trained model's output probabilities contain far more information than the hard label alone. When BERT predicts a masked word and assigns 70% to "cat," 15% to "kitten," 10% to "dog," and 5% to "pet," those reveal that "kitten" is semantically close to "cat" while "dog" is a plausible but less likely alternative. A hard label that just says "cat" throws all of that relational knowledge away.

Think of it like a multiple-choice exam. The hard label says "the answer is A." But the teacher's soft probabilities say "A is very likely, B is somewhat plausible because it shares properties with A, C is unlikely but not absurd, and D is completely wrong." The student who sees the teacher's soft reasoning learns relationships between concepts, not just the right answer. Hinton called this "dark knowledge" — the knowledge hidden in the wrong answers.

Open in Lab
Compare hard labels vs soft targets. Raise the temperature to see how the probability distribution softens, revealing hidden relationships.
The demo wakes as you arrive…

The T controls how soft the probabilities become. At T=1 (standard ), the distribution is peaked — the top class dominates. As T increases, the distribution flattens, giving more weight to second and third choices. During distillation, both teacher and student use the same elevated temperature T so the student can learn from the teacher's full ranking, not just its top prediction.

pi=exp⁡(zi/T)∑jexp⁡(zj/T)p_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}
Softmax with temperature — z_i is the logit for class i. When T=1 this is standard softmax. As T increases, the output distribution becomes smoother, revealing the teacher's soft rankings among classes.

Architecture: half the layers, same width

DistilBERT keeps the same hidden size (768), the same number of attention heads (12), and the same feed-forward dimension (3072) as BERT-base — but uses only 6 encoder layers instead of 12. This halves the depth while preserving the width.

Why reduce depth rather than width? Because each layer of captures different linguistic phenomena: lower layers handle syntax, middle layers handle semantics, and upper layers handle task-level patterns. By taking every other layer from the teacher, you preserve the breadth of attention patterns at each layer. Shrinking the hidden size, on the other hand, would compress every layer's capacity uniformly, which hurts more.

Two additional simplifications reduce parameter count further: the token-type embeddings (used in BERT for sentence-pair tasks) are removed, and the pooler (the linear layer on top of [CLS]) is removed. These are not needed during pre-training and can be added back during fine-tuning if necessary.

Open in Lab
Explore how DistilBERT selects every other layer from BERT and initializes the student.
The demo wakes as you arrive…

The triple loss: three teachers in one

DistilBERT's training objective combines three loss functions, each teaching the student a different aspect of the teacher's knowledge. Think of it as a student learning from three channels simultaneously: what the teacher says (soft targets), what the textbook says (hard labels from MLM), and how the teacher thinks ( alignment).

Open in Lab
The three loss components and how they flow through the teacher-student architecture.
The demo wakes as you arrive…

1. Distillation loss (L_ce): The student's soft predictions (at temperature T) are compared to the teacher's soft predictions using (equivalent to cross-entropy with soft targets). This is the primary channel for transferring dark knowledge — the relational structure between vocabulary items that the teacher has learned.

2. Masked language modeling loss (L_mlm): The standard BERT training objective. The student predicts masked tokens from their hard ground-truth labels. This anchors the student to the actual data and prevents it from just mimicking the teacher's mistakes.

3. Cosine embedding loss (L_cos): The student's hidden states are aligned with the teacher's hidden states using . This ensures that the student does not just match the teacher's output predictions but also builds similar internal representations — preserving the geometry of the representation space.

L=α⋅Lce+β⋅Lmlm+γ⋅Lcos\mathcal{L} = \alpha \cdot \mathcal{L}_{\text{ce}} + \beta \cdot \mathcal{L}_{\text{mlm}} + \gamma \cdot \mathcal{L}_{\text{cos}}
Triple loss — The total training loss is a weighted combination of: the distillation cross-entropy L_ce (soft targets from the teacher), the masked language modeling loss L_mlm (hard labels from the data), and the cosine embedding loss L_cos (hidden state alignment).
Lce=−∑ipiTlog⁡qiTwherepiT=softmax(ziteacher/T),qiT=softmax(zistudent/T)\mathcal{L}_{\text{ce}} = -\sum_{i} p_i^T \log q_i^T \quad \text{where} \quad p_i^T = \text{softmax}(z_i^{\text{teacher}} / T), \quad q_i^T = \text{softmax}(z_i^{\text{student}} / T)
Distillation loss (KL divergence with temperature) — The teacher's logits and student's logits are both softened with temperature T before computing cross-entropy. Higher T reveals the ranking structure in the teacher's predictions.

Training recipe: borrowing from RoBERTa

DistilBERT borrows several training optimizations from RoBERTa that were shown to improve BERT's pre-training:

— instead of masking the same positions across epochs, the mask pattern is regenerated each time a sequence is fed to the model. This creates more diverse training examples from the same data.

No — BERT's original NSP objective was shown to be unnecessary by RoBERTa. DistilBERT drops it entirely, simplifying the training pipeline.

Large batches — training uses very large batches of up to 4,000 examples, following the finding that larger batches improve pre-training quality.

Code: distillation training loop

Simplified DistilBERT training steppython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn.functional as F

def distillation_step(student, teacher, batch, T=2.0, alpha=0.5):
    """One training step of DistilBERT distillation."""
    # Forward pass through both models
    student_out = student(batch["input_ids"], batch["attention_mask"])
    with torch.no_grad():
        teacher_out = teacher(batch["input_ids"], batch["attention_mask"])

    # 1) Distillation loss: KL divergence on softened logits
    s_logits = student_out.logits / T
    t_logits = teacher_out.logits / T
    loss_ce = F.kl_div(
        F.log_softmax(s_logits, dim=-1),
        F.softmax(t_logits, dim=-1),
        reduction="batchmean"
    ) * (T ** 2)  # Scale by T^2 to balance gradients

    # 2) MLM loss: hard labels from ground truth
    loss_mlm = F.cross_entropy(
        student_out.logits.view(-1, vocab_size),
        batch["labels"].view(-1),
        ignore_index=-100
    )

    # 3) Cosine embedding loss: align hidden states
    loss_cos = 1 - F.cosine_similarity(
        student_out.hidden_states[-1],
        teacher_out.hidden_states[-1],
        dim=-1
    ).mean()

    # Triple loss combination
    loss = alpha * loss_ce + (1 - alpha) * loss_mlm + loss_cos
    return loss

Results: 97% of the taste, 60% of the calories

DistilBERT was evaluated on the GLUE benchmark (9 language understanding tasks), the IMDb sentiment classification task, and SQuAD v1.1 (question answering).

On GLUE, DistilBERT scores 77.0 — retaining 97% of BERT-base's 79.5. It matches or exceeds the ELMo baseline on all tasks. On IMDb, the gap is even smaller: only 0.6% behind BERT. On SQuAD v1.1, DistilBERT is within 3.9 F1 points of BERT, and adding a second distillation step during fine-tuning narrows the gap further.

Speed-wise, excluding , DistilBERT is 71% faster than BERT on CPU. The entire model weighs ~207 MB — small enough for on-device deployment with further .

Open in Lab
GLUE task-by-task comparison of DistilBERT, BERT-base, and ELMo.
The demo wakes as you arrive…

On-device deployment: NLP in your pocket

The paper includes a proof-of-concept experiment deploying DistilBERT on a mobile device (an iPhone 7 Plus). Using a question-answering pipeline based on SQuAD, the model ran inference on the device with acceptable latency — demonstrating that the compression achieved by distillation makes real-time on-device NLP feasible.

This was a significant result in 2019: before DistilBERT, running a Transformer-based language model on an edge device was considered impractical. DistilBERT showed that the pre-train → distill → fine-tune → deploy pipeline could bring state-of-the-art NLP to constrained environments.

Model compression timeline

  1. 2015

    Knowledge Distillation (Hinton et al.)

    Introduced the teacher-student framework with soft targets and temperature scaling. Showed that dark knowledge in wrong predictions helps students learn.

  2. 2018

    BERT (Devlin et al.)

    Pre-trained bidirectional encoder with MLM and NSP. Set the standard for NLP but required significant compute for inference.

  3. 2019

    DistilBERT (Sanh et al.)

    Task-agnostic distillation during pre-training. 40% smaller, 60% faster, 97% of BERT's performance. First general-purpose distilled Transformer.

  4. 2020

    TinyBERT (Jiao et al.)

    Added attention transfer and intermediate layer distillation. Achieved 96% of BERT with 7.5× fewer parameters by distilling both attention maps and hidden states.

  5. 2020

    MiniLM (Wang et al.)

    Self-attention distillation from the teacher's last layer. Allowed flexible student architectures — width and depth can both differ from the teacher.

  6. 2021

    DeiT (Touvron et al.)

    Extended distillation to Vision Transformers. A distillation token learns from a CNN teacher, enabling data-efficient training of ViTs without massive datasets.

DistilBERT's legacy is not just a smaller model — it's a proof of concept that changed the field's approach to deployment. Before DistilBERT, the default was to distill after fine-tuning, producing one compressed model per task. DistilBERT showed that distilling during pre-training gives you a universal compressed backbone, just like BERT itself but cheaper. This pattern — pre-train large, distill once, fine-tune many — became the standard for practical NLP systems.

CitationSanh, Debut, Chaumond, Wolf. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. EMC @ NeurIPS 2019, 2019.

Terms in this paper