Model Compression2015intermediate10 min read

Distilling the Knowledge in a Neural Network

تقطير المعرفة في الشبكات العصبية

Hinton, G. · Vinyals, O. · Dean, J. — NeurIPS Deep Learning Workshop

The problem

Large ensembles of deep neural networks achieve the highest , but deploying many big models together is impractical — too slow, too expensive, too memory-hungry for mobile devices or real-time systems. Meanwhile, a single small model from scratch on the same data yields noticeably worse . How do you compress the collective knowledge of a cumbersome into one compact model that can actually be deployed?

The contribution

: train a small "student" network to mimic the soft probability outputs of a large "teacher" (or ensemble) by raising the temperature. At high temperature, the teacher's outputs reveal "" — the relative probabilities of wrong classes — which encodes similarity structure the hard labels throw away. The student optimizes a weighted sum of two objectives: matching the teacher's via , and matching the ground-truth hard labels via cross-entropy. The paper also introduces specialist models — sub-ensembles that handle confusable classes — and shows that an ensemble of specialists can be distilled into one general student.

The impact

Knowledge distillation became the standard method for . DistilBERT used it to compress BERT to 60% its size while retaining 97% performance. DeiT applied teacher- student distillation to train Vision Transformers without massive datasets. , self-distillation, and speculative decoding all trace intellectual lineage to this paper. It proved that the knowledge in a trained network is not its weights — it is the distribution over outputs, including the "wrong" answers.

Imagine a seasoned professor teaching a new instructor. Instead of handing over the textbook (the raw data), the professor shares graded exam papers: "This student wrote 7, but notice the stroke leans toward a 1, and the top curl almost looks like a 9."

Those nuances — the partial credit, the near-misses — carry more insight than a simple answer key ever could. Knowledge distillation works the same way: a big trained model plays the professor, passing its soft judgments to a small so it can learn not just the answers, but the reasoning behind the ambiguity.

The problem: big models know too much to deploy

By 2015, the best-performing systems on benchmarks like ImageNet were ensembles — often 5 to 10 independently trained deep networks whose predictions were averaged. This brute- force approach consistently outperformed any single model.

But there is a sharp divide between the training phase and the deployment phase. During training, we want to extract every drop of knowledge from the data — computational cost and latency do not matter. During deployment, we need a model that responds in milliseconds on a phone or serves millions of requests per second. An ensemble of ten giant networks cannot do either.

The naive solution — train one small model directly on the data — consistently generalizes worse. The small model sees only the hard labels ("this is a 3"), and misses the rich structure the ensemble discovered: that this particular 3 was also slightly similar to an 8 and not at all like a 7.

Open in Lab
Compare hard labels (one-hot) vs. soft labels from a teacher. Notice how soft labels preserve similarity structure between classes.
The demo wakes as you arrive…

Dark knowledge: the information hiding in wrong answers

When a model trained on digit recognition sees a picture of the digit 2, it might output: "probability of 2 = 0.93, probability of 3 = 0.05, probability of 7 = 0.01, everything else near zero." The hard label says only "this is a 2." But the soft output says much more — it says the model sees some resemblance between 2 and 3 (both have curves), a faint similarity to 7 (open top stroke), and almost none to 0 or 1.

Hinton calls these subtle probability ratios among incorrect classes dark knowledge. They encode structural relationships between categories — relationships that hard labels completely destroy. A 2 that almost looks like a 3 carries more information than a 2 that looks nothing like anything else.

The problem is that a well-trained model's softmax outputs are extremely peaked — the correct class gets probability near 1.0, and all that rich relational information is crushed into tiny, negligible tail probabilities. How do we uncover the hidden structure?

Temperature: softening the softmax

The key mechanism is a TT applied to the softmax function. In the standard softmax (T=1T = 1), the output distribution is sharp — the largest dominates. As TT increases, the distribution softens: all classes get more balanced probability mass, and the relative differences between wrong-class probabilities become visible.

Think of TT as a knob on a microscope: at T=1T = 1 you see the image with naked eyes — only the dominant shape (correct class) is visible. Raise TT and you zoom into the fine grain, revealing subtle texture differences (dark knowledge) that were invisible at normal scale.

qi=exp⁡(zi/T)∑jexp⁡(zj/T)q_i = \frac{\exp(z_i / T)}{\sum_j \exp(z_j / T)}
Softmax with Temperature — Temperature controls how confident or uncertain the output probability distribution appears. Higher temperatures spread probability more evenly across multiple classes, revealing relative similarities between them. Lower temperatures concentrate probability on the most likely class, producing sharper and more decisive predictions. Standard softmax is a special case obtained with the default temperature setting.
Open in Lab
Drag the temperature slider and watch how the probability distribution changes. Notice how higher T reveals the hidden structure between classes.
The demo wakes as you arrive…

The distillation objective: two losses, one student

The student model is trained with a weighted combination of two terms:

Soft loss (distillation loss): the KL divergence between the teacher's soft output (at temperature TT) and the student's soft output (also at temperature TT). This loss transfers the dark knowledge — it tells the student not just "what" but "how much like what else."

Hard loss (standard loss): the cross-entropy between the student's output (at T=1T=1) and the ground-truth one-hot labels. This anchors the student to real-world correctness.

The total loss is:

L=α⋅T2⋅KL ⁣(σ(zT/T)  ∥  σ(zS/T))+(1−α)⋅CE ⁣(y,  σ(zS))\mathcal{L} = \alpha \cdot T^2 \cdot \text{KL}\!\left(\sigma(z_T / T) \;\|\; \sigma(z_S / T)\right) + (1 - \alpha) \cdot \text{CE}\!\left(y,\; \sigma(z_S)\right)
Combined Distillation Loss — This objective combines two learning signals. One encourages the student model to imitate the teacher's output distribution, allowing it to learn nuanced relationships between classes rather than only the correct answer. The other trains the student using the ground-truth labels from the dataset. A weighting factor controls the balance between these two sources of supervision, while a temperature setting softens the teacher's predictions to reveal richer information about its confidence and class preferences.
Open in Lab
Follow the data flow from the teacher to the student. Click each stage to see what happens at that step.
The demo wakes as you arrive…

Why it works: information-rich supervision

The central insight is about information density per training sample. A hard label provides a single bit of information: which class is correct. A soft probability vector from the teacher provides many bits: the relative similarity of the input to every class, as discovered by a model that has already generalized well.

This means the student gets a much richer gradient signal per sample. Hinton showed that with soft targets, you can train with far fewer samples or far fewer epochs than training on hard labels alone. In fact, on MNIST, a distilled model using only 3% of the training set (transferred from a teacher) outperformed a model trained on 100% of the data with hard labels.

There is also a effect. prevent the student from becoming over-confident on any single class, because even for a clear "3", the teacher assigns small but non-zero probability to visually similar digits. This acts as implicit regularization — it penalizes over-confident predictions — very much like the technique later formalized as label smoothing.

Open in Lab
Select different digit images and see what the teacher model's soft predictions reveal about class similarities that hard labels miss.
The demo wakes as you arrive…

The recipe: step by step

The knowledge distillation process follows a clear pipeline:

Step 1 — Train the teacher. Build a large model (or ensemble) and train it to on the full dataset. This is the computationally expensive part, done once.

Step 2 — Generate soft targets. Run the training data through the teacher at temperature TT (typically TT = 3–20). Save the soft probability vectors.

Step 3 — Train the student. Train a smaller architecture on the combined loss: match the teacher's soft targets (KL divergence at temperature TT) while also matching the true labels (cross-entropy at T=1T = 1).

Step 4 — Deploy the student at T=1T = 1 for . The student now captures the teacher's generalization ability in a fraction of the parameters.

Open in Lab
Watch the student model train. The loss curve shows how it converges faster with soft targets than with hard labels alone.
The demo wakes as you arrive…

Specialist models: dividing the confusion

Large classification tasks (like JFT with 15,000 classes) have clusters of easily confused categories — different dog breeds, different car models, different mushroom species. The generalist teacher wastes capacity distinguishing a German Shepherd from a Siberian Husky when a specialized sub-model could do it better.

Hinton proposes specialist models: small networks each trained on a confusable cluster (e.g. "all dog breeds" or "all mushroom species"), plus a random sample of other classes to prevent to the cluster. During inference, the generalist model first identifies which specialist is relevant, then the specialist refines the prediction.

The key result: an ensemble of one generalist plus several specialists was distilled into a single student that outperformed a comparable ensemble of purely generalist models — with far less total compute for training.

Open in Lab
Explore how specialist models handle confusable class clusters, and how their knowledge is combined during distillation.
The demo wakes as you arrive…

Results: small model, big performance

On MNIST, a single distilled model matched the accuracy of the original ensemble. More remarkably, the distilled model handled the digit "3" correctly even when all 3s were removed from its transfer set — the dark knowledge about 3's similarity to other digits (via the teacher's soft targets on non-3 examples) was sufficient.

On a large-scale internal speech task at Google (the system), an ensemble of 10 models was distilled into a single model of the same architecture. The single distilled model approached the ensemble's word error rate while being 10× cheaper to serve.

The experiments on JFT demonstrated that a mixture of specialist and generalist knowledge, once distilled, could match or exceed baselines trained with more compute.

Open in Lab
Compare the accuracy and efficiency of distilled models vs. ensembles and individually trained baselines.
The demo wakes as you arrive…

What knowledge distillation unlocked

  1. 2015

    Knowledge Distillation (this paper)

    Showed that soft targets from a teacher encode dark knowledge — similarity structure between classes — and that a small student can inherit it via temperature-scaled softmax.

  2. 2017

    Label Smoothing

    Replaced hard labels with a mixture of the one-hot and a uniform distribution. Conceptually a special case of distillation where the "teacher" is a uniform distribution.

  3. 2019

    DistilBERT

    Applied distillation to compress BERT. 60% of BERT's size, 97% of its performance, 60% faster inference. Brought Transformers to resource-constrained deployment.

  4. 2021

    DeiT — Data-efficient Image Transformers

    Used a CNN teacher to distill knowledge into a Vision Transformer, eliminating the need for JFT-scale pretraining. Introduced a distillation token alongside the class token.

  5. 2021

    DINO — Self-distillation with no labels

    Self-distillation — the student and teacher are the same architecture, and the teacher is an exponential moving average of the student. No labels needed. Learned features that spontaneously segment objects.

  6. 2023

    Speculative Decoding

    A small draft model proposes tokens quickly, and a large verifier model checks them in parallel. The same teacher-student asymmetry as distillation, applied to inference speed.

Distillation's deepest lesson is philosophical: the knowledge in a is not its weights or its architecture — it is the function it has learned, expressed as the distribution over outputs for any given input. Weights are a container; the probability distribution is the content. Once you accept that, you see that knowledge can be poured from one container to another — from a big model to a small one, from an ensemble to a single model, from a teacher to a student.

CitationHinton, Vinyals, Dean. Distilling the Knowledge in a Neural Network. NeurIPS Deep Learning Workshop, 2015.

Terms in this paper