Computer Vision2021intermediate11 min read

Training Data-Efficient Image Transformers & Distillation Through Attention

تدريب محوِّلات الصور الكفوءة بيانياً والتقطير عبر الانتباه

Touvron, H. · Cord, M. · Douze, M. · Massa, F. · Sablayrolles, A. · Jégou, H. — ICML

The problem

By 2020 the had shown that pure -based architectures could rival CNNs on — but only after pre- on JFT-300M, a private Google dataset of 300 million labeled images, using enormous compute. Trained on ImageNet alone, ViT's accuracy dropped dramatically. The community faced a dilemma: transformers seemed fundamentally data-hungry, and without massive private datasets and GPU clusters, they were impractical for most researchers and practitioners.

The contribution

DeiT demonstrated that a Vision can be trained competitively on ImageNet-1k alone — on a single 8-GPU machine in under 3 days — by combining careful (RandAugment, Mixup, CutMix), (, repeated augmentation), and a novel strategy. The key innovation is a : a learnable appended alongside the class that learns to reproduce the teacher's predictions through , rather than through a separate loss head. With a RegNet-based CNN teacher, DeiT-B⚗ achieves 85.2% top-1 on ImageNet, surpassing ViT-B pre-trained on JFT-300M (84.15%).

The impact

DeiT democratized vision transformers. It proved that transformers do not inherently need hundreds of millions of images — the right training recipe is enough. Its distillation token became a template for injecting CNN inductive biases into transformers. The paper directly enabled the wave of efficient ViT variants (Swin, PVT, CaiT, DeiT-III) and made vision transformers accessible to any lab with a single GPU node, transforming them from a big-tech luxury into a community-wide research tool.

Imagine two students taking the same exam. Student A has access to a private library of 300 million books and a team of personal tutors. Student B has only the public textbook — but an exceptionally smart study plan: flashcards from every angle (data augmentation), random chapter skipping to build resilience (stochastic depth), and a wise upperclassman (a CNN teacher) who slips hints through a special seat in every classroom. That seat is the distillation token — it sits alongside the regular class token, absorbs knowledge from the teacher through self-attention, and feeds it back to the student.

DeiT proved that Student B can match — or even beat — Student A.

The problem: ViT is data-hungry

The original Vision Transformer (ViT) by Dosovitskiy et al. splits an image into 16×16 patches, projects each patch into a , and processes the sequence through a standard Transformer encoder — the same architecture used in NLP. This is elegant but comes with a cost: transformers lack the inductive biases that make CNNs data-efficient.

A CNN's convolutional filters enforce locality (each filter sees a small region), (the same filter scans everywhere), and hierarchical feature extraction (early layers detect edges, later layers detect objects). These biases mean CNNs can learn useful patterns even with modest amounts of data.

A Vision Transformer has none of these built-in shortcuts. Every patch attends to every other patch from layer one — there is no notion of "nearby" or "far away" baked into the architecture. Without these priors, the model must learn spatial structure entirely from data. ViT needed JFT-300M (300 million private images) to match CNN performance. Trained on ImageNet-1k alone (1.28 million images), its accuracy dropped sharply.

Open in Lab
Compare the data requirements: ViT needs 300M private images, while DeiT matches it with ImageNet-1k (1.28M) and the right training recipe.
The demo wakes as you arrive…

Solution part 1: the training recipe

DeiT does not change the ViT architecture at all — DeiT-B is architecturally identical to ViT-B (12 layers, 768 embedding dimensions, 12 heads, 86M parameters). The breakthrough comes entirely from how it is trained.

The authors found that transformers, lacking the inductive biases of CNNs, are far more sensitive to the training procedure. Through systematic ablation, they assembled a recipe combining several data augmentation and regularization techniques — each one partially compensates for the missing inductive biases.

Open in Lab
Toggle each augmentation technique to see how it transforms the input and its impact on accuracy.
The demo wakes as you arrive…

A critical finding from the ablation: removing stochastic depth or repeated augmentation caused the model to fail entirely (accuracy collapsed to ~4%). These are not optional improvements — they are essential for transformer training on limited data. Stochastic depth acts as implicit ensembling, and repeated augmentation ensures each image is seen with enough variety within an epoch to learn robust features despite the smaller dataset.

Solution part 2: distillation through attention

Beyond the training recipe, DeiT introduces a fundamentally new way to distill knowledge into a transformer. Traditional trains a student to mimic the teacher's output distribution by minimizing the between their outputs. This works, but it treats the student as a generic function approximator — the distillation signal does not interact with the student's internal representations.

DeiT's insight is that transformers offer something convnets do not: an attention-based mechanism that can naturally integrate the distillation signal into the model's internal processing. The authors add a single new learnable token — the distillation token — that sits alongside the class token and the patch tokens. Through self-attention, it interacts with every other token at every layer.

Open in Lab
The distillation token sits alongside the class token and patch tokens, learning from the teacher through self-attention at every layer.
The demo wakes as you arrive…

The architecture works as follows: an input image of 224×224 pixels is split into 196 patches of 16×16. Each patch is linearly projected into a 768-dimensional embedding. Two special tokens are prepended: the class token (whose output predicts the true label) and the distillation token (whose output reproduces the teacher's prediction). All 198 tokens (196 patches + 2 special tokens) pass through 12 transformer layers. At the end, the class token's embedding is fed to a head trained with the true label, and the distillation token's embedding is fed to a separate head trained with the teacher's label.

Hard vs soft distillation

The paper compares two distillation approaches:

Soft distillation minimizes the KL divergence between the student's and teacher's softmax distributions at a temperature τ. The total loss is a weighted combination of the standard cross-entropy with true labels and the KL divergence with the teacher's soft output.

Hard distillation is simpler: the teacher's hard prediction (argmax) is treated as a second ground truth. The loss becomes the average of the cross-entropy with the true label and the cross-entropy with the teacher's hard label.

Lsoft=(1−λ) LCE ⁣(ψ(Zs), y)  +  λ τ2 KL ⁣(ψ(Zs/τ), ψ(Zt/τ))\mathcal{L}_{\text{soft}} = (1-\lambda)\,\mathcal{L}_{\text{CE}}\!\bigl(\psi(Z_s),\, y\bigr) \;+\; \lambda\,\tau^2\,\text{KL}\!\bigl(\psi(Z_s/\tau),\, \psi(Z_t/\tau)\bigr)
Soft distillation loss — Zₛ and Zₜ are student and teacher logits; ψ is softmax; τ is temperature; λ balances cross-entropy and KL terms. Higher τ produces softer probability distributions, revealing more inter-class structure from the teacher.
Lhard=12 LCE ⁣(ψ(Zs), y)  +  12 LCE ⁣(ψ(Zs), yt)\mathcal{L}_{\text{hard}} = \frac{1}{2}\,\mathcal{L}_{\text{CE}}\!\bigl(\psi(Z_s),\, y\bigr) \;+\; \frac{1}{2}\,\mathcal{L}_{\text{CE}}\!\bigl(\psi(Z_s),\, y_t\bigr)
Hard distillation loss — yₜ = argmax Zₜ(c) is the teacher's hard prediction. This is parameter-free and conceptually simpler — the teacher's prediction is treated exactly like a ground-truth label. Surprisingly, hard distillation outperforms soft distillation for transformers.
Open in Lab
Compare hard vs soft distillation: slide the temperature to see how soft labels distribute probability mass across classes.
The demo wakes as you arrive…

Why a CNN teacher beats a transformer teacher

A surprising finding: a CNN teacher (RegNet) produces a better student than a transformer teacher of comparable accuracy. DeiT-B distilled from a RegNetY-16GF (82.9% teacher accuracy) reaches 83.4% — surpassing its own teacher and far exceeding what a transformer teacher of similar accuracy can achieve.

This happens because distillation transfers inductive biases, not just label information. The CNN teacher's predictions encode locality and translation invariance — exactly the biases that transformers lack. Through the distillation token, these biases flow into the transformer student via self-attention, compensating for the architectural gap without modifying the architecture itself.

The evidence is clear in the disagreement analysis: the distillation token's predictions are more correlated with the CNN teacher than the class token's predictions, while the class token remains more similar to a transformer trained without distillation. The two tokens learn complementary perspectives — combining them via late fusion gives the best results.

Open in Lab
CNN teachers outperform transformer teachers of similar accuracy. The student surpasses its own teacher.
The demo wakes as you arrive…

Token analysis: class vs distillation

A fascinating property emerges when examining the learned tokens. The class token and the distillation token start very different (cosine similarity ≈ 0.06 at initialization) and gradually converge through the layers, reaching cosine similarity 0.93 at the final layer — high, but still not identical. Each token captures partially different information: the class token focuses on the true label, while the distillation token focuses on the teacher's perspective.

The authors verified that this complementarity is real, not an artifact: when they replaced the distillation token with a second class token (trained on the same true label), the two class tokens converged to identical representations (cosine similarity 0.999) and provided no improvement. The complementarity comes specifically from the different training targets.

At test time, the final prediction is made by adding the softmax outputs from both tokens' classification heads — a simple late fusion that captures the best of both perspectives.

Open in Lab
Watch how class and distillation tokens start far apart and converge layer by layer, while remaining complementary at the output.
The demo wakes as you arrive…

The DeiT model family

DeiT comes in three sizes, all using 12 layers with a per-head dimension of 64:

DeiT-Ti (Tiny): 192 embedding dimensions, 3 heads, 5M parameters, 2536 images/second — the counterpart of ResNet-18. DeiT-S (Small): 384 embedding dimensions, 6 heads, 22M parameters, 940 images/second — the counterpart of ResNet-50. DeiT-B (Base): 768 embedding dimensions, 12 heads, 86M parameters, 292 images/second — identical to ViT-B.

With distillation (denoted by ⚗), DeiT-B⚗ at resolution 384 reaches 85.2% top-1 accuracy after 1000 epochs of training, surpassing ViT-B trained on JFT-300M (84.15%) despite using 234× fewer training images.

Open in Lab
Explore the DeiT model family: compare sizes, throughput, and accuracy with and without distillation.
The demo wakes as you arrive…

Transfer learning results

DeiT models transfer competitively to downstream tasks. When fine-tuned on CIFAR-10, DeiT-B⚗ reaches 99.1% accuracy — comparable to EfficientNet-B7 (98.9%). On fine-grained benchmarks like Stanford Cars (92.9%) and Oxford Flowers (98.8%), DeiT matches or exceeds convolutional state of the art.

The authors also apply resolution : models are pre-trained at 224×224, then fine-tuned at 384×384 by interpolating positional embeddings with bicubic interpolation. This preserves the ℓ₂-norm of the embeddings (unlike bilinear interpolation, which shrinks them), avoiding a sharp accuracy drop before fine-tuning. This resolution change increases the number of patches from 196 to 576, giving the model a more detailed view of each image.

DeiT in code

Distillation token: the key architectural additionpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

class DeiT(nn.Module):
    """Vision Transformer with a distillation token."""
    def __init__(self, img_size=224, patch_size=16, embed_dim=768,
                 depth=12, num_heads=12, num_classes=1000):
        super().__init__()
        num_patches = (img_size // patch_size) ** 2  # 196
        # Patch embedding: project 16x16x3 → 768
        self.patch_embed = nn.Conv2d(3, embed_dim,
            kernel_size=patch_size, stride=patch_size)
        # Two special tokens
        self.cls_token = nn.Parameter(torch.zeros(1, 1, embed_dim))
        self.dist_token = nn.Parameter(torch.zeros(1, 1, embed_dim))
        # Positional embeddings for 196 patches + 2 tokens
        self.pos_embed = nn.Parameter(
            torch.zeros(1, num_patches + 2, embed_dim))
        # Transformer blocks
        self.blocks = nn.ModuleList([
            TransformerBlock(embed_dim, num_heads) for _ in range(depth)
        ])
        self.norm = nn.LayerNorm(embed_dim)
        # Two classification heads
        self.head = nn.Linear(embed_dim, num_classes)       # true label
        self.head_dist = nn.Linear(embed_dim, num_classes)  # teacher label

    def forward(self, x):
        B = x.shape[0]
        # Patchify: (B,3,224,224) → (B,196,768)
        x = self.patch_embed(x).flatten(2).transpose(1, 2)
        # Prepend cls + distillation tokens → (B,198,768)
        cls = self.cls_token.expand(B, -1, -1)
        dist = self.dist_token.expand(B, -1, -1)
        x = torch.cat([cls, dist, x], dim=1)
        x = x + self.pos_embed
        # Pass through transformer layers
        for block in self.blocks:
            x = block(x)
        x = self.norm(x)
        # Separate outputs
        cls_out = self.head(x[:, 0])       # class token
        dist_out = self.head_dist(x[:, 1]) # distillation token
        # Late fusion at test time
        return (cls_out + dist_out) / 2

What DeiT unlocked

  1. 2020

    ViT: Vision Transformer

    Dosovitskiy et al. show transformers can classify images, but require JFT-300M (300M private images) and massive compute to match CNN performance.

  2. 2021

    DeiT: Data-efficient training + distillation token

    Trains ViT competitively on ImageNet-1k alone using data augmentation, regularization, and a novel distillation token. 85.2% top-1 accuracy with 234× fewer images than ViT.

  3. 2021

    Swin Transformer: hierarchical + shifted windows

    Introduces hierarchical feature maps and shifted-window self-attention, combining CNN-like multiscale with transformer power. Becomes the dominant vision backbone.

  4. 2021

    CaiT: Going deeper with transformers

    Touvron et al. build on DeiT with class-attention layers, enabling deeper vision transformers with better stability.

  5. 2022

    DeiT-III: Revenge of the ViT

    Further refined training strategies push vanilla ViT to new heights without any architectural changes — validating DeiT's core thesis that training recipe matters more than architecture.

DeiT's deepest contribution is a shift in mindset. Before DeiT, the community assumed that vision transformers needed architectural innovations (local attention, hierarchical stages, hybrids) to work on limited data. DeiT showed that the vanilla ViT architecture is sufficient — the was the training procedure, not the model. This insight liberated an entire generation of researchers: instead of redesigning the transformer, they could focus on how to train it.

CitationTouvron, Cord, Douze, Massa, Sablayrolles, Jégou. Training Data-Efficient Image Transformers & Distillation Through Attention. ICML, 2021.

Terms in this paper