Computer Vision2021intermediate14 min read
Emerging Properties in Self-Supervised Vision Transformers
خصائص ناشئة في محوِّلات الرؤية ذاتية الإشراف
Caron, M. · Touvron, H. · Misra, I. · Jégou, H. · Mairal, J. · Bojanowski, P. · Joulin, A. — ICCV
The problem
By 2021, Vision Transformers (ViTs) had shown strong results on image , but only when trained with supervision on massive labeled datasets like ImageNet-21k or JFT-300M. Meanwhile, the secret weapon of Transformers in NLP — self-supervised pre- — had not been fully explored for vision. Existing self-supervised methods (SimCLR, MoCo, BYOL) were designed around convolutional networks and relied on contrastive losses with negative samples, large memory banks, or specialized predictors. Nobody had asked the key question: does combining with ViTs unlock properties that neither supervised ViTs nor self-supervised convnets can achieve?
The contribution
DINO ( with NO labels): a simple self-supervised framework where a student network learns to match the output distribution of a teacher network — and the teacher is nothing but an of the student itself. Two global crops go through the teacher; the student processes both global and local crops, learning local-to-global correspondences via . and sharpening prevent without needing negative samples, contrastive losses, or memory banks. The key discovery: self-supervised ViTs trained with DINO exhibit emergent properties — their maps explicitly encode boundaries, and their features serve as excellent k-NN classifiers (78.3% top-1 on ImageNet with -S) without any .
The impact
DINO proved that self-supervised ViTs produce qualitatively different representations from supervised ones — representations that "see" object boundaries without ever being told where objects are. It became the foundation for DINOv2, which scales the approach with curated data and combined objectives to produce general-purpose visual features rivaling CLIP. DINO's self-distillation paradigm influenced a wave of methods (iBOT, MSN, EsViT) and its maps became the standard demonstration that Transformers can learn spatial understanding from images alone.
Imagine a dance class with a peculiar setup: the teacher and the student are the same person at different moments in time. The student dances today; the teacher is a smoothed memory of every dance the student has ever performed. The student watches the full stage through a panoramic window (global crops), but also practices through small peepholes (local crops) — seeing only a hand or a foot at a time. The rule: whatever the student glimpses through a peephole must match what the teacher sees through the full window. Over weeks of practice, the student internalizes the whole choreography from fragments alone — and something magical happens: the student's gaze naturally locks onto the dancer and ignores the background, even though no one ever said "look at the dancer, not the curtains."
The problem: supervised ViTs miss hidden structure
When Vision Transformers are trained with labeled data — say, 1000 ImageNet classes — they learn to classify correctly, but their internal attention maps look noisy and scattered. Ask a supervised ViT "what are you looking at?" and you'll get a messy heat map that activates everywhere, not clean object boundaries.
Meanwhile, in NLP, Transformers thrived precisely because of self-supervised pre-training (BERT's masked tokens, GPT's next-word prediction). The authors of DINO asked: what happens when we bring self-supervised training to ViTs? Not just "does it work?" but "does it reveal something new?"
The answer surprised everyone. Self-supervised ViTs developed emergent properties — abilities that nobody explicitly trained them for. Their attention heads learned to segment objects from backgrounds. Their feature spaces organized so cleanly that a simple lookup, with no training at all, could classify images nearly as well as a trained linear classifier.
The DINO framework: self-distillation with no labels
DINO stands for self-DIstillation with NO labels. The name captures the entire idea: take — a technique where a student network learns from a teacher network — and remove the two things you'd think are essential: (1) the labels, and (2) the separately-trained teacher.
In classical knowledge distillation, the teacher is a large, pre-trained model that provides soft probability targets for a smaller student. DINO flips this: the teacher and student share the same architecture, and the teacher is simply an exponential moving average of the student's own weights. The student updates by ; the teacher updates by slowly absorbing the student's parameters. There are no labels anywhere in the pipeline.
Why does this work at all? Because the teacher, being a momentum-smoothed version of the student, provides a more stable and consistent target than the student's own rapidly changing output. It's like learning to draw by comparing your sketch not to a photograph, but to a running average of your best past sketches — a smoothed version of yourself that's always slightly better because it's less noisy.
Multi-crop: teaching local-to-global understanding
The strategy in DINO creates an asymmetry between what the teacher and student see. From each training image, DINO generates:
2 global crops at 224×224 resolution, each covering more than 50% of the image. These are the teacher's inputs — it always sees the big picture.
Several local crops (typically 6–8) at 96×96 resolution, each covering only 5–50% of the image. These small patches might show just a wing, a wheel, or a patch of fur.
The student processes all crops (global + local). The teacher processes only the global crops. The loss asks: can the student, looking through a peephole at a small patch, produce the same output as the teacher looking at the whole scene?
This local-to-global correspondence is the engine of DINO's learning. The student must infer global semantics from local details — "this texture and shape could only belong to a bird, so my output should match the teacher's bird-like distribution." Over training, the model develops a hierarchy of features from textures to parts to objects, without any labels.
The loss: matching probability distributions
Both networks produce a K-dimensional vector (where K = 65,536 prototypes by default). This vector is passed through a with to produce a probability distribution. The student's temperature τ_s is relatively high (e.g. 0.1) producing softer distributions, while the teacher's temperature τ_t is lower (e.g. 0.04) producing sharper, more confident distributions. This sharpening is critical — it tells the student "the teacher is very sure; pay attention to where it puts its probability mass."
The loss is the cross-entropy between the teacher's output (the target) and the student's output (the prediction), summed over all cross-view pairs. Crucially, self-pairs are excluded: if the teacher sees global crop 1, the student's loss is computed on global crop 2 and all local crops — never on global crop 1 itself. This prevents trivial identity solutions.
Avoiding collapse: centering and sharpening
Without labels or negative samples, what stops the network from collapsing to a trivial solution — say, outputting the same constant vector for every image? This is the representation collapse problem, the bane of self-.
Contrastive methods (SimCLR, MoCo) solve this with explicit negative pairs: "make this image's representation different from that image's." BYOL avoids negatives using a predictor head and Stop . DINO takes a third path with two complementary mechanisms:
Centering subtracts a running mean of the teacher's outputs across the batch. This prevents any single dimension from dominating — without centering, one dimension could grow huge while all others shrink to zero (single-dimension collapse). But centering alone would push toward a uniform distribution where every gets equal probability.
Sharpening counteracts this by using a low temperature τ_t in the teacher's softmax. Low temperature makes the distribution peaky — the teacher is forced to commit to a few prototypes, not spread probability everywhere.
Together, centering prevents dimension dominance while sharpening prevents uniform collapse. They are opposing forces in perfect balance, like two walls holding up an arch: remove either one and the structure falls.
The momentum teacher: a smoother self
The teacher network never receives gradients. Instead, after each student update, the teacher's weights are nudged toward the student's via exponential moving average. The momentum parameter λ starts at 0.996 and is increased to 1.0 during training following a cosine schedule.
Why does a slower copy of the student make a better teacher? Because gradient descent is noisy — each batch pushes the student in a slightly different direction. The EMA teacher absorbs these noisy updates gradually, producing a smoother, more stable version of the student. In early training, λ is lower (0.996) so the teacher adapts faster; as training stabilizes, λ approaches 1.0 so the teacher changes very slowly, providing an increasingly consistent target.
This is the same principle as Polyak averaging in optimization: the average of many noisy iterates is closer to the optimum than any single iterate.
Architecture: ViT backbone + projection head
DINO works with both convolutional networks (ResNet-50) and Vision Transformers, but the emergent properties only appear with ViTs. The architecture has two parts:
The is a standard ViT. An image is split into non-overlapping patches (16×16 or 8×8 pixels), each patch is linearly embedded, a is prepended, positional embeddings are added, and the sequence passes through layers. The [CLS] token's output becomes the image-level representation.
The sits on top of the backbone. It takes the [CLS] output and maps it to the K-dimensional space where centering and softmax are applied. The projection head is a 3-layer MLP with GELU activations, followed by L2 normalization and a weight-normalized fully connected layer. This head is used during pre-training but is typically discarded for downstream tasks.
The paper experiments with ViT-Small (ViT-S/16 with 21M parameters) and ViT-Base (ViT-B/16 with 86M parameters). A key finding: using smaller patches (8×8 instead of 16×16) significantly improves feature quality, because the model can attend to finer spatial details.
Emergent properties: what DINO sees without being told
The paper's most striking result is not a benchmark number — it's a qualitative discovery. When you visualize the self-attention maps of the last layer's [CLS] token, DINO-trained ViTs produce clean, object-aware segmentation masks. Different attention heads specialize: one head might attend to the foreground object, another to background regions, another to fine-grained parts.
This emergent segmentation was not trained for. The model never saw a single segmentation mask. Yet it learned to separate objects from backgrounds because that's what best serves the self-distillation objective: to match the teacher's view from a local crop, the student needs to understand what the object is regardless of where it appears in the frame.
Quantitatively, the features also shine. A simple k-nearest neighbors classifier on frozen DINO features achieves 78.3% top-1 on ImageNet with ViT-S — no fine-tuning, no trained head, just look up the nearest neighbors in the training set. With (a single trainable linear layer on frozen features), ViT-B reaches 80.1% top-1. These results demonstrate that DINO features organize into a well-structured metric space where semantic similarity maps cleanly to geometric distance.
Putting it together: the training loop
Simplified to show the idea — not the real implementation.
# DINO pseudocode — PyTorch-like
# student and teacher share the same architecture
student = ViT()
teacher = copy(student) # same weights initially
teacher.requires_grad_(False) # no gradients for teacher
for x in dataloader:
# Generate multi-crop views
global_views = [augment_global(x), augment_global(x)]
local_views = [augment_local(x) for _ in range(n_local)]
all_views = global_views + local_views
# Teacher processes only global views
t_out = [teacher(v) for v in global_views]
# Student processes all views
s_out = [student(v) for v in all_views]
# Cross-entropy loss over cross-view pairs
loss = 0
for t_view in t_out:
t_probs = softmax((t_view - center) / tau_t) # centered + sharpened
for s_view in s_out:
if s_view is not t_view: # exclude self-pairs
s_probs = softmax(s_view / tau_s)
loss += cross_entropy(t_probs, s_probs)
loss.backward()
optimizer.step() # update student only
# EMA update for teacher
with torch.no_grad():
for tp, sp in zip(teacher.parameters(), student.parameters()):
tp.data = lambda_ * tp.data + (1 - lambda_) * sp.data
# Update center
center = m * center + (1 - m) * mean(t_out)What matters most: ablation insights
The paper's reveals which components are essential and which are helpful but optional:
Remove centering → training collapses. The output becomes uniform — every image produces the same probability distribution. This is the uniform-collapse mode.
Remove sharpening (high teacher temperature) → training collapses to single-dimension dominance. One prototype captures all the probability mass; the rest are dead.
Remove multi-crop → accuracy drops 2–3%. The model still works but learns less about local-to-global relationships.
Replace EMA with copy (no momentum) → the teacher changes too fast and training becomes unstable or collapses.
Add a predictor (like BYOL) → no improvement. Centering and sharpening already provide sufficient asymmetry.
The lesson: DINO's simplicity is not accidental. Each component plays a specific, non-redundant role, and removing any one of the collapse-prevention mechanisms (centering, sharpening, or EMA) causes failure.
Key results
DINO achieves competitive results across multiple evaluation protocols:
k-NN classification on ImageNet (no training, just neighbor lookup): 78.3% top-1 with ViT-S/16, 77.4% with ViT-B/16. These numbers are remarkable — they mean the feature space is so well-organized that simple distance works as a classifier.
Linear probing (single frozen-feature linear layer): 77.0% with ViT-S/16, 80.1% with ViT-B/16. The ViT-B result matches or exceeds most supervised baselines of its era.
Semantic segmentation (attention-based, no training): the self-attention maps of the [CLS] token in the last layer produce segmentation masks that correlate strongly with ground-truth object boundaries — a property unique to DINO-trained ViTs.
Smaller patches help: moving from 16×16 to 8×8 patches (ViT-S/8) raises k-NN accuracy to 78.3% and dramatically improves the spatial resolution of attention maps, enabling finer segmentation.
Context and legacy
2020
BYOL — Bootstrap Your Own Latent
Showed self-supervised learning works without negative samples using a momentum encoder and a predictor. DINO builds directly on this insight.
2020
ViT — Vision Transformer
Demonstrated that pure Transformer architectures can match CNNs on image classification, but required massive supervised pre-training.
2021
DINO — Self-Distillation with No Labels
Combined self-supervised learning with ViTs, discovering emergent segmentation in attention maps and achieving strong k-NN classification without any labels.
2021
MAE — Masked Autoencoders
An alternative self-supervised approach for ViTs: mask 75% of patches and reconstruct them. Different philosophy from DINO — generative vs. discriminative.
2023
DINOv2 — Scaling DINO
Combined DINO with iBOT's masked-patch objective, trained on curated LVD-142M dataset. Produced general-purpose visual features rivaling CLIP without text.
DINO's legacy is twofold. First, it showed that the combination of self-supervision and ViTs is more than the sum of its parts: it unlocks qualitatively new behaviors (emergent segmentation) that neither ingredient produces alone. Second, it provided a remarkably simple recipe — no contrastive loss, no memory bank, no predictor — that scales cleanly to larger models and datasets, as DINOv2 later demonstrated.
CitationCaron, Touvron, Misra, Jégou, Mairal, Bojanowski, Joulin. Emerging Properties in Self-Supervised Vision Transformers. ICCV, 2021.
Terms in this paper
- Self-Distillationالتقطير الذاتي
- Self-Supervised Learningالتعلم ذاتي الإشراف
- Vision Transformer (ViT)محوِّل الرؤية (ViT)
- Knowledge Distillationتقطير المعرفة
- Momentum Encoderمُرمِّز الزخم
- Exponential Moving Average (EMA)المتوسط المتحرك الأُسِّي
- Multi-Crop Augmentationتعزيز الاقتصاص المتعدد
- Centeringالتوسيط
- Teacher Modelنموذج المعلم
- Student Modelنموذج الطالب
- Projection Headرأس الإسقاط
- k-Nearest Neighborsخوارزمية الجيران الأقرب (KNN)
- Semantic Segmentationالتجزئة الدلالية للصورة
- Linear Probingالاختبار الخطي
- Representation Collapseانهيار التمثيلات