Computer Vision2023intermediate11 min read

DINOv2: Learning Robust Visual Features Without Supervision

DINOv2: تعلُّم سمات بصرية متينة بدون إشراف

Oquab, M. · Darcet, T. · Moutakanni, T. · Vo, H. · Szafraniec, M. · Khalidov, V. · Fernandez, P. · Haziza, D. · Massa, F. · El-Nouby, A. · Assran, M. · Ballas, N. · Galuba, W. · Howes, R. · Huang, P.-Y. · Li, S.-W. · Misra, I. · Rabbat, M. · Sharma, V. · Synnaeve, G. · Xu, H. · Jegou, H. · Mairal, J. · Labatut, P. · Joulin, A. · Bojanowski, P. — TMLR

The problem

By 2023, NLP had already achieved the dream of general-purpose pretrained models — train once on massive data, then use the features everywhere without . Computer vision lagged behind. Self-supervised methods like DINO and MAE showed promise, but they were trained on small curated datasets like ImageNet-1k. Scaling them to uncurated web data led to poor feature quality. Text-guided models like CLIP learned strong features, but they depended on caption data and missed fine-grained pixel-level information. There was no purely visual, label-free model whose frozen features could rival supervised and text-guided models across image-level and pixel-level tasks.

The contribution

DINOv2: a family of Vision Transformers (-S/B/L/g) pretrained entirely with self-supervision on a curated dataset of 142M images (LVD-142M), built by an automatic pipeline that retrieves diverse images from uncurated web data using visual similarity to curated sources. The training combines image-level self-distillation (DINO loss), patch-level (iBOT loss), and a KoLeo regularizer — with engineering innovations (FlashAttention, sequence packing, FSDP, efficient ) that make training 2× faster and 3× more memory-efficient. Smaller models are distilled from a 1B-parameter ViT-g teacher. The resulting frozen features match or surpass the best available weakly-supervised features (OpenCLIP-G) across classification, segmentation, depth estimation, and retrieval — without any fine-tuning.

The impact

DINOv2 proved that can produce truly general-purpose visual features — the computer vision equivalent of GPT-style pretrained representations. Its frozen features became the default for Depth Anything and numerous downstream applications in segmentation, retrieval, and 3D understanding. It demonstrated that matters more than data scale for self-supervised pretraining, and that distillation from large teachers is more effective than training smaller models from scratch. DINOv2 shifted the field's expectation: a good visual backbone should not require fine-tuning.

Imagine an art student who visits every museum in the world — not with a teacher pointing out masterpieces, but alone, studying everything independently. After years of self-guided learning, you hand them any image — a medical scan, a satellite photo, a painting — and they can instantly describe its structure, segment its regions, and estimate the depth of every surface. No further training needed.

DINOv2 is that student. Trained entirely without human labels on 142 million carefully curated images, it learns visual features so rich and universal that they work "out of the box" for almost any vision task — from classifying species to estimating the depth of a room.

The gap: why vision lacked a foundation model

In NLP, the recipe was clear by 2023: pretrain a large model on massive text using self-supervision, then use its features everywhere — no fine-tuning needed. GPT and BERT proved that frozen features from a single model could power dozens of tasks.

Computer vision had no equivalent. Self-supervised methods like DINO and MAE showed promising results, but they were confined to small curated datasets like ImageNet-1k with only 1.3M images. When researchers tried scaling these methods to billions of uncurated web images, the feature quality dropped significantly — the noise and redundancy in raw web data hurt more than the extra scale helped.

Text-guided models like CLIP took a different path: they used image-caption pairs to learn visual features aligned with language. This produced strong image-level features, but captions are a lossy description of images — they miss fine-grained spatial information needed for tasks like and depth estimation.

DINOv2 asked a fundamental question: can we build a visual using only images — no labels, no captions — if we invest in curating the right data and scaling the right training recipe?

Open in Lab
Compare three paradigms: supervised (needs labels), text-guided (needs captions), and self-supervised (images only). Toggle to see what each approach requires and what it can do.
The demo wakes as you arrive…

Curating the data: quality over quantity

The key insight of DINOv2 is that curated data matters more than raw scale for self-supervised pretraining. The authors built an automatic data pipeline — requiring no human labels or metadata — to assemble a diverse, deduplicated, and balanced dataset called LVD-142M.

Think of it as building a library. A library with a billion random books is less useful than one with 142 million carefully selected books that cover every subject without duplication. The pipeline works in three stages:

Stage 1 — Seeding. Start with existing curated datasets (ImageNet-22k, Google Landmarks, fine-grained datasets) as "anchor" images that define what good, diverse data looks like.

Stage 2 — Retrieval. From a pool of 1.2 billion uncurated web images, retrieve images that are visually similar to the curated anchors using in a self-supervised space. For each anchor, the 4 nearest neighbors are retrieved.

Stage 3 — Deduplication and rebalancing. Remove near-duplicates using copy detection, and cluster the data to ensure no single visual concept dominates the dataset.

The result: 142 million images that are diverse, balanced, and free of redundancy — assembled in under two days on 20 GPU nodes.

Open in Lab
Explore the three-stage data curation pipeline. Click each stage to see how uncurated web images are filtered into a high-quality training set.
The demo wakes as you arrive…

The training recipe: three losses, one model

DINOv2 combines three complementary training objectives. Think of them as three different "exercises" that force the model to develop both global understanding (what is in the image) and local understanding (where everything is, at pixel level).

The architecture follows a student-teacher setup. The student network sees cropped and masked versions of an image; the teacher network sees the full, unmasked image. The teacher is not a separate model — it is an of the student's own past weights. This creates a stable learning target that evolves gradually.

Loss 1: DINO (image-level). Both the student and teacher extract a — a single vector summarizing the entire image. The student must match the teacher's summary. This forces the model to learn high-level semantic understanding: "this is a dog", "this is a forest." The loss is a cross-entropy between the student and teacher prototype scores:

LDINO=−∑ptlog⁡ps\mathcal{L}_{\text{DINO}} = -\sum p_t \log p_s
DINO loss — image-level self-distillation — The teacher produces soft probability targets ptp_t over learned prototypes; the student produces its own predictions psp_s. The cross-entropy pushes the student to match the teacher's view of the image, despite seeing different crops. The teacher is an exponential moving average (EMA) of the student.

Loss 2: iBOT (patch-level). Random patches are masked from the student's input. The student must predict the teacher's representation of those hidden patches — like filling in missing pieces of a jigsaw puzzle. This forces the model to learn fine-grained spatial features needed for tasks like segmentation and depth estimation:

LiBOT=−∑i∈maskedptilog⁡psi\mathcal{L}_{\text{iBOT}} = -\sum_{i \in \text{masked}} p_{t_i} \log p_{s_i}
iBOT loss — patch-level masked image modeling — For each masked patch ii, the teacher provides soft targets ptip_{t_i} from the visible version of that patch. The student, seeing only a mask token, must reconstruct the teacher's representation. This trains pixel-level understanding — the model must infer local structure from context.

Loss 3: KoLeo regularizer. This prevents the feature space from collapsing — a common failure mode in self-supervised learning where all representations converge to a small region. KoLeo encourages features to spread uniformly across the embedding space, like ensuring books in a library occupy every shelf rather than piling up in one corner:

LKoLeo=−1n∑i=1nlog⁡(min⁡j≠i∥xi−xj∥)\mathcal{L}_{\text{KoLeo}} = -\frac{1}{n}\sum_{i=1}^{n}\log\left(\min_{j \neq i}\|x_i - x_j\|\right)
KoLeo regularizer — prevents representation collapse — For each feature vector xix_i in a batch, we compute its distance to the nearest neighbor xjx_j. Minimizing the negative log of these distances pushes points apart, ensuring the feature space is well-utilized. Without this, self-supervised models risk collapsing all representations to a small cluster.
Open in Lab
Explore how the three losses work together. Toggle each loss to see what it teaches the model and which tasks benefit from it.
The demo wakes as you arrive…

Engineering at scale: making it feasible

Training a 1 billion parameter Vision Transformer on 142 million images requires solving several engineering challenges. The DINOv2 team contributed four key innovations that together made the training 2× faster and 3× more memory-efficient compared to iBOT:

FlashAttention. A custom implementation of memory-efficient that avoids materializing the full attention matrix. This is critical because 's memory scales quadratically with sequence length.

Sequence packing. The DINO algorithm uses both large crops (224×224) and small crops (98×98). Instead of forwarding them separately (wasting compute on padding), they are concatenated into one long sequence with a block-diagonal attention mask — preventing attention between different crops while processing them together.

Efficient stochastic depth. Instead of masking dropped residual blocks (computing them and then throwing away the result), the skipped blocks are never computed at all. With a 40% drop rate, this saves substantial compute and memory.

FSDP (Fully-Sharded Data Parallel). The 1B parameter model with its optimizer states requires 16 GB of memory. FSDP shards this across GPUs, with mixed-precision communication that cuts bandwidth by 50% compared to standard distributed training.

Open in Lab
See how each engineering optimization contributes to the 2× speedup and 3× memory reduction. Click each technique to learn its impact.
The demo wakes as you arrive…

Distillation: one large teacher, many smaller students

Training the ViT-g model (1.1B parameters) from scratch is extremely expensive. For smaller models — ViT-S, ViT-B, ViT-L — DINOv2 uses instead of training from scratch.

The idea is simple: use the already-trained ViT-g as a frozen teacher and train smaller student models to reproduce its outputs. The same student-teacher framework used during pretraining is reused for distillation, with a few changes: the teacher is now frozen (the large ViT-g), masking and stochastic depth are removed for the student, and the final model is the of the student.

The surprising result: distilled models outperform models trained from scratch on every single benchmark. A distilled ViT-L nearly matches the ViT-g teacher, and sometimes even surpasses it. This makes distillation not just cheaper, but better.

Open in Lab
Watch knowledge flow from the frozen ViT-g teacher to smaller student models. Compare the performance of distilled vs. scratch-trained models.
The demo wakes as you arrive…

Results: features that work everywhere

The defining property of DINOv2 is that its frozen features — extracted without any task-specific training — perform competitively or better than the best available features across a remarkable range of tasks:

Image classification. On ImageNet-1k linear evaluation, DINOv2 ViT-g achieves 86.5% top-1 accuracy — surpassing OpenCLIP-G (86.2%) and EVA-CLIP (86.4%), both of which used text supervision. On robustness benchmarks like ImageNet-A, DINOv2 reaches 75.9% vs 63.8% for OpenCLIP-G.

Semantic segmentation. With just a linear layer on top of frozen features, DINOv2 reaches 53.0 mIoU on ADE20k — matching MAE with full fine-tuning and a complex UperNet decoder (53.6). With a full decoder, it reaches 60.2 mIoU.

Depth estimation. DINOv2 features, with a simple DPT head, match or exceed dedicated depth estimation models — while being a completely general-purpose backbone. This capability directly enabled Depth Anything.

Instance retrieval. On Oxford-Hard, DINOv2 achieves 52.3 mAP vs 19.7 for OpenCLIP-G — nearly 3× better. This shows the features capture fine-grained identity, not just category-level similarity.

The fact that a single self-supervised model with frozen features can match or beat specialized models across all these tasks is the central achievement of DINOv2.

Open in Lab
Compare DINOv2 against OpenCLIP-G and iBOT across classification, segmentation, depth, and retrieval. Click a method to highlight its profile.
The demo wakes as you arrive…

Seeing what the model sees: PCA of patch features

One of the most striking demonstrations of DINOv2's quality is visualizing its patch features using . When you take the patch tokens from the last layer and project them to their first three principal components, mapping each to a color channel (RGB), something remarkable emerges: the model has learned a consistent semantic mapping across images.

The same body part of a dog — say, the head — maps to the same color across different dog breeds, poses, and even between a real dog and a painting of one. Background pixels cluster separately from foreground. This shows that the model has learned, without any labels, to segment objects and match semantic parts across dramatically different images.

This is not just a pretty visualization — it proves the features encode rich spatial and semantic structure that traditional supervised features often lack.

Open in Lab
See how DINOv2 patch features map semantically similar parts to similar colors across different images. Click image pairs to explore.
The demo wakes as you arrive…

From DINO to DINOv2: the road to visual foundation features

  1. 2020

    BYOL and SimCLR

    Established contrastive and self-distillation frameworks for self-supervised visual learning, showing that augmentation-based objectives can learn strong features from images alone.

  2. 2021

    DINO (Caron et al.)

    Introduced self-distillation with Vision Transformers, showing that ViT features trained with a momentum teacher contain emergent object segmentation in the attention maps — no labels needed.

  3. 2021

    VICReg (Bardes et al.)

    Proposed variance-invariance-covariance regularization to prevent representation collapse without negative pairs or momentum encoders, influencing DINOv2's feature spreading strategy.

  4. 2022

    MAE (He et al.)

    Showed that masked autoencoders learn powerful features by reconstructing masked image patches. Features require fine-tuning but the patch-level objective inspired DINOv2's iBOT loss component.

  5. 2022

    iBOT (Zhou et al.)

    Combined DINO-style self-distillation with masked image modeling at the patch level, creating the direct precursor to DINOv2's training recipe.

  6. 2023

    DINOv2 (this paper)

    Scaled self-supervised pretraining with curated data, engineering optimizations, and distillation to produce the first self-supervised model whose frozen features match weakly-supervised models across image and pixel-level tasks.

  7. 2024

    Depth Anything (Yang et al.)

    Built directly on DINOv2's frozen backbone to create a state-of-the-art monocular depth estimation model, validating DINOv2's promise as a universal visual feature extractor.

CitationOquab, Darcet, Moutakanni, Vo, Szafraniec, Khalidov, Fernandez, Haziza, Massa, El-Nouby, Assran, Ballas, Galuba, Howes, Huang, Li, Misra, Rabbat, Sharma, Synnaeve, Xu, Jegou, Mairal, Labatut, Joulin, Bojanowski. DINOv2: Learning Robust Visual Features Without Supervision. Transactions on Machine Learning Research (TMLR), 2023.

Terms in this paper