Computer Vision2023intermediate11 min read
DINOv2: Learning Robust Visual Features Without Supervision
DINOv2: تعلُّم سمات بصرية متينة بدون إشراف
Oquab, M. · Darcet, T. · Moutakanni, T. · Vo, H. · Szafraniec, M. · Khalidov, V. · Fernandez, P. · Haziza, D. · Massa, F. · El-Nouby, A. · Assran, M. · Ballas, N. · Galuba, W. · Howes, R. · Huang, P.-Y. · Li, S.-W. · Misra, I. · Rabbat, M. · Sharma, V. · Synnaeve, G. · Xu, H. · Jegou, H. · Mairal, J. · Labatut, P. · Joulin, A. · Bojanowski, P. — TMLR
The problem
By 2023, NLP had already achieved the dream of general-purpose pretrained models — train once on massive data, then use the features everywhere without . Computer vision lagged behind. Self-supervised methods like DINO and MAE showed promise, but they were trained on small curated datasets like ImageNet-1k. Scaling them to uncurated web data led to poor feature quality. Text-guided models like CLIP learned strong features, but they depended on caption data and missed fine-grained pixel-level information. There was no purely visual, label-free model whose frozen features could rival supervised and text-guided models across image-level and pixel-level tasks.
The contribution
DINOv2: a family of Vision Transformers (-S/B/L/g) pretrained entirely with self-supervision on a curated dataset of 142M images (LVD-142M), built by an automatic pipeline that retrieves diverse images from uncurated web data using visual similarity to curated sources. The training combines image-level self-distillation (DINO loss), patch-level (iBOT loss), and a KoLeo regularizer — with engineering innovations (FlashAttention, sequence packing, FSDP, efficient ) that make training 2× faster and 3× more memory-efficient. Smaller models are distilled from a 1B-parameter ViT-g teacher. The resulting frozen features match or surpass the best available weakly-supervised features (OpenCLIP-G) across classification, segmentation, depth estimation, and retrieval — without any fine-tuning.
The impact
DINOv2 proved that can produce truly general-purpose visual features — the computer vision equivalent of GPT-style pretrained representations. Its frozen features became the default for Depth Anything and numerous downstream applications in segmentation, retrieval, and 3D understanding. It demonstrated that matters more than data scale for self-supervised pretraining, and that distillation from large teachers is more effective than training smaller models from scratch. DINOv2 shifted the field's expectation: a good visual backbone should not require fine-tuning.
Imagine an art student who visits every museum in the world — not with a teacher pointing out masterpieces, but alone, studying everything independently. After years of self-guided learning, you hand them any image — a medical scan, a satellite photo, a painting — and they can instantly describe its structure, segment its regions, and estimate the depth of every surface. No further training needed.
DINOv2 is that student. Trained entirely without human labels on 142 million carefully curated images, it learns visual features so rich and universal that they work "out of the box" for almost any vision task — from classifying species to estimating the depth of a room.
The gap: why vision lacked a foundation model
In NLP, the recipe was clear by 2023: pretrain a large model on massive text using self-supervision, then use its features everywhere — no fine-tuning needed. GPT and BERT proved that frozen features from a single model could power dozens of tasks.
Computer vision had no equivalent. Self-supervised methods like DINO and MAE showed promising results, but they were confined to small curated datasets like ImageNet-1k with only 1.3M images. When researchers tried scaling these methods to billions of uncurated web images, the feature quality dropped significantly — the noise and redundancy in raw web data hurt more than the extra scale helped.
Text-guided models like CLIP took a different path: they used image-caption pairs to learn visual features aligned with language. This produced strong image-level features, but captions are a lossy description of images — they miss fine-grained spatial information needed for tasks like and depth estimation.
DINOv2 asked a fundamental question: can we build a visual using only images — no labels, no captions — if we invest in curating the right data and scaling the right training recipe?
Curating the data: quality over quantity
The key insight of DINOv2 is that curated data matters more than raw scale for self-supervised pretraining. The authors built an automatic data pipeline — requiring no human labels or metadata — to assemble a diverse, deduplicated, and balanced dataset called LVD-142M.
Think of it as building a library. A library with a billion random books is less useful than one with 142 million carefully selected books that cover every subject without duplication. The pipeline works in three stages:
Stage 1 — Seeding. Start with existing curated datasets (ImageNet-22k, Google Landmarks, fine-grained datasets) as "anchor" images that define what good, diverse data looks like.
Stage 2 — Retrieval. From a pool of 1.2 billion uncurated web images, retrieve images that are visually similar to the curated anchors using in a self-supervised space. For each anchor, the 4 nearest neighbors are retrieved.
Stage 3 — Deduplication and rebalancing. Remove near-duplicates using copy detection, and cluster the data to ensure no single visual concept dominates the dataset.
The result: 142 million images that are diverse, balanced, and free of redundancy — assembled in under two days on 20 GPU nodes.
The training recipe: three losses, one model
DINOv2 combines three complementary training objectives. Think of them as three different "exercises" that force the model to develop both global understanding (what is in the image) and local understanding (where everything is, at pixel level).
The architecture follows a student-teacher setup. The student network sees cropped and masked versions of an image; the teacher network sees the full, unmasked image. The teacher is not a separate model — it is an of the student's own past weights. This creates a stable learning target that evolves gradually.
Loss 1: DINO (image-level). Both the student and teacher extract a — a single vector summarizing the entire image. The student must match the teacher's summary. This forces the model to learn high-level semantic understanding: "this is a dog", "this is a forest." The loss is a cross-entropy between the student and teacher prototype scores:
Loss 2: iBOT (patch-level). Random patches are masked from the student's input. The student must predict the teacher's representation of those hidden patches — like filling in missing pieces of a jigsaw puzzle. This forces the model to learn fine-grained spatial features needed for tasks like segmentation and depth estimation:
Loss 3: KoLeo regularizer. This prevents the feature space from collapsing — a common failure mode in self-supervised learning where all representations converge to a small region. KoLeo encourages features to spread uniformly across the embedding space, like ensuring books in a library occupy every shelf rather than piling up in one corner:
Engineering at scale: making it feasible
Training a 1 billion parameter Vision Transformer on 142 million images requires solving several engineering challenges. The DINOv2 team contributed four key innovations that together made the training 2× faster and 3× more memory-efficient compared to iBOT:
FlashAttention. A custom implementation of memory-efficient that avoids materializing the full attention matrix. This is critical because 's memory scales quadratically with sequence length.
Sequence packing. The DINO algorithm uses both large crops (224×224) and small crops (98×98). Instead of forwarding them separately (wasting compute on padding), they are concatenated into one long sequence with a block-diagonal attention mask — preventing attention between different crops while processing them together.
Efficient stochastic depth. Instead of masking dropped residual blocks (computing them and then throwing away the result), the skipped blocks are never computed at all. With a 40% drop rate, this saves substantial compute and memory.
FSDP (Fully-Sharded Data Parallel). The 1B parameter model with its optimizer states requires 16 GB of memory. FSDP shards this across GPUs, with mixed-precision communication that cuts bandwidth by 50% compared to standard distributed training.
Distillation: one large teacher, many smaller students
Training the ViT-g model (1.1B parameters) from scratch is extremely expensive. For smaller models — ViT-S, ViT-B, ViT-L — DINOv2 uses instead of training from scratch.
The idea is simple: use the already-trained ViT-g as a frozen teacher and train smaller student models to reproduce its outputs. The same student-teacher framework used during pretraining is reused for distillation, with a few changes: the teacher is now frozen (the large ViT-g), masking and stochastic depth are removed for the student, and the final model is the of the student.
The surprising result: distilled models outperform models trained from scratch on every single benchmark. A distilled ViT-L nearly matches the ViT-g teacher, and sometimes even surpasses it. This makes distillation not just cheaper, but better.
Results: features that work everywhere
The defining property of DINOv2 is that its frozen features — extracted without any task-specific training — perform competitively or better than the best available features across a remarkable range of tasks:
Image classification. On ImageNet-1k linear evaluation, DINOv2 ViT-g achieves 86.5% top-1 accuracy — surpassing OpenCLIP-G (86.2%) and EVA-CLIP (86.4%), both of which used text supervision. On robustness benchmarks like ImageNet-A, DINOv2 reaches 75.9% vs 63.8% for OpenCLIP-G.
Semantic segmentation. With just a linear layer on top of frozen features, DINOv2 reaches 53.0 mIoU on ADE20k — matching MAE with full fine-tuning and a complex UperNet decoder (53.6). With a full decoder, it reaches 60.2 mIoU.
Depth estimation. DINOv2 features, with a simple DPT head, match or exceed dedicated depth estimation models — while being a completely general-purpose backbone. This capability directly enabled Depth Anything.
Instance retrieval. On Oxford-Hard, DINOv2 achieves 52.3 mAP vs 19.7 for OpenCLIP-G — nearly 3× better. This shows the features capture fine-grained identity, not just category-level similarity.
The fact that a single self-supervised model with frozen features can match or beat specialized models across all these tasks is the central achievement of DINOv2.
Seeing what the model sees: PCA of patch features
One of the most striking demonstrations of DINOv2's quality is visualizing its patch features using . When you take the patch tokens from the last layer and project them to their first three principal components, mapping each to a color channel (RGB), something remarkable emerges: the model has learned a consistent semantic mapping across images.
The same body part of a dog — say, the head — maps to the same color across different dog breeds, poses, and even between a real dog and a painting of one. Background pixels cluster separately from foreground. This shows that the model has learned, without any labels, to segment objects and match semantic parts across dramatically different images.
This is not just a pretty visualization — it proves the features encode rich spatial and semantic structure that traditional supervised features often lack.
From DINO to DINOv2: the road to visual foundation features
2020
BYOL and SimCLR
Established contrastive and self-distillation frameworks for self-supervised visual learning, showing that augmentation-based objectives can learn strong features from images alone.
2021
DINO (Caron et al.)
Introduced self-distillation with Vision Transformers, showing that ViT features trained with a momentum teacher contain emergent object segmentation in the attention maps — no labels needed.
2021
VICReg (Bardes et al.)
Proposed variance-invariance-covariance regularization to prevent representation collapse without negative pairs or momentum encoders, influencing DINOv2's feature spreading strategy.
2022
MAE (He et al.)
Showed that masked autoencoders learn powerful features by reconstructing masked image patches. Features require fine-tuning but the patch-level objective inspired DINOv2's iBOT loss component.
2022
iBOT (Zhou et al.)
Combined DINO-style self-distillation with masked image modeling at the patch level, creating the direct precursor to DINOv2's training recipe.
2023
DINOv2 (this paper)
Scaled self-supervised pretraining with curated data, engineering optimizations, and distillation to produce the first self-supervised model whose frozen features match weakly-supervised models across image and pixel-level tasks.
2024
Depth Anything (Yang et al.)
Built directly on DINOv2's frozen backbone to create a state-of-the-art monocular depth estimation model, validating DINOv2's promise as a universal visual feature extractor.
CitationOquab, Darcet, Moutakanni, Vo, Szafraniec, Khalidov, Fernandez, Haziza, Massa, El-Nouby, Assran, Ballas, Galuba, Howes, Huang, Li, Misra, Rabbat, Sharma, Synnaeve, Xu, Jegou, Mairal, Labatut, Joulin, Bojanowski. DINOv2: Learning Robust Visual Features Without Supervision. Transactions on Machine Learning Research (TMLR), 2023.
Terms in this paper
- Self-Supervised Learningالتعلم ذاتي الإشراف
- Vision Transformer (ViT)محوِّل الرؤية (ViT)
- Knowledge Distillationتقطير المعرفة
- Data Augmentationتعزيز البيانات
- Feature Extractionاستخلاص السمات
- Linear Probingالاختبار الخطي
- Semantic Segmentationالتجزئة الدلالية للصورة
- Contrastive Learningالتعلم التبايُني
- Masked Image Modelingنمذجة الصورة المُقنَّعة
- Backboneالبنية الأساسية
- Foundation Modelنموذج أساسي
- Teacher Modelنموذج المعلم
- Student Modelنموذج الطالب
- Exponential Moving Average (EMA)المتوسط المتحرك الأُسِّي
- Cosine Similarityتشابه جيب التمام