Computer Vision2024intermediate11 min read
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
Depth Anything: إطلاق قوة البيانات غير الموسومة واسعة النطاق
Yang, L. · Kang, B. · Huang, Z. · Xu, X. · Feng, J. · Zhao, H. — CVPR
The problem
— predicting how far each pixel is from the camera using a single image — is essential for robotics, autonomous driving, and augmented reality. But labeled depth datasets are expensive to build: they require LiDAR sensors, stereo rigs, or structure-from-motion pipelines, and even the best efforts produce only a few million labeled images. MiDaS, the previous leading model, trained on mixed labeled datasets but was still limited by data coverage, failing in foggy weather, ultra-remote distances, or unusual indoor scenes. The question was: can we break the data barrier without building more expensive labeled datasets?
The contribution
Three key ideas. First, a that collects 62 million diverse unlabeled images from eight public datasets and automatically annotates them with pseudo depth labels from a . Second, the discovery that naively using pseudo labels fails — but challenging the student with strong (, Gaussian blur, CutMix) during forces it to learn robust, generalizable representations. Third, an auxiliary loss that transfers rich semantic priors from a frozen DINOv2 , producing an encoder that excels at both depth estimation and .
The impact
Depth Anything set new state-of-the-art results in monocular depth estimation, surpassing MiDaS v3.1 across six benchmarks — even when MiDaS used those benchmarks for training. Its encoder also advanced semantic segmentation on Cityscapes and ADE20K. The model demonstrated that scaling unlabeled data with the right training strategy can outperform scaling labeled data, inspiring Depth Anything V2 and becoming a standard depth for ControlNet-based image generation, 3D reconstruction, and robotics.
Imagine a master art restorer training an apprentice. The restorer has studied a small number of perfectly documented paintings (labeled data), but the world is full of undocumented canvases in attics and garages. The restorer first examines each undocumented painting and writes down notes about its depth and perspective (pseudo labels). Then, instead of handing the apprentice clean originals, she gives them photos taken through frosted glass, with random patches swapped between paintings. The apprentice must figure out the true depth structure despite the distortions — and ends up developing sharper perception than the restorer herself.
Depth Anything follows exactly this strategy: a teacher model labels millions of free images, and a learns from deliberately corrupted versions, emerging with stronger than any model trained on labeled data alone.
The ceiling of labeled data
Monocular depth estimation asks a deceptively simple question: given a single photograph, how far is each pixel from the camera? Unlike stereo vision, which uses two cameras to triangulate distance, monocular estimation relies on a single viewpoint — the same information a human uses to judge depth from a flat painting.
Before Depth Anything, models like MiDaS trained on collections of labeled datasets gathered from LiDAR sensors, stereo matching, and structure-from-motion. These sources are expensive, limited in diversity, and hard to scale. MiDaS v3.1 assembled about 1.5 million labeled images from up to 12 datasets — impressive, but still a tiny sample of the visual world. When encountering fog, low light, unusual indoor layouts, or ultra-remote landscapes, MiDaS often failed catastrophically.
The core insight of Depth Anything is that unlabeled monocular images are essentially free: they exist on the internet in billions, cover every conceivable scene, and require no special sensors. The challenge is how to use them effectively.
The data engine: from labeled seeds to 62 million images
The pipeline has two stages. In the first stage, a teacher model is trained on 1.5 million labeled images from six public datasets (BlendedMVS, DIML, HRWSI, IRS, MegaDepth, and TartanAir). The encoder uses DINOv2 pre-trained weights for strong visual features. Training uses an that aligns each prediction's scale and shift with the ground truth — essential because different datasets measure depth in different units.
In the second stage, this teacher annotates 62 million unlabeled images collected from eight large-scale public datasets: SA-1B, Open Images, BDD100K, ImageNet-21K, LSUN, Objects365, Places365, and Google Landmarks. Each unlabeled image is simply passed through the teacher to produce a dense pseudo depth map — one forward pass per image, no stereo matching or expensive reconstruction needed.
The result is a combined training set of labeled and pseudo-labeled images, spanning indoor, outdoor, driving, aerial, and artistic scenes. This diversity is the key to generalization.
The affine-invariant loss: comparing apples to oranges
Different depth datasets use different scales and offsets. A LiDAR dataset might report depth in meters, while a synthetic dataset uses arbitrary units. Training on mixed datasets requires a loss function that ignores these differences and focuses on the relative ordering of depth values.
The solution is to normalize both the prediction and ground truth before computing the error. Each depth map is centered by subtracting its median and scaled by dividing by its mean absolute deviation. This affine alignment ensures that only the shape of the depth map matters, not its absolute values.
The key insight: challenge the student, do not coddle it
Here is the surprise the authors discovered. When they naively combined labeled images with pseudo-labeled images and trained a student, performance did not improve. The student was essentially memorizing what the teacher already knew — no new knowledge was gained.
The fix was counterintuitive: make the student's job harder. During training on unlabeled images, the student receives heavily corrupted versions: strong color jittering, Gaussian blur, and CutMix (random rectangular patches swapped between images). But the teacher's pseudo labels come from clean, undistorted images. The student must reconstruct the clean depth from a damaged input.
Think of it like a musician practicing with a metronome set slightly too fast. The difficulty forces development of skills that would not emerge from comfortable repetition. The student is compelled to seek invariant visual cues — edges, textures, spatial relationships — that survive the perturbations, building representations that generalize far better to unseen scenes.
CutMix for depth: stitching two worlds together
CutMix was originally designed for image classification. Depth Anything adapts it for depth estimation. A random rectangular region from one unlabeled image is pasted onto another, creating a composite image . The student must predict depth for the whole composite, but the loss is computed separately in the two regions — each region compared to its own teacher .
This is mathematically precise. The composite is formed as , where is a binary mask with a rectangle set to 1. The loss on each region uses its own affine normalization, because the two source images can have completely different depth distributions:
Semantic-assisted perception: borrowing knowledge from DINOv2
Depth and semantics are deeply linked. Knowing that a region is "sky" tells the model it is infinitely far; knowing it is a "table" constrains its possible depth. Prior works tried auxiliary semantic segmentation tasks, but the authors found these failed when the depth model was already strong — mapping rich visual features to a small set of class labels discards too much information.
The solution was elegant: instead of predicting semantic classes, align the student's feature space with DINOv2's rich, continuous feature space. DINOv2 has extraordinary semantic capabilities (image retrieval, segmentation) even with frozen weights. A simple loss transfers this semantic knowledge:
Putting it all together: the overall loss
The final training objective is an average of three losses, each serving a distinct role. The labeled loss teaches accurate depth from human annotations. The unlabeled loss expands coverage to millions of diverse scenes through pseudo labels and hard perturbations. The feature alignment loss transfers semantic understanding from DINOv2.
Architecture: DINOv2 encoder + DPT decoder
Depth Anything uses a DINOv2 encoder with a DPT ( Transformer) , following MiDaS. The encoder extracts multi-scale features using a , and the DPT decoder fuses features from multiple layers to produce a dense depth prediction at the input resolution.
The model comes in three sizes: ViT-S (24.8M parameters), ViT-B (97.5M), and ViT-L (335.3M). Remarkably, the small ViT-S model — less than 1/10 the size of MiDaS — outperforms MiDaS on several benchmarks. The ViT-B model already clearly surpasses MiDaS ViT-L, demonstrating that the training strategy matters more than raw model size.
Key implementation details: images are resized with the shorter side to 518 pixels and cropped to 518 × 518 during training. The encoder is 5e-6, with the decoder at 10× higher. The labeled-to-unlabeled ratio in each batch is 1:2. The teacher always uses the ViT-L encoder for maximum annotation quality, even when training smaller student models.
Results: state of the art across the board
Zero-shot relative depth. Depth Anything ViT-L surpasses MiDaS ViT-L on all six unseen benchmarks (KITTI, NYUv2, Sintel, DDAD, ETH3D, DIODE). On KITTI, the AbsRel error drops from 0.127 to 0.076 and accuracy jumps from 0.850 to 0.947 — and this is comparing against a MiDaS that actually trained on KITTI data, while Depth Anything never saw it.
Metric depth. the Depth Anything encoder for metric depth estimation using ZoeDepth's framework sets new records on both NYUv2 (: 0.964 → 0.984) and KITTI (: 0.978 → 0.982). The gains also transfer to zero-shot metric depth on unseen indoor and outdoor datasets.
Semantic segmentation. The encoder also advances semantic segmentation: 86.2 mIoU on Cityscapes (vs. 84.6 for ConvNeXt-XL) and 59.4 mIoU on ADE20K (vs. 58.3 for ViT-Adapter with BEiT-L). This confirms the encoder's potential as a universal multi-task backbone for both middle-level and high-level visual perception.
Ablation: what each ingredient contributes
The ablation studies tell a clear story. Starting from labeled-only training (baseline AbsRel 0.085 on KITTI), adding unlabeled images without perturbations gives virtually no gain (0.085). Adding strong perturbations drops the error to 0.081 — the unlabeled data finally contributes. Adding the feature alignment loss drops it further to 0.076, a 10.6% relative improvement. The pattern is consistent across all six benchmarks.
The tolerance margin also matters. Without it (, full alignment), the mean AbsRel across benchmarks is 0.188. With , it drops to 0.175. This confirms that blindly aligning features hurts depth discrimination within objects.
Another key finding: applying feature alignment only to unlabeled data works better than applying it to labeled data. The authors explain that labeled data has high-quality ground truth, so the semantic loss can interfere with learning from precise annotations. For pseudo-labeled data, however, the semantic constraint combats noise in the pseudo labels while adding knowledge.
Context and legacy
2020
MiDaS — pioneering multi-dataset depth
Introduced affine-invariant loss for combining multiple depth datasets, enabling the first generation of zero-shot monocular depth models.
2023
DINOv2 — universal visual features
Self-supervised Vision Transformer producing features that excel at semantic tasks with frozen weights. Depth Anything inherits these features via alignment.
2023
ZoeDepth — relative to metric depth
Showed that a strong relative depth model can be fine-tuned for metric depth, providing the framework Depth Anything later uses for metric evaluation.
2024
Depth Anything (this paper)
Unlocked 62M unlabeled images via a data engine and hard perturbation strategy, surpassing MiDaS across all benchmarks with strong multi-task capabilities.
2024
Depth Anything V2
Replaced pseudo labels from a teacher with synthetic data, further improving accuracy and becoming the standard depth model for downstream applications.
CitationYang, Kang, Huang, Xu, Feng, Zhao. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. CVPR, 2024.
Terms in this paper
- Foundation Modelنموذج أساسي
- Semi-Supervised Learningالتعلم شبه المُوجَّه
- Teacher Modelنموذج المعلم
- Student Modelنموذج الطالب
- Data Augmentationتعزيز البيانات
- Data Engineمحرّك بيانات
- Pseudo Labelتوسيمة زائفة
- Self-Trainingالتدريب الذاتي
- Zero-Shotالنمط الصفري
- Transfer Learningنقل التعلم
- Feature Alignmentمحاذاة السمات
- Cosine Similarityتشابه جيب التمام
- Encoderالمُرمِّز
- Decoderمفكّ الترميز
- Backboneالبنية الأساسية