Computer Vision2024intermediate11 min read

Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

Depth Anything: إطلاق قوة البيانات غير الموسومة واسعة النطاق

Yang, L. · Kang, B. · Huang, Z. · Xu, X. · Feng, J. · Zhao, H. — CVPR

The problem

— predicting how far each pixel is from the camera using a single image — is essential for robotics, autonomous driving, and augmented reality. But labeled depth datasets are expensive to build: they require LiDAR sensors, stereo rigs, or structure-from-motion pipelines, and even the best efforts produce only a few million labeled images. MiDaS, the previous leading model, trained on mixed labeled datasets but was still limited by data coverage, failing in foggy weather, ultra-remote distances, or unusual indoor scenes. The question was: can we break the data barrier without building more expensive labeled datasets?

The contribution

Three key ideas. First, a that collects 62 million diverse unlabeled images from eight public datasets and automatically annotates them with pseudo depth labels from a . Second, the discovery that naively using pseudo labels fails — but challenging the student with strong (, Gaussian blur, CutMix) during forces it to learn robust, generalizable representations. Third, an auxiliary loss that transfers rich semantic priors from a frozen DINOv2 , producing an encoder that excels at both depth estimation and .

The impact

Depth Anything set new state-of-the-art results in monocular depth estimation, surpassing MiDaS v3.1 across six benchmarks — even when MiDaS used those benchmarks for training. Its encoder also advanced semantic segmentation on Cityscapes and ADE20K. The model demonstrated that scaling unlabeled data with the right training strategy can outperform scaling labeled data, inspiring Depth Anything V2 and becoming a standard depth for ControlNet-based image generation, 3D reconstruction, and robotics.

Imagine a master art restorer training an apprentice. The restorer has studied a small number of perfectly documented paintings (labeled data), but the world is full of undocumented canvases in attics and garages. The restorer first examines each undocumented painting and writes down notes about its depth and perspective (pseudo labels). Then, instead of handing the apprentice clean originals, she gives them photos taken through frosted glass, with random patches swapped between paintings. The apprentice must figure out the true depth structure despite the distortions — and ends up developing sharper perception than the restorer herself.

Depth Anything follows exactly this strategy: a teacher model labels millions of free images, and a learns from deliberately corrupted versions, emerging with stronger than any model trained on labeled data alone.

The ceiling of labeled data

Monocular depth estimation asks a deceptively simple question: given a single photograph, how far is each pixel from the camera? Unlike stereo vision, which uses two cameras to triangulate distance, monocular estimation relies on a single viewpoint — the same information a human uses to judge depth from a flat painting.

Before Depth Anything, models like MiDaS trained on collections of labeled datasets gathered from LiDAR sensors, stereo matching, and structure-from-motion. These sources are expensive, limited in diversity, and hard to scale. MiDaS v3.1 assembled about 1.5 million labeled images from up to 12 datasets — impressive, but still a tiny sample of the visual world. When encountering fog, low light, unusual indoor layouts, or ultra-remote landscapes, MiDaS often failed catastrophically.

The core insight of Depth Anything is that unlabeled monocular images are essentially free: they exist on the internet in billions, cover every conceivable scene, and require no special sensors. The challenge is how to use them effectively.

Open in Lab
Compare the scale of labeled vs unlabeled data. Depth Anything uses 1.5M labeled images and 62M unlabeled images — a 40× expansion in data coverage.
The demo wakes as you arrive…

The data engine: from labeled seeds to 62 million images

The pipeline has two stages. In the first stage, a teacher model is trained on 1.5 million labeled images from six public datasets (BlendedMVS, DIML, HRWSI, IRS, MegaDepth, and TartanAir). The encoder uses DINOv2 pre-trained weights for strong visual features. Training uses an that aligns each prediction's scale and shift with the ground truth — essential because different datasets measure depth in different units.

In the second stage, this teacher annotates 62 million unlabeled images collected from eight large-scale public datasets: SA-1B, Open Images, BDD100K, ImageNet-21K, LSUN, Objects365, Places365, and Google Landmarks. Each unlabeled image is simply passed through the teacher to produce a dense pseudo depth map — one forward pass per image, no stereo matching or expensive reconstruction needed.

The result is a combined training set of labeled and pseudo-labeled images, spanning indoor, outdoor, driving, aerial, and artistic scenes. This diversity is the key to generalization.

Open in Lab
Walk through the data engine pipeline: labeled images train a teacher, which then annotates millions of unlabeled images for student training.
The demo wakes as you arrive…

The affine-invariant loss: comparing apples to oranges

Different depth datasets use different scales and offsets. A LiDAR dataset might report depth in meters, while a synthetic dataset uses arbitrary units. Training on mixed datasets requires a loss function that ignores these differences and focuses on the relative ordering of depth values.

The solution is to normalize both the prediction and ground truth before computing the error. Each depth map is centered by subtracting its median and scaled by dividing by its mean absolute deviation. This affine alignment ensures that only the shape of the depth map matters, not its absolute values.

d^i=di−median(d)s(d),s(d)=1HW∑i=1HW∣di−median(d)∣\hat{d}_i = \frac{d_i - \text{median}(d)}{s(d)}, \quad s(d) = \frac{1}{HW}\sum_{i=1}^{HW}|d_i - \text{median}(d)|
Affine normalization — aligning scale and shift — Each depth map dd is centered at zero by subtracting its median and scaled to unit deviation. The loss is then the mean absolute error between the normalized prediction and normalized ground truth: Ll=1HW∑∣d^i∗−d^i∣\mathcal{L}_l = \frac{1}{HW}\sum|\hat{d}_i^* - \hat{d}_i|. This allows joint training on datasets with incompatible depth scales.

The key insight: challenge the student, do not coddle it

Here is the surprise the authors discovered. When they naively combined labeled images with pseudo-labeled images and trained a student, performance did not improve. The student was essentially memorizing what the teacher already knew — no new knowledge was gained.

The fix was counterintuitive: make the student's job harder. During training on unlabeled images, the student receives heavily corrupted versions: strong color jittering, Gaussian blur, and CutMix (random rectangular patches swapped between images). But the teacher's pseudo labels come from clean, undistorted images. The student must reconstruct the clean depth from a damaged input.

Think of it like a musician practicing with a metronome set slightly too fast. The difficulty forces development of skills that would not emerge from comfortable repetition. The student is compelled to seek invariant visual cues — edges, textures, spatial relationships — that survive the perturbations, building representations that generalize far better to unseen scenes.

Open in Lab
Toggle perturbation types to see how the student receives corrupted inputs while learning from clean pseudo labels. The gap between what it sees and what it must predict forces deeper learning.
The demo wakes as you arrive…

CutMix for depth: stitching two worlds together

CutMix was originally designed for image classification. Depth Anything adapts it for depth estimation. A random rectangular region from one unlabeled image is pasted onto another, creating a composite image uabu_{ab}. The student must predict depth for the whole composite, but the loss is computed separately in the two regions — each region compared to its own teacher .

This is mathematically precise. The composite is formed as uab=ua⊙M+ub⊙(1−M)u_{ab} = u_a \odot M + u_b \odot (1-M), where MM is a binary mask with a rectangle set to 1. The loss on each region uses its own affine normalization, because the two source images can have completely different depth distributions:

Lu=∑MHW ρ ⁣(S(uab) ⁣⊙ ⁣M,  T(ua) ⁣⊙ ⁣M)+∑(1 ⁣− ⁣M)HW ρ ⁣(S(uab) ⁣⊙ ⁣(1 ⁣− ⁣M),  T(ub) ⁣⊙ ⁣(1 ⁣− ⁣M))\mathcal{L}_u = \frac{\sum M}{HW}\,\rho\!\bigl(S(u_{ab})\!\odot\!M,\;T(u_a)\!\odot\!M\bigr) + \frac{\sum(1\!-\!M)}{HW}\,\rho\!\bigl(S(u_{ab})\!\odot\!(1\!-\!M),\;T(u_b)\!\odot\!(1\!-\!M)\bigr)
CutMix depth loss — separate affine alignment per region — SS is the student, TT is the teacher, and ρ\rho is the affine-invariant mean absolute error. The loss on each region is weighted by the fraction of pixels it covers. CutMix is applied with 50% probability; the remaining samples use standard color distortions only.

Semantic-assisted perception: borrowing knowledge from DINOv2

Depth and semantics are deeply linked. Knowing that a region is "sky" tells the model it is infinitely far; knowing it is a "table" constrains its possible depth. Prior works tried auxiliary semantic segmentation tasks, but the authors found these failed when the depth model was already strong — mapping rich visual features to a small set of class labels discards too much information.

The solution was elegant: instead of predicting semantic classes, align the student's feature space with DINOv2's rich, continuous feature space. DINOv2 has extraordinary semantic capabilities (image retrieval, segmentation) even with frozen weights. A simple loss transfers this semantic knowledge:

Lfeat=1−1HW∑i=1HWcos⁡(fi, fi′)\mathcal{L}_{feat} = 1 - \frac{1}{HW}\sum_{i=1}^{HW}\cos(f_i,\,f'_i)
Feature alignment loss — inheriting semantic priors — fif_i is the student's feature at pixel ii, fi′f'_i is the corresponding feature from a frozen DINOv2 encoder. The cosine similarity pulls the two feature spaces together. A crucial detail: a tolerance margin α=0.85\alpha = 0.85 exempts pixels whose similarity already exceeds the threshold, because DINOv2 produces similar features for different parts of the same object (e.g. car front and rear), but these parts can have very different depths.
Open in Lab
See how the tolerance margin α prevents the depth encoder from losing part-level depth discrimination while still inheriting DINOv2's semantics.
The demo wakes as you arrive…

Putting it all together: the overall loss

The final training objective is an average of three losses, each serving a distinct role. The labeled loss Ll\mathcal{L}_l teaches accurate depth from human annotations. The unlabeled loss Lu\mathcal{L}_u expands coverage to millions of diverse scenes through pseudo labels and hard perturbations. The feature alignment loss Lfeat\mathcal{L}_{feat} transfers semantic understanding from DINOv2.

L=Ll+Lu+Lfeat\mathcal{L} = \mathcal{L}_l + \mathcal{L}_u + \mathcal{L}_{feat}
Overall training loss — The three losses are averaged with equal weight. Ll\mathcal{L}_l operates on labeled images with ground truth, Lu\mathcal{L}_u on unlabeled images with pseudo labels (after strong perturbations), and Lfeat\mathcal{L}_{feat} enforces feature-level consistency between the student and a frozen DINOv2 encoder.
Open in Lab
The complete Depth Anything pipeline. Click each component to see how labeled data, unlabeled data, perturbations, and feature alignment work together.
The demo wakes as you arrive…

Architecture: DINOv2 encoder + DPT decoder

Depth Anything uses a DINOv2 encoder with a DPT ( Transformer) , following MiDaS. The encoder extracts multi-scale features using a , and the DPT decoder fuses features from multiple layers to produce a dense depth prediction at the input resolution.

The model comes in three sizes: ViT-S (24.8M parameters), ViT-B (97.5M), and ViT-L (335.3M). Remarkably, the small ViT-S model — less than 1/10 the size of MiDaS — outperforms MiDaS on several benchmarks. The ViT-B model already clearly surpasses MiDaS ViT-L, demonstrating that the training strategy matters more than raw model size.

Key implementation details: images are resized with the shorter side to 518 pixels and cropped to 518 × 518 during training. The encoder is 5e-6, with the decoder at 10× higher. The labeled-to-unlabeled ratio in each batch is 1:2. The teacher always uses the ViT-L encoder for maximum annotation quality, even when training smaller student models.

Results: state of the art across the board

Zero-shot relative depth. Depth Anything ViT-L surpasses MiDaS ViT-L on all six unseen benchmarks (KITTI, NYUv2, Sintel, DDAD, ETH3D, DIODE). On KITTI, the AbsRel error drops from 0.127 to 0.076 and δ1\delta_1 accuracy jumps from 0.850 to 0.947 — and this is comparing against a MiDaS that actually trained on KITTI data, while Depth Anything never saw it.

Metric depth. the Depth Anything encoder for metric depth estimation using ZoeDepth's framework sets new records on both NYUv2 (δ1\delta_1: 0.964 → 0.984) and KITTI (δ1\delta_1: 0.978 → 0.982). The gains also transfer to zero-shot metric depth on unseen indoor and outdoor datasets.

Semantic segmentation. The encoder also advances semantic segmentation: 86.2 mIoU on Cityscapes (vs. 84.6 for ConvNeXt-XL) and 59.4 mIoU on ADE20K (vs. 58.3 for ViT-Adapter with BEiT-L). This confirms the encoder's potential as a universal multi-task backbone for both middle-level and high-level visual perception.

Open in Lab
Compare zero-shot depth estimation performance across six benchmarks. Toggle between Depth Anything sizes and MiDaS.
The demo wakes as you arrive…

Ablation: what each ingredient contributes

The ablation studies tell a clear story. Starting from labeled-only training (baseline AbsRel 0.085 on KITTI), adding unlabeled images without perturbations gives virtually no gain (0.085). Adding strong perturbations drops the error to 0.081 — the unlabeled data finally contributes. Adding the feature alignment loss drops it further to 0.076, a 10.6% relative improvement. The pattern is consistent across all six benchmarks.

The tolerance margin α\alpha also matters. Without it (α=1.0\alpha = 1.0, full alignment), the mean AbsRel across benchmarks is 0.188. With α=0.85\alpha = 0.85, it drops to 0.175. This confirms that blindly aligning features hurts depth discrimination within objects.

Another key finding: applying feature alignment only to unlabeled data works better than applying it to labeled data. The authors explain that labeled data has high-quality ground truth, so the semantic loss can interfere with learning from precise annotations. For pseudo-labeled data, however, the semantic constraint combats noise in the pseudo labels while adding knowledge.

Context and legacy

  1. 2020

    MiDaS — pioneering multi-dataset depth

    Introduced affine-invariant loss for combining multiple depth datasets, enabling the first generation of zero-shot monocular depth models.

  2. 2023

    DINOv2 — universal visual features

    Self-supervised Vision Transformer producing features that excel at semantic tasks with frozen weights. Depth Anything inherits these features via alignment.

  3. 2023

    ZoeDepth — relative to metric depth

    Showed that a strong relative depth model can be fine-tuned for metric depth, providing the framework Depth Anything later uses for metric evaluation.

  4. 2024

    Depth Anything (this paper)

    Unlocked 62M unlabeled images via a data engine and hard perturbation strategy, surpassing MiDaS across all benchmarks with strong multi-task capabilities.

  5. 2024

    Depth Anything V2

    Replaced pseudo labels from a teacher with synthetic data, further improving accuracy and becoming the standard depth model for downstream applications.

CitationYang, Kang, Huang, Xu, Feng, Zhao. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. CVPR, 2024.

Terms in this paper