Self-Supervised Learning2021intermediate11 min read

VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning

VICReg: تنظيم التباين والثبات والتغاير للتعلّم الذاتي الإشراف

Bardes, A. · Ponce, J. · LeCun, Y. — ICLR 2022

The problem

methods based on joint embedding architectures train two networks to produce similar representations for different views of the same image. The central challenge is : the networks can cheat by outputting constant vectors for every input, achieving perfect agreement while learning nothing useful. By 2021, solutions existed — contrastive losses that require massive negative-pair batches, momentum encoders, Stop Gradients, tricks — but all relied on implicit, hard-to-interpret mechanisms. None offered a clear, principled, and modular explanation of why collapse is prevented.

The contribution

VICReg introduces a loss function with three explicit, interpretable terms: an term (MSE between paired embeddings), a term ( that keeps each embedding dimension's standard deviation above a threshold), and a term (penalizing off-diagonal covariance to decorrelate dimensions). Together, these prevent both complete collapse (all outputs identical) and informational collapse (dimensions redundantly encoding the same information). VICReg matches state-of-the-art performance on ImageNet and downstream tasks without requiring weight sharing, momentum encoders, Stop Gradients, large batch sizes, or feature-wise normalization — making it the most architecturally flexible self-supervised method of its time.

The impact

VICReg demonstrated that collapse prevention in self-supervised learning does not require architectural tricks — simple, explicit regularization suffices. Its variance term has been adopted into other methods (BYOL, SimSiam) to improve stability. The modular loss design enabled multi-modal extensions — pairing image with text, audio with spectrogram — because branches need not share architecture. VICReg's principles directly influenced DINOv2, which scaled self-supervised vision to foundation-model level. The paper's conceptual clarity — naming the three failure modes and solving each with one term — became a template for designing self-supervised objectives.

Imagine an orchestra where every musician must play the same melody from different sheet music arrangements. Three things can go wrong: (1) all musicians play a single note forever — collapse. (2) They all play the melody correctly, but every instrument sounds identical — . (3) They play different notes, but disagree on the melody — no invariance.

VICReg is the conductor who enforces three rules: play the same melody (invariance), each instrument must actually make sound (variance), and each instrument must play its own part, not copy its neighbor (covariance). With these three rules, the orchestra produces a rich, harmonious performance — without needing anyone to watch over anyone else.

The collapse problem: why agreement alone is not enough

Modern self-supervised learning for images works by creating two views of the same image — typically through random cropping, color jittering, and blurring — and training an encoder to produce similar representations for both views. The intuition is sound: if two crops of the same photo yield similar embeddings, the encoder must be capturing the semantic content, not the crop-specific details.

But there is a devastating shortcut. If the encoder outputs the exact same constant vector for every input, the similarity between views is perfect — zero distance. The loss is minimized, the optimizer is happy, and the representations are completely useless. This is representation collapse.

There are actually two forms of collapse. Complete collapse is when all outputs converge to a single point — every image maps to the same vector. Informational collapse is subtler: the vectors vary, but all dimensions carry the same information, as if an 8192-dimensional vector really has only one degree of freedom. Both are fatal for downstream tasks.

Open in Lab
Click each tab to see what healthy representations, complete collapse, and informational collapse look like in embedding space.
The demo wakes as you arrive…

Before VICReg, solutions to collapse fell into two camps. Contrastive methods like SimCLR added a repulsive force: push apart embeddings of different images. This works but requires huge batches of negative pairs — memory-hungry and expensive. Architectural tricks like BYOL's momentum encoder or SimSiam's Stop Gradient avoided collapse through implicit biases, but nobody fully understood why they worked. Both approaches tied collapse prevention to specific architectural choices, limiting flexibility.

VICReg asked a different question: can we prevent collapse explicitly — with clear regularization terms that have interpretable meanings — without constraining the architecture?

The VICReg architecture: encoder, expander, and two views

VICReg uses a joint embedding architecture with two branches. Each branch has two parts: an encoder fθf_\theta (typically a ResNet-50) that produces representations yy, and an expander hϕh_\phi (a 3-layer MLP with 8192-dimensional layers) that maps representations into embeddings zz where the loss function operates.

The training flow is simple. Given an image ii, two random transformations tt and t′t' produce two views x=t(i)x = t(i) and x′=t′(i)x' = t'(i). Each view passes through the encoder then the expander, producing embeddings z=hϕ(fθ(x))z = h_\phi(f_\theta(x)) and z′=hϕ(fθ(x′))z' = h_\phi(f_\theta(x')). The VICReg loss is computed on zz and z′z'.

After pretraining, the expander is discarded. Only the encoder's representations are used for downstream tasks. The expander's role is to provide a higher-dimensional space where of dimensions can more effectively reduce dependencies (not just correlations) in the representation vectors.

Open in Lab
The VICReg pipeline: one image, two augmented views, shared encoder and expander, then three loss terms applied to the embeddings.
The demo wakes as you arrive…

The three loss terms: Variance, Invariance, Covariance

The VICReg loss function is the sum of three terms, each preventing a specific failure mode. Think of them as three guards, each watching for a different kind of cheating. Let Z=[z1,…,zn]Z = [z_1, \ldots, z_n] be a batch of nn embeddings of dimension dd, and let zjz^j denote the vector of all values at dimension jj across the batch.

Term 1 — Invariance (s): the paired embeddings must agree. This is the simplest term: it minimizes the between embeddings of two views of the same image. Without this term, the network has no reason to learn anything about the input.

s(Z,Z′)=1n∑i=1n∥zi−zi′∥22s(Z, Z') = \frac{1}{n} \sum_{i=1}^{n} \|z_i - z'_i\|_2^2
Invariance term — bring paired embeddings together — Simple MSE between each pair of embeddings from two views of the same image. No normalization is applied — unlike BYOL or SimSiam which use cosine similarity or l2-normalized MSE.

Term 2 — Variance (v): each dimension must stay alive. This is VICReg's key innovation. For each dimension jj of the embedding, the standard deviation across the batch is computed. If it falls below a target γ=1\gamma = 1, a penalty kicks in. This is a hinge loss — there is no penalty when the standard deviation is already above the threshold.

Why standard deviation and not variance? Because when embeddings are near collapse (all values close to the mean), the variance is close to zero, and its gradient also approaches zero — the optimizer cannot push back. The standard deviation has a steeper gradient near zero, providing the escape velocity the optimizer needs.

v(Z)=1d∑j=1dmax⁡ ⁣(0,  γ−S(zj,ϵ)),S(x,ϵ)=Var(x)+ϵv(Z) = \frac{1}{d} \sum_{j=1}^{d} \max\!\bigl(0,\; \gamma - S(z^j, \epsilon)\bigr), \qquad S(x,\epsilon) = \sqrt{\mathrm{Var}(x) + \epsilon}
Variance term — keep each dimension spread out — A hinge loss on the regularized standard deviation of each embedding dimension across the batch. The target γ=1\gamma = 1 means each dimension should have standard deviation at least 1. The small ϵ\epsilon prevents division-by-zero.

Term 3 — Covariance (c): dimensions must not copy each other. Even if every dimension has high variance, the representations are useless if all dimensions move together — the 8192-dimensional embedding would effectively be one-dimensional. The covariance term penalizes off-diagonal entries of the , pushing correlations between different dimensions toward zero.

c(Z)=1d∑i≠j[C(Z)]i,j2,C(Z)=1n−1∑i=1n(zi−zˉ)(zi−zˉ)⊤c(Z) = \frac{1}{d} \sum_{i \neq j} [C(Z)]_{i,j}^2, \qquad C(Z) = \frac{1}{n-1} \sum_{i=1}^{n}(z_i - \bar{z})(z_i - \bar{z})^\top
Covariance term — decorrelate embedding dimensions — The sum of squared off-diagonal entries of the covariance matrix over the batch. Driving these to zero means each embedding dimension carries unique information. This mechanism is borrowed from the Barlow Twins method.
Open in Lab
Adjust the three loss coefficients and watch how embeddings behave. Turn off variance and watch them collapse; turn off covariance and watch them become redundant.
The demo wakes as you arrive…

Putting it together: the full VICReg objective

The overall loss combines the three terms with weighting coefficients. Crucially, the variance and covariance terms are applied to each branch independently — not across branches. This independence is what allows VICReg to work with completely different architectures and even different input modalities on each branch.

ℓ(Z,Z′)=λ s(Z,Z′)+μ[v(Z)+v(Z′)]+ν[c(Z)+c(Z′)]\ell(Z, Z') = \lambda \, s(Z, Z') + \mu \bigl[v(Z) + v(Z')\bigr] + \nu \bigl[c(Z) + c(Z')\bigr]
Full VICReg loss — weighted sum of three terms — In practice, λ=μ=25\lambda = \mu = 25 and ν=1\nu = 1. The high weight on invariance and variance relative to covariance ensures the network first learns to agree and stay alive before worrying about redundancy.
VICReg loss — PyTorch pseudocodepython

Simplified to show the idea — not the real implementation.

# z_a, z_b: [N, D] embeddings from two augmented views
# --- Invariance loss ---
sim_loss = mse_loss(z_a, z_b)
# --- Variance loss ---
std_a = torch.sqrt(z_a.var(dim=0) + 1e-4)
std_b = torch.sqrt(z_b.var(dim=0) + 1e-4)
var_loss = torch.mean(F.relu(1 - std_a)) + torch.mean(F.relu(1 - std_b))
# --- Covariance loss ---
z_a = z_a - z_a.mean(dim=0)
z_b = z_b - z_b.mean(dim=0)
cov_a = (z_a.T @ z_a) / (N - 1)
cov_b = (z_b.T @ z_b) / (N - 1)
cov_loss = off_diagonal(cov_a).pow(2).sum() / D + off_diagonal(cov_b).pow(2).sum() / D
# --- Total loss ---
loss = 25 * sim_loss + 25 * var_loss + 1 * cov_loss

What VICReg does not need: architectural freedom

Most self-supervised methods at the time required at least one architectural constraint to prevent collapse. SimCLR needs large batches of negative pairs. BYOL needs a momentum encoder — a copy of the network whose weights slowly track the main encoder. SimSiam needs a Stop Gradient operation on one branch. Barlow Twins needs batch-wise normalization of the embeddings before computing cross-correlations. SwAV needs an online clustering step with Sinkhorn-Knopp balancing.

VICReg needs none of these. Because variance and covariance are computed on each branch independently, the two branches can have different architectures, different weights, and even process different input modalities (images on one branch, text or audio on the other). This modularity is VICReg's most practical advantage.

Open in Lab
Compare VICReg's simplicity against other self-supervised methods. Each card shows what constraints each method requires to avoid collapse.
The demo wakes as you arrive…

Relation to Barlow Twins: shared roots, different philosophy

VICReg borrows its covariance decorrelation mechanism from Barlow Twins. Both penalize off-diagonal entries of a covariance-like matrix. But they differ in a fundamental way.

Barlow Twins computes a cross-correlation matrix between the two branches' embeddings, then pushes it toward the identity matrix. This has two consequences: the diagonal (matching dimensions across branches) should be 1, and the off-diagonal (different dimensions) should be 0. The normalization required to make correlations meaningful couples the two branches together.

VICReg instead computes the covariance matrix of each branch separately, and uses a separate variance term to prevent shrinkage. This decoupling means the branches can have completely different output statistics — a critical advantage for multi-modal learning where image embeddings and text embeddings naturally live at different scales.

Open in Lab
See how the covariance matrix changes during training. Healthy training drives off-diagonal entries to zero while maintaining diagonal variance.
The demo wakes as you arrive…

Results: competitive without tricks

VICReg achieves 73.2% top-1 accuracy on ImageNet linear evaluation — matching Barlow Twins exactly and coming within 1% of BYOL (74.3%). On semi-supervised evaluation with 1% and 10% of labels, VICReg matches or slightly exceeds Barlow Twins. On transfer tasks including Places205 scene classification, VOC07 detection, and COCO instance segmentation, performance is consistently competitive.

The more revealing result is on multi-modal learning. When pairing images with text captions on MS-COCO, VICReg outperforms both the contrastive baseline (VSE++) and Barlow Twins on retrieval tasks — because its independent branch regularization handles the differing statistics of image and text encoders more gracefully.

Another key finding: adding VICReg's variance term to BYOL improved its performance by 0.9% at 100 epochs, suggesting that even methods with implicit collapse prevention suffer from subtle variance shrinkage that an explicit variance floor can fix.

Key design choices and ablations

The paper provides thorough ablation studies revealing which components truly matter. Without variance regularization, the representations immediately collapse regardless of what other components are present. Without covariance regularization, the network achieves only 57.5% — functional but far from competitive. Both are necessary; neither alone suffices.

The expander dimension has a dramatic effect: performance climbs from 55.9% at dimension 256 to 68.6% at 8192, saturating around 16384. This confirms the intuition that decorrelation in a higher-dimensional space is more effective at reducing dependencies.

VICReg is robust to batch size, dropping only 1.3% when batch size is reduced from 2048 to 128 — a significant advantage over contrastive methods like SimCLR that rely on large in-batch negative sets. No normalization of the embeddings is needed: removing standardization or l2-normalization actually helps performance, which is unique among self-supervised methods.

Context and legacy: from contrastive to explicit regularization

  1. 2020

    SimCLR & MoCo v2

    Contrastive methods achieve strong results but require large batches (SimCLR) or memory banks (MoCo) to provide enough negative pairs.

  2. 2020

    BYOL — no negatives needed

    Bootstrap Your Own Latent shows that collapse can be avoided without negative pairs, using a momentum encoder and Stop Gradient. But why this works remains unclear.

  3. 2021

    Barlow Twins — decorrelation objective

    Zbontar et al. propose redundancy reduction via cross-correlation, preventing informational collapse. But the cross-branch computation limits multi-modal use.

  4. 2021

    VICReg (this paper)

    Bardes, Ponce and LeCun unify collapse prevention into three explicit, interpretable terms. No architectural constraints needed. Multi-modal by design.

  5. 2023

    DINOv2 — scaling SSL vision

    Meta AI's DINOv2 builds on VICReg's principles to train foundation vision models at scale, producing general-purpose visual features without fine-tuning.

CitationBardes, Ponce, LeCun. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. ICLR, 2022.

Terms in this paper