Self-Supervised Learning2021intermediate12 min read

Barlow Twins: Self-Supervised Learning via Redundancy Reduction

Barlow Twins: التعلّم ذاتي الإشراف عبر تقليل التكرار

Zbontar, J. · Jing, L. · Misra, I. · LeCun, Y. · Deny, S. — ICML

The problem

methods learn visual representations by making embeddings invariant to image distortions. But there's a lurking danger: the network can cheat by mapping every image to the same constant vector — a problem called . By 2020, avoiding collapse required intricate tricks: large batches of negative pairs (SimCLR), momentum encoders (MoCo, BYOL), stop gradients (SimSiam), or online clustering (SwAV). Each trick works, but none addresses the root cause — and each adds complexity that is hard to justify from first principles.

The contribution

Barlow Twins proposes a that naturally prevents collapse without negative pairs, momentum encoders, or Stop Gradients. It computes the matrix between the embeddings of two augmented views of the same batch, then penalizes deviation from the identity matrix. The diagonal terms enforce (same image, same features), while the off-diagonal terms enforce (different features carry different information). This is a direct application of neuroscientist H. Barlow's -reduction principle. The method achieves 73.2% top-1 on ImageNet linear evaluation with ResNet-50 and 55.0% in the 1% semi-supervised regime — the best among competing methods.

The impact

Barlow Twins demonstrated that a principled information-theoretic objective can replace the ad-hoc tricks used to prevent collapse. It directly inspired VICReg, which generalizes the idea into separate variance, invariance, and covariance terms. The redundancy-reduction principle has since been extended to multi-modal learning, reinforcement learning, and graph . The paper also showed that self-supervised methods can benefit from very high-dimensional projector outputs — a counter-intuitive finding that reshaped how practitioners design projection heads.

Imagine a team of scouts reporting on the same landscape. Each scout writes a checklist of features: "river present," "mountain visible," "forest dense." Two problems can arise. First, if all scouts copy each other word for word, you learn nothing new from having a team — that's collapse. Second, if every item on a scout's checklist says the same thing in different words, the report is redundant.

Barlow Twins solves both: it asks each pair of scouts to agree on the same items (diagonal = 1) but to talk about different things (off-diagonal = 0). The result is a team where every scout confirms the landscape is real, and every checklist item adds unique information.

The lurking danger: representation collapse

Self-supervised learning in typically works by creating two distorted views of the same image — through random cropping, color jittering, blurring, and other augmentations — and training a network to produce similar embeddings for both views. The intuition is clear: if the network can recognize that a cropped, blurred photo and a color-shifted version are the same dog, it must have learned meaningful features.

But there is a trivial shortcut: the network can map every image to the exact same point in space. This constant solution perfectly satisfies "both views should have similar embeddings" — because all embeddings are identical. This is called representation collapse, and it produces a completely useless model.

By 2020, researchers had developed several clever tricks to avoid collapse. SimCLR uses large batches of negative pairs to push apart embeddings of different images. BYOL adds a that updates slowly. SimSiam uses to prevent one branch from collapsing. Each works — but each adds complexity, and none explains why it works from first principles.

Open in Lab
Toggle between methods to see how each avoids collapse. Barlow Twins is the only one that does it through the loss function alone.
The demo wakes as you arrive…

The idea: make the cross-correlation matrix an identity

Barlow Twins starts the same way as other self-supervised methods: take an image, create two randomly augmented views YAY^A and YBY^B, and pass each through the same network fθf_\theta followed by a projector to get embeddings ZAZ^A and ZBZ^B.

The key innovation is what happens next. Instead of comparing individual pairs of embeddings, Barlow Twins computes the cross-correlation matrix C\mathcal{C} between the two sets of embeddings across the entire batch. This is a D×DD \times D matrix (where DD is the embedding dimension), and each entry Cij\mathcal{C}_{ij} measures how correlated dimension ii of ZAZ^A is with dimension jj of ZBZ^B across all samples in the batch.

The objective is beautifully simple: push C\mathcal{C} toward the identity matrix. That means each diagonal element Cii\mathcal{C}_{ii} should be 1 (the same feature in both views should be perfectly correlated), and each off-diagonal element Cij\mathcal{C}_{ij} for i≠ji \neq j should be 0 (different features should be uncorrelated).

Open in Lab
Follow the data flow: one image becomes two augmented views, both pass through the same encoder and projector, then the cross-correlation matrix is computed and pushed toward the identity.
The demo wakes as you arrive…

The loss function: invariance meets redundancy reduction

The loss has two complementary terms. The first is the invariance term: it penalizes diagonal elements of the cross-correlation matrix for deviating from 1. When Cii=1\mathcal{C}_{ii} = 1, it means that the ii-th feature in View A is perfectly correlated with the ii-th feature in View B — the representation is invariant to the augmentation. Think of it as asking: "Do both views agree on what each feature says about the image?"

The second is the term: it penalizes off-diagonal elements for deviating from 0. When Cij=0\mathcal{C}_{ij} = 0 for i≠ji \neq j, different embedding dimensions carry independent information. Think of it as asking: "Is each feature telling you something new, or just repeating what another feature already said?"

A hyperparameter λ\lambda balances the two terms. The authors found λ=5×10−3\lambda = 5 \times 10^{-3} works best — a gentle nudge toward decorrelation rather than a hard constraint.

LBT≜∑i(1−Cii)2⏟invariance term+λ∑i∑j≠iCij2⏟redundancy reduction term\mathcal{L}_{\text{BT}} \triangleq \underbrace{\sum_{i} (1 - \mathcal{C}_{ii})^{2}}_{\text{invariance term}} + \lambda \underbrace{\sum_{i} \sum_{j \neq i} \mathcal{C}_{ij}^{2}}_{\text{redundancy reduction term}}
Barlow Twins loss — invariance plus redundancy reduction — The first sum pushes diagonal entries of the cross-correlation matrix toward 1 (invariance to augmentations). The second sum pushes off-diagonal entries toward 0 (decorrelation between features). Together they force the cross-correlation matrix to approximate the identity matrix.

The cross-correlation matrix itself is computed by normalizing each embedding dimension to have zero mean and unit variance across the batch, then multiplying the two normalized embedding matrices:

Cij≜∑bzb,iA zb,jB∑b(zb,iA)2  ∑b(zb,jB)2\mathcal{C}_{ij} \triangleq \frac{\sum_{b} z^{A}_{b,i} \, z^{B}_{b,j}} {\sqrt{\sum_{b} (z^{A}_{b,i})^{2}} \;\sqrt{\sum_{b} (z^{B}_{b,j})^{2}}}
Cross-correlation matrix entry between embedding dimensions ii and jj — Each entry is a Pearson correlation coefficient measured across the batch dimension bb. The matrix C\mathcal{C} has shape D×DD \times D where DD is the embedding dimension (typically 8192), and values range from -1 to 1.
Open in Lab
Watch the cross-correlation matrix evolve during training. The diagonal should approach 1 (white) and the off-diagonal should approach 0 (dark).
The demo wakes as you arrive…

The deeper reason: Barlow's redundancy-reduction principle

The method is named after neuroscientist Horace Barlow, who proposed in 1961 that the goal of sensory processing is to recode highly redundant sensory inputs into a — a representation where each component is statistically independent of the others.

Barlow observed that the raw signals hitting our eyes are massively redundant: neighboring pixels in a natural image are highly correlated. The visual system's job, he argued, is to strip away this redundancy and produce a compact, efficient code where each neuron carries unique information. This is exactly what the off-diagonal term in the Barlow Twins loss does: it decorrelates the embedding dimensions, producing a factorial code.

More formally, the loss can be connected to the principle: find a representation that preserves maximum information about the input while being minimally informative about the specific augmentations applied. The invariance term maximizes information about the input; the redundancy reduction term ensures this information is encoded efficiently.

Architecture: symmetric twins with a big projector

The architecture of Barlow Twins is refreshingly simple and symmetric. Both branches use the exact same network — no predictor head, no momentum encoder, no Stop Gradient. The shared network consists of two parts:

Encoder (): A standard ResNet-50 (without its classification head), producing 2048-dimensional representations. These representations are what you use for downstream tasks after pre-training.

Projector: A 3-layer MLP with 8192 units in each layer, with and ReLU after the first two layers. The projector's output — the 8192-dimensional embeddings — is what enters the loss function.

A striking finding is that Barlow Twins benefits enormously from very high-dimensional projector outputs. While SimCLR and BYOL plateau at 256 or 2048 dimensions, Barlow Twins keeps improving up to 16,384 dimensions. This makes sense intuitively: more dimensions mean more entries in the cross-correlation matrix, which means the redundancy reduction term has more room to enforce independence.

Open in Lab
Unlike other methods, Barlow Twins keeps improving with higher projector dimensions. Click each method to compare.
The demo wakes as you arrive…

The idea in code

Barlow Twins — PyTorch-style pseudocodepython

Simplified to show the idea — not the real implementation.

# f: encoder + projector network
# lambda_: weight on the off-diagonal terms
# N: batch size, D: embedding dimensionality

for x in loader:
    # Step 1: create two augmented views of the same batch
    y_a, y_b = augment(x)

    # Step 2: compute embeddings through the SAME network
    z_a = f(y_a)  # shape: (N, D)
    z_b = f(y_b)  # shape: (N, D)

    # Step 3: normalize along the batch dimension
    z_a = (z_a - z_a.mean(0)) / z_a.std(0)  # zero mean, unit var per feature
    z_b = (z_b - z_b.mean(0)) / z_b.std(0)

    # Step 4: compute D×D cross-correlation matrix
    c = (z_a.T @ z_b) / N

    # Step 5: loss = push c toward identity matrix
    c_diff = (c - eye(D)).pow(2)
    # Off-diagonal terms are weighted by lambda
    off_diagonal(c_diff).mul_(lambda_)
    loss = c_diff.sum()

    # Step 6: standard backprop — gradients flow through BOTH branches
    loss.backward()
    optimizer.step()

Key advantages: simplicity and robustness

Three properties set Barlow Twins apart from the competition:

No negative pairs needed. Unlike SimCLR, which requires large batches to sample enough negative pairs, Barlow Twins operates purely on the statistics of embedding dimensions. The can be as small as 256 with almost no performance drop — compared to a ~4% drop for SimCLR.

No asymmetry needed. Unlike BYOL (momentum encoder) or SimSiam (Stop Gradient + predictor), Barlow Twins uses a perfectly symmetric architecture. Gradients flow through both branches identically. Adding asymmetry actually hurts performance slightly.

Benefits from high-dimensional projectors. While other methods plateau, Barlow Twins keeps improving as the projector output grows — from 256 to 8192 dimensions and beyond. More dimensions give the redundancy reduction term more room to enforce decorrelation.

Open in Lab
Barlow Twins maintains performance at small batch sizes where SimCLR degrades significantly.
The demo wakes as you arrive…

Barlow Twins vs. the field: a conceptual map

Self-supervised learning methods can be grouped by how they avoid collapse:

Contrastive methods (SimCLR, MoCo) use negative pairs — they explicitly push apart embeddings of different images. This works but requires either very large batches or memory banks of stored embeddings.

Asymmetric methods (BYOL, SimSiam) avoid collapse through architectural tricks: momentum encoders, predictor networks, Stop Gradients. They don't need negative pairs, but they don't explain why collapse is avoided.

Clustering methods (SwAV, DeepCluster) avoid collapse through online clustering and prototype assignment. Effective, but adds a non-differentiable component.

Redundancy reduction (Barlow Twins, and later VICReg) avoids collapse through the loss function itself, by decorrelating embedding dimensions. No negative pairs, no asymmetry, no clustering — just a principled information-theoretic objective.

Open in Lab
Adjust the diagonal and off-diagonal terms independently to see their effect on the learned representations.
The demo wakes as you arrive…

Results: competitive simplicity

Barlow Twins achieves strong results across multiple evaluation protocols:

Linear evaluation on ImageNet: 73.2% top-1 accuracy — comparable to BYOL (74.3%) and SwAV (75.3%), and ahead of SimCLR (69.3%) and MoCo v2 (71.1%).

Semi-supervised (1% labels): 55.0% top-1 — the best among all competing methods, surpassing BYOL (53.2%) and SwAV (53.9%). This is where Barlow Twins truly shines: when labeled data is extremely scarce, its representations prove most robust.

Semi-supervised (10% labels): 69.7% top-1 — on par with SwAV (70.2%) and ahead of BYOL (68.8%).

: Competitive performance on Places-205, VOC07, iNaturalist, and COCO , demonstrating that the learned features generalize well beyond ImageNet.

The most remarkable aspect is not any single number, but that these results come from a method that requires no negative pairs, no asymmetric architecture, and no special training tricks — just an encoder, a projector, and a principled loss function.

Context: the self-supervised learning landscape

  1. 2020

    SimCLR (Chen et al.)

    Simple contrastive framework using large batches of negative pairs. Requires batch sizes of 4096+ for best results, with an InfoNCE loss.

  2. 2020

    BYOL (Grill et al.)

    Bootstrap Your Own Latent — no negative pairs, but relies on a momentum encoder and predictor network for asymmetry.

  3. 2021

    SimSiam (Chen & He)

    Simplified siamese network — no negative pairs, no momentum encoder, but requires Stop Gradient on one branch plus a predictor head.

  4. 2021

    Barlow Twins (this paper)

    Redundancy reduction via cross-correlation — no negative pairs, no asymmetry, no Stop Gradient. Collapse is prevented by the loss function alone.

  5. 2021

    VICReg (Bardes et al.)

    Generalizes Barlow Twins into three explicit terms — Variance (prevent collapse), Invariance (match views), Covariance (decorrelate features). Uses the covariance matrix instead of the cross-correlation matrix.

CitationZbontar, Jing, Misra, LeCun, Deny. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. ICML, 2021.

Terms in this paper