Self-Supervised Learning2020intermediate13 min read

Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning

BYOL: مقاربة جديدة للتعلُّم الذاتي الإشراف بتمهيد التمثيلات الكامنة

Grill, J.-B. · Strub, F. · Altché, F. · Tallec, C. · Richemond, P. H. · Buchatskaya, E. · Doersch, C. · Pires, B. A. · Guo, Z. D. · Azar, M. G. · Piot, B. · Kavukcuoglu, K. · Munos, R. · Valko, M. — NeurIPS

The problem

By 2020, state-of-the-art methods for images — SimCLR, MoCo, and their variants — all relied on . They pulled together representations of different views of the same image (positive pairs) while pushing apart representations of different images (negative pairs). This worked well, but came with serious practical costs: SimCLR needed enormous batch sizes (4096+) to get enough negatives, MoCo needed a memory bank, and both were fragile to the choice of image augmentations. The field assumed negative pairs were essential to prevent — where the network outputs the same vector for every input. No one had shown otherwise at scale.

The contribution

BYOL introduced a self-supervised method that achieves state-of-the-art image representations without any negative pairs. It uses two networks — an online network and a — where the online network learns to predict the target's representations of augmented views. The target network is updated as a slow of the online network. An asymmetric predictor on the online branch prevents collapse. With a standard ResNet-50, BYOL reached 74.3% top-1 on ImageNet linear evaluation — surpassing all prior self-supervised methods — and 79.6% with a larger ResNet-200 (2×).

The impact

BYOL shattered the assumption that negative pairs are necessary for self-. It opened a new family of non-contrastive methods and directly inspired SimSiam (which removed the EMA entirely) and DINO (which combined self-distillation with Vision Transformers). The insight that a slow-moving target network plus an asymmetric predictor prevents collapse became a foundational design pattern in . BYOL also demonstrated remarkable robustness to and augmentation choices, making self-supervised learning more accessible to practitioners with limited compute.

Imagine two apprentice painters sitting back-to-back. Each sees a different angle of the same still-life arrangement. The first apprentice paints what she sees, then tries to predict what the second apprentice's painting looks like. But the second apprentice isn't painting in real time — she's working from a smoothed average of all the first apprentice's past techniques. Neither ever looks at paintings of other still-life arrangements to know what to avoid.

Over hundreds of sessions, the first apprentice's predictions become so accurate that her internal model of the scene captures its true 3D structure — all without a single "wrong example." That is BYOL: learning by predicting a slowly evolving version of yourself, not by contrasting against others.

The contrastive dilemma: why negative pairs are expensive

Before BYOL, the dominant paradigm in self-supervised visual representation learning was contrastive learning. The recipe was straightforward: take an image, create two augmented views (random crops, color jitter, blur), and train the network to make their representations similar. But this alone is not enough — the network could cheat by mapping every image to the same constant vector. To prevent this collapse, contrastive methods introduced negative pairs: representations of different images that must be pushed apart.

Think of it like a crowded conference room. Positive pairs are colleagues who should cluster together. Negative pairs are strangers who should spread out. Without the strangers, everyone would just pile into one corner.

The catch is that you need many negative pairs for this to work well. SimCLR used batch sizes of 4096 or more so that each image had thousands of negatives to push against. MoCo maintained a separate memory bank of past representations. Both approaches demanded significant computational resources, and their performance was highly sensitive to the augmentation strategy — removing color jittering from SimCLR, for example, caused accuracy to plummet by over 20 percentage points.

Open in Lab
Toggle between contrastive learning and BYOL to see the fundamental difference: contrastive methods need negative pairs to avoid collapse, while BYOL uses a predictor and moving-average target instead.
The demo wakes as you arrive…

BYOL's architecture: two networks, one goal

BYOL uses two neural networks that learn from each other: the online network and the target network. Both share the same architecture but have different weights.

The pipeline works like this: take an image and create two augmented views vv and v′v'. Feed vv to the online network and v′v' to the target network. The online network has three stages: an fθf_\theta (e.g., ResNet-50) that produces a representation yθy_\theta, a projector gθg_\theta (an MLP) that maps it to a projection zθz_\theta, and a predictor qθq_\theta (another MLP) that predicts the target's projection. The target network has only the encoder fξf_\xi and projector gξg_\xi — no predictor. This asymmetry is crucial.

The online network is trained by . The target network is never trained by gradients — instead, its weights ξ\xi are updated as a slow exponential moving average of the online weights θ\theta. At the end of training, the projector and predictor are discarded, and only the encoder fθf_\theta is kept as the learned representation.

Open in Lab
Explore BYOL's architecture. The online network (top) has an encoder, projector, and predictor. The target network (bottom) mirrors the first two but receives its weights via exponential moving average — no gradients flow to it.
The demo wakes as you arrive…

The loss: predicting the target projection

BYOL's training objective is intuitive: make the online network's prediction match the target network's projection. Both outputs are first ℓ2\ell_2-normalized (projected onto the unit hypersphere), then compared using — which, after normalization, is equivalent to a negative .

The purpose of normalization is to prevent the network from trivially solving the task by simply growing the magnitude of its outputs. By forcing everything onto the unit sphere, BYOL ensures that the network must encode directional information — the angle between vectors matters, not their length.

Lθ,ξ  ≜  ∥qˉθ(zθ)−zˉξ′∥22  =  2−2⋅⟨qθ(zθ), zξ′⟩∥qθ(zθ)∥2  ⋅  ∥zξ′∥2\mathcal{L}_{\theta,\xi} \;\triangleq\; \left\| \bar{q}_\theta(z_\theta) - \bar{z}'_\xi \right\|_2^2 \;=\; 2 - 2 \cdot \frac{\langle q_\theta(z_\theta),\, z'_\xi \rangle} {\|q_\theta(z_\theta)\|_2 \;\cdot\; \|z'_\xi\|_2}
BYOL loss — normalized MSE equals negative cosine similarity — The overline denotes ℓ2\ell_2-normalized vectors. The left side is the squared distance on the unit sphere between the online prediction qθ(zθ)q_\theta(z_\theta) and the target projection zξ′z'_\xi. The right side shows this equals 2−2cos⁡(α)2 - 2\cos(\alpha), where α\alpha is the angle between them. Minimizing this loss aligns the two vectors on the sphere.

The loss is symmetrized: we also feed v′v' to the online network and vv to the target network, computing a second loss L~θ,ξ\tilde{\mathcal{L}}_{\theta,\xi}. The total loss is their sum. Critically, gradients flow only through θ\theta (the online network). The target parameters ξ\xi are updated by a EMA — this is what makes BYOL fundamentally different from methods that minimize a joint loss over both networks.

Lθ,ξBYOL=Lθ,ξ+L~θ,ξ\mathcal{L}^{\text{BYOL}}_{\theta,\xi} = \mathcal{L}_{\theta,\xi} + \tilde{\mathcal{L}}_{\theta,\xi}
Symmetrized BYOL loss — By computing the loss in both directions — online predicting target, and the reverse — BYOL ensures that the representation captures information from both augmented views equally. This symmetrization improves stability and final performance.

The target network: a slowly evolving teacher

The target network's weights are never updated by gradients. Instead, after each training step, they are set to an exponential moving average of the online network's weights. This is controlled by a τ\tau that starts high (e.g., 0.996) and increases toward 1 during training following a cosine schedule.

Think of the target network as a patient mentor. It doesn't react to every new idea the student (online network) has. Instead, it slowly absorbs the student's knowledge over time, providing stable, consistent guidance. This stability is what prevents the online network from chasing its own noise and collapsing.

ξ←τ ξ+(1−τ) θ\xi \leftarrow \tau\,\xi + (1-\tau)\,\theta
EMA update rule for the target network — After each gradient step on θ\theta, the target weights ξ\xi are nudged slightly toward θ\theta. With τ=0.99\tau = 0.99, the target keeps 99% of its old weights and absorbs 1% from the online network. This creates a smoothed, delayed version of the online network that evolves much more slowly — like a time-averaged photograph versus a single snapshot.
Open in Lab
Adjust the decay rate τ and watch how the target network tracks the online network. High τ means slow, stable tracking. Low τ means rapid but noisy following.
The demo wakes as you arrive…

Why doesn't BYOL collapse?

This is the central mystery of BYOL. Without negative pairs to push representations apart, why doesn't the network converge to a trivial constant solution — outputting the same vector for every input?

The answer lies in the interplay of two components: the asymmetric predictor and the EMA target network. Neither alone is sufficient. Remove the predictor and keep the EMA? The network collapses. Remove the EMA and keep the predictor? Also collapse. Both must work together.

The intuition comes from thinking about what the predictor is actually learning. When the predictor is near-optimal, BYOL's loss becomes the conditional variance of the target projection given the online projection. A constant (collapsed) output would maximize this variance, not minimize it — because the target would still vary while the prediction stays fixed. So the gradient pushes the online network to capture more information, not less.

The EMA target network's role is to ensure the predictor stays near-optimal. If the target changed as fast as the online network (τ = 0), the predictor would always be playing catch-up and never be close to optimal. The slow EMA gives it time to converge, maintaining the anti-collapse dynamics.

∇θE ⁣[∥q∗(zθ)−zξ′∥22]=∇θE ⁣[∑iVar(zξ,i′∣zθ)]\nabla_\theta \mathbb{E}\!\left[\left\|q^*(z_\theta) - z'_\xi\right\|_2^2\right] = \nabla_\theta \mathbb{E}\!\left[\sum_i \mathrm{Var}(z'_{\xi,i} \mid z_\theta)\right]
Optimal predictor reduces BYOL loss to conditional variance — When the predictor q∗q^* is optimal (equal to the conditional expectation), the loss becomes the sum of conditional variances of each target feature given the online projection. A collapsed (constant) zθz_\theta would have higher conditional variance than an informative one — so gradients push toward retaining more information, not less.
Open in Lab
Toggle the predictor and EMA on/off to see when collapse happens. Only the combination of both prevents the representation from becoming constant.
The demo wakes as you arrive…

Data augmentation: creating meaningful views

BYOL creates two augmented views of each image using a composition of standard transformations: random cropping and resizing, random horizontal flip, color jittering (brightness, contrast, saturation, hue), random grayscale conversion, Gaussian blur, and solarization. The two augmentation pipelines T\mathcal{T} and T′\mathcal{T}' differ slightly in their parameter distributions.

A key finding is that BYOL is remarkably more robust to the choice of augmentations than contrastive methods. When restricted to only random crops — removing all color augmentations — BYOL's accuracy drops by 13.1 percentage points (to 59.4%), while SimCLR's drops by 27.6 points (to 40.3%). This makes sense: contrastive methods can exploit shortcuts (like matching color histograms) when those are the only distinguishing features between positive and negative pairs. BYOL, without negatives, is always incentivized to preserve all information in the target representation, not just the information that distinguishes images.

Open in Lab
Compare how BYOL and SimCLR degrade as augmentations are progressively removed. BYOL maintains much stronger performance with minimal augmentations.
The demo wakes as you arrive…

Training recipe: putting it all together

BYOL's training procedure can be summarized in five steps per iteration. First, sample an image and create two augmented views. Second, pass each view through the respective network (online gets one view, target gets the other). Third, compute the symmetrized loss — both directions. Fourth, update the online network parameters θ\theta by gradient descent on the loss. Fifth, update the target network parameters ξ\xi using the EMA rule.

The encoder is a ResNet-50 (or larger) pretrained on ImageNet. The projector is a 2-layer MLP with : 4096→2564096 \to 256. The predictor has the same architecture: 4096→2564096 \to 256. Training uses LARS optimizer with a cosine schedule, batch size of 4096, and 1000 epochs. The target decay rate follows a cosine schedule from τbase=0.996\tau_{\text{base}} = 0.996 to 1.

BYOL training step — simplified pseudocodepython

Simplified to show the idea — not the real implementation.

# Augment the same image twice
v1, v2 = augment(image, T), augment(image, T_prime)

# Online network: encoder → projector → predictor
y1 = encoder_online(v1)      # representation
z1 = projector_online(y1)     # projection
p1 = predictor(z1)            # prediction

# Target network: encoder → projector (no predictor!)
with stop_gradient():
    y2 = encoder_target(v2)
    z2 = projector_target(y2)


# Normalized MSE loss (= negative cosine similarity)
loss = norm_mse(p1, z2) + norm_mse(predictor(projector_online(encoder_online(v2))),
                                   projector_target(encoder_target(v1)))


# Update online network by gradient descent
theta = optimizer_step(theta, grad(loss, theta))

# Update target network by EMA (no gradients!)
xi = tau * xi + (1 - tau) * theta

Results: state of the art without negatives

Under linear evaluation on ImageNet (training a linear classifier on frozen representations), BYOL with a standard ResNet-50 achieved 74.3% top-1 accuracy — a 1.3% improvement over the previous self-supervised state of the art (InfoMin Aug. at 73.0%). With a ResNet-200 (2×), BYOL reached 79.6%, surpassing even the supervised ResNet-50 baseline of 76.5%.

On semi-supervised benchmarks using only 1% of ImageNet labels, BYOL outperformed SimCLR by 4.9 percentage points (53.2% vs 48.3%). On across 12 diverse datasets — from fine-grained recognition (Birds, Cars, Aircraft) to textures (DTD) to scenes (SUN397) — BYOL outperformed SimCLR on every single benchmark and matched or exceeded the supervised ImageNet baseline on most of them.

Beyond classification, BYOL's representation transferred well to (+1.1 mIoU over SimCLR on VOC), (+2.3 AP50 on VOC), and depth estimation on NYU v2. This breadth of strong transfer results demonstrates that BYOL learns genuinely general-purpose visual features.

Open in Lab
Compare BYOL's ImageNet accuracy against contrastive methods across different ResNet architectures. BYOL consistently leads without using negative pairs.
The demo wakes as you arrive…

Key ablations: what matters and what doesnt

The authors ran extensive ablations to isolate what makes BYOL work. The most illuminating findings concern the decay rate τ\tau. Setting τ=1\tau = 1 (never updating the target) gives only 18.8% accuracy — the target is a frozen random network, so the online network can only learn limited structure. Setting τ=0\tau = 0 (copying the online network instantly) gives 0.3% accuracy — complete collapse, because the target changes too fast for the predictor to track. All values between 0.9 and 0.999 give strong performance (68–72%), with τ=0.99\tau = 0.99 performing best at 72.5% over 300 epochs.

Another critical ablation showed that BYOL's batch size robustness is dramatically better than SimCLR's. Reducing the batch from 4096 to 256, SimCLR's accuracy drops by about 3 points, while BYOL barely changes. This is because BYOL's loss depends only on the — the batch size only affects batch normalization statistics, not the learning signal itself.

Legacy: the non-contrastive revolution

  1. 2018

    MoCo — momentum contrastive learning

    Introduced a momentum-updated encoder and a queue of negative representations, making contrastive learning practical with smaller batches. Established the momentum encoder pattern that BYOL would later repurpose.

  2. 2020

    SimCLR — simple contrastive framework

    Showed that a simple framework with large batches, strong augmentations, and a projection head could match complex contrastive methods. But required batch sizes of 4096+ for best results.

  3. 2020

    BYOL — this paper

    Proved that self-supervised learning works without negative pairs, using an asymmetric predictor and EMA target network to prevent collapse. Achieved 74.3% on ImageNet with ResNet-50, surpassing all contrastive methods.

  4. 2021

    SimSiam — simplifying further

    Showed that even the EMA target network is not strictly necessary — a simple Stop Gradient on one branch suffices to prevent collapse. Stripped BYOL to its minimal essentials.

  5. 2021

    DINO — self-distillation meets Vision Transformers

    Combined BYOL's teacher-student framework with Vision Transformers, discovering that self-supervised ViTs learn attention maps that naturally segment objects. Brought BYOL's principles into the Transformer era.

BYOL's impact extends beyond its specific algorithm. It established that the "contrastive = negative pairs" equation is false. You can learn excellent representations by simply predicting a slowly evolving version of your own outputs. This insight simplified the self-supervised learning pipeline, reduced compute requirements, and opened the door to methods like VICReg, Barlow Twins, and the entire non-contrastive learning family. Every modern self-supervised method that avoids negative pairs traces its intellectual ancestry back to BYOL.

CitationGrill, Strub, Altché, Tallec, Richemond, Buchatskaya, Doersch, Pires, Guo, Azar, Piot, Kavukcuoglu, Munos, Valko. Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning. NeurIPS, 2020.

Terms in this paper