Self-Supervised Learning2020intermediate13 min read
Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning
BYOL: مقاربة جديدة للتعلُّم الذاتي الإشراف بتمهيد التمثيلات الكامنة
Grill, J.-B. · Strub, F. · Altché, F. · Tallec, C. · Richemond, P. H. · Buchatskaya, E. · Doersch, C. · Pires, B. A. · Guo, Z. D. · Azar, M. G. · Piot, B. · Kavukcuoglu, K. · Munos, R. · Valko, M. — NeurIPS
The problem
By 2020, state-of-the-art methods for images — SimCLR, MoCo, and their variants — all relied on . They pulled together representations of different views of the same image (positive pairs) while pushing apart representations of different images (negative pairs). This worked well, but came with serious practical costs: SimCLR needed enormous batch sizes (4096+) to get enough negatives, MoCo needed a memory bank, and both were fragile to the choice of image augmentations. The field assumed negative pairs were essential to prevent — where the network outputs the same vector for every input. No one had shown otherwise at scale.
The contribution
BYOL introduced a self-supervised method that achieves state-of-the-art image representations without any negative pairs. It uses two networks — an online network and a — where the online network learns to predict the target's representations of augmented views. The target network is updated as a slow of the online network. An asymmetric predictor on the online branch prevents collapse. With a standard ResNet-50, BYOL reached 74.3% top-1 on ImageNet linear evaluation — surpassing all prior self-supervised methods — and 79.6% with a larger ResNet-200 (2×).
The impact
BYOL shattered the assumption that negative pairs are necessary for self-. It opened a new family of non-contrastive methods and directly inspired SimSiam (which removed the EMA entirely) and DINO (which combined self-distillation with Vision Transformers). The insight that a slow-moving target network plus an asymmetric predictor prevents collapse became a foundational design pattern in . BYOL also demonstrated remarkable robustness to and augmentation choices, making self-supervised learning more accessible to practitioners with limited compute.
Imagine two apprentice painters sitting back-to-back. Each sees a different angle of the same still-life arrangement. The first apprentice paints what she sees, then tries to predict what the second apprentice's painting looks like. But the second apprentice isn't painting in real time — she's working from a smoothed average of all the first apprentice's past techniques. Neither ever looks at paintings of other still-life arrangements to know what to avoid.
Over hundreds of sessions, the first apprentice's predictions become so accurate that her internal model of the scene captures its true 3D structure — all without a single "wrong example." That is BYOL: learning by predicting a slowly evolving version of yourself, not by contrasting against others.
The contrastive dilemma: why negative pairs are expensive
Before BYOL, the dominant paradigm in self-supervised visual representation learning was contrastive learning. The recipe was straightforward: take an image, create two augmented views (random crops, color jitter, blur), and train the network to make their representations similar. But this alone is not enough — the network could cheat by mapping every image to the same constant vector. To prevent this collapse, contrastive methods introduced negative pairs: representations of different images that must be pushed apart.
Think of it like a crowded conference room. Positive pairs are colleagues who should cluster together. Negative pairs are strangers who should spread out. Without the strangers, everyone would just pile into one corner.
The catch is that you need many negative pairs for this to work well. SimCLR used batch sizes of 4096 or more so that each image had thousands of negatives to push against. MoCo maintained a separate memory bank of past representations. Both approaches demanded significant computational resources, and their performance was highly sensitive to the augmentation strategy — removing color jittering from SimCLR, for example, caused accuracy to plummet by over 20 percentage points.
BYOL's architecture: two networks, one goal
BYOL uses two neural networks that learn from each other: the online network and the target network. Both share the same architecture but have different weights.
The pipeline works like this: take an image and create two augmented views and . Feed to the online network and to the target network. The online network has three stages: an (e.g., ResNet-50) that produces a representation , a projector (an MLP) that maps it to a projection , and a predictor (another MLP) that predicts the target's projection. The target network has only the encoder and projector — no predictor. This asymmetry is crucial.
The online network is trained by . The target network is never trained by gradients — instead, its weights are updated as a slow exponential moving average of the online weights . At the end of training, the projector and predictor are discarded, and only the encoder is kept as the learned representation.
The loss: predicting the target projection
BYOL's training objective is intuitive: make the online network's prediction match the target network's projection. Both outputs are first -normalized (projected onto the unit hypersphere), then compared using — which, after normalization, is equivalent to a negative .
The purpose of normalization is to prevent the network from trivially solving the task by simply growing the magnitude of its outputs. By forcing everything onto the unit sphere, BYOL ensures that the network must encode directional information — the angle between vectors matters, not their length.
The loss is symmetrized: we also feed to the online network and to the target network, computing a second loss . The total loss is their sum. Critically, gradients flow only through (the online network). The target parameters are updated by a EMA — this is what makes BYOL fundamentally different from methods that minimize a joint loss over both networks.
The target network: a slowly evolving teacher
The target network's weights are never updated by gradients. Instead, after each training step, they are set to an exponential moving average of the online network's weights. This is controlled by a that starts high (e.g., 0.996) and increases toward 1 during training following a cosine schedule.
Think of the target network as a patient mentor. It doesn't react to every new idea the student (online network) has. Instead, it slowly absorbs the student's knowledge over time, providing stable, consistent guidance. This stability is what prevents the online network from chasing its own noise and collapsing.
Why doesn't BYOL collapse?
This is the central mystery of BYOL. Without negative pairs to push representations apart, why doesn't the network converge to a trivial constant solution — outputting the same vector for every input?
The answer lies in the interplay of two components: the asymmetric predictor and the EMA target network. Neither alone is sufficient. Remove the predictor and keep the EMA? The network collapses. Remove the EMA and keep the predictor? Also collapse. Both must work together.
The intuition comes from thinking about what the predictor is actually learning. When the predictor is near-optimal, BYOL's loss becomes the conditional variance of the target projection given the online projection. A constant (collapsed) output would maximize this variance, not minimize it — because the target would still vary while the prediction stays fixed. So the gradient pushes the online network to capture more information, not less.
The EMA target network's role is to ensure the predictor stays near-optimal. If the target changed as fast as the online network (τ = 0), the predictor would always be playing catch-up and never be close to optimal. The slow EMA gives it time to converge, maintaining the anti-collapse dynamics.
Data augmentation: creating meaningful views
BYOL creates two augmented views of each image using a composition of standard transformations: random cropping and resizing, random horizontal flip, color jittering (brightness, contrast, saturation, hue), random grayscale conversion, Gaussian blur, and solarization. The two augmentation pipelines and differ slightly in their parameter distributions.
A key finding is that BYOL is remarkably more robust to the choice of augmentations than contrastive methods. When restricted to only random crops — removing all color augmentations — BYOL's accuracy drops by 13.1 percentage points (to 59.4%), while SimCLR's drops by 27.6 points (to 40.3%). This makes sense: contrastive methods can exploit shortcuts (like matching color histograms) when those are the only distinguishing features between positive and negative pairs. BYOL, without negatives, is always incentivized to preserve all information in the target representation, not just the information that distinguishes images.
Training recipe: putting it all together
BYOL's training procedure can be summarized in five steps per iteration. First, sample an image and create two augmented views. Second, pass each view through the respective network (online gets one view, target gets the other). Third, compute the symmetrized loss — both directions. Fourth, update the online network parameters by gradient descent on the loss. Fifth, update the target network parameters using the EMA rule.
The encoder is a ResNet-50 (or larger) pretrained on ImageNet. The projector is a 2-layer MLP with : . The predictor has the same architecture: . Training uses LARS optimizer with a cosine schedule, batch size of 4096, and 1000 epochs. The target decay rate follows a cosine schedule from to 1.
Simplified to show the idea — not the real implementation.
# Augment the same image twice
v1, v2 = augment(image, T), augment(image, T_prime)
# Online network: encoder → projector → predictor
y1 = encoder_online(v1) # representation
z1 = projector_online(y1) # projection
p1 = predictor(z1) # prediction
# Target network: encoder → projector (no predictor!)
with stop_gradient():
y2 = encoder_target(v2)
z2 = projector_target(y2)
# Normalized MSE loss (= negative cosine similarity)
loss = norm_mse(p1, z2) + norm_mse(predictor(projector_online(encoder_online(v2))),
projector_target(encoder_target(v1)))
# Update online network by gradient descent
theta = optimizer_step(theta, grad(loss, theta))
# Update target network by EMA (no gradients!)
xi = tau * xi + (1 - tau) * thetaResults: state of the art without negatives
Under linear evaluation on ImageNet (training a linear classifier on frozen representations), BYOL with a standard ResNet-50 achieved 74.3% top-1 accuracy — a 1.3% improvement over the previous self-supervised state of the art (InfoMin Aug. at 73.0%). With a ResNet-200 (2×), BYOL reached 79.6%, surpassing even the supervised ResNet-50 baseline of 76.5%.
On semi-supervised benchmarks using only 1% of ImageNet labels, BYOL outperformed SimCLR by 4.9 percentage points (53.2% vs 48.3%). On across 12 diverse datasets — from fine-grained recognition (Birds, Cars, Aircraft) to textures (DTD) to scenes (SUN397) — BYOL outperformed SimCLR on every single benchmark and matched or exceeded the supervised ImageNet baseline on most of them.
Beyond classification, BYOL's representation transferred well to (+1.1 mIoU over SimCLR on VOC), (+2.3 AP50 on VOC), and depth estimation on NYU v2. This breadth of strong transfer results demonstrates that BYOL learns genuinely general-purpose visual features.
Key ablations: what matters and what doesnt
The authors ran extensive ablations to isolate what makes BYOL work. The most illuminating findings concern the decay rate . Setting (never updating the target) gives only 18.8% accuracy — the target is a frozen random network, so the online network can only learn limited structure. Setting (copying the online network instantly) gives 0.3% accuracy — complete collapse, because the target changes too fast for the predictor to track. All values between 0.9 and 0.999 give strong performance (68–72%), with performing best at 72.5% over 300 epochs.
Another critical ablation showed that BYOL's batch size robustness is dramatically better than SimCLR's. Reducing the batch from 4096 to 256, SimCLR's accuracy drops by about 3 points, while BYOL barely changes. This is because BYOL's loss depends only on the — the batch size only affects batch normalization statistics, not the learning signal itself.
Legacy: the non-contrastive revolution
2018
MoCo — momentum contrastive learning
Introduced a momentum-updated encoder and a queue of negative representations, making contrastive learning practical with smaller batches. Established the momentum encoder pattern that BYOL would later repurpose.
2020
SimCLR — simple contrastive framework
Showed that a simple framework with large batches, strong augmentations, and a projection head could match complex contrastive methods. But required batch sizes of 4096+ for best results.
2020
BYOL — this paper
Proved that self-supervised learning works without negative pairs, using an asymmetric predictor and EMA target network to prevent collapse. Achieved 74.3% on ImageNet with ResNet-50, surpassing all contrastive methods.
2021
SimSiam — simplifying further
Showed that even the EMA target network is not strictly necessary — a simple Stop Gradient on one branch suffices to prevent collapse. Stripped BYOL to its minimal essentials.
2021
DINO — self-distillation meets Vision Transformers
Combined BYOL's teacher-student framework with Vision Transformers, discovering that self-supervised ViTs learn attention maps that naturally segment objects. Brought BYOL's principles into the Transformer era.
BYOL's impact extends beyond its specific algorithm. It established that the "contrastive = negative pairs" equation is false. You can learn excellent representations by simply predicting a slowly evolving version of your own outputs. This insight simplified the self-supervised learning pipeline, reduced compute requirements, and opened the door to methods like VICReg, Barlow Twins, and the entire non-contrastive learning family. Every modern self-supervised method that avoids negative pairs traces its intellectual ancestry back to BYOL.
CitationGrill, Strub, Altché, Tallec, Richemond, Buchatskaya, Doersch, Pires, Guo, Azar, Piot, Kavukcuoglu, Munos, Valko. Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning. NeurIPS, 2020.
Terms in this paper
- Self-Supervised Learningالتعلم ذاتي الإشراف
- Contrastive Learningالتعلم التبايُني
- Representation Learningتعلم التمثيلات الرقمية
- Target Networkشبكة الهدف
- Exponential Moving Averageالمتوسط المتحرك الأُسّي
- Projection Headرأس الإسقاط
- Positive Pairالزوج الإيجابي
- Negative Pairالزوج السلبي
- Data Augmentationتعزيز البيانات
- Cosine Similarityتشابه جيب التمام
- Backboneالبنية الأساسية
- Encoderالمُرمِّز
- Linear Probeالمسبار الخطي
- Downstream Taskالمهمة اللاحقة
- Batch Normalizationتسوية الدفعات الحسابية