Representation Learning2021intermediate12 min read
Exploring Simple Siamese Representation Learning
استكشاف تعلُّم التمثيلات بشبكات سيامية بسيطة
Chen, X. · He, K. — CVPR
The problem
By 2020, self-supervised visual representation learning had made remarkable progress through methods like SimCLR, MoCo, SwAV, and BYOL. But each relied on a specific trick to avoid — the failure mode where a network outputs the same constant vector for every input. SimCLR needed negative pairs and large batches. MoCo needed a and a queue. SwAV needed online . BYOL needed a momentum . The question remained: which of these components is truly essential, and which are incidental?
The contribution
SimSiam: a minimalist that learns meaningful representations without negative pairs, without large batches, and without a momentum encoder. The architecture is simply a shared encoder with a prediction on one branch and a on the other. The paper demonstrates that the Stop Gradient operation alone is the essential ingredient for preventing collapse, and provides a hypothesis that SimSiam implicitly solves an Expectation-Maximization-like alternating optimization problem. SimSiam achieves 71.3% top-1 accuracy on ImageNet linear evaluation (800 epochs) and competitive results on detection and segmentation.
The impact
SimSiam stripped to its conceptual bones, proving that the Siamese architecture itself — not the bells and whistles around it — is the core reason these methods work. It served as a unifying hub connecting SimCLR, BYOL, SwAV, and MoCo: each can be described as "SimSiam plus one extra component." This reframing shifted the field's understanding and influenced subsequent work like Barlow Twins, VICReg, and other non-contrastive methods that continued simplifying the recipe.
Imagine a sculptor and a photographer both studying the same statue. The sculptor makes a clay copy; the photographer takes a snapshot. A judge compares the two and says "make them more alike." But here's the twist: only the sculptor gets to hear the feedback. The photographer's snapshot is frozen — treated as a fixed target.
Without this one-way feedback rule, both artists would take the laziest shortcut: they'd both produce a featureless grey blob, identical for every statue. That blob is "collapse."
SimSiam discovered that this asymmetric feedback — one side learns, the other stays still — is all you need to prevent collapse. No need for negative examples, no need for thousands of different statues in each batch, no need for a slowly-updating "shadow photographer." Just the one-way rule.
The enemy: representation collapse
The dream of self-supervised learning is simple: teach a network to understand images without any human labels. The strategy? Show the network two augmented views of the same image — cropped differently, color-jittered, maybe blurred — and train it to produce similar representations for both. If it can match views of the same image, it must be learning something meaningful about the image's content.
But there is a devastating shortcut. The network could learn to output the exact same vector for every input. Two views of a cat? Same vector. Two views of a car? Same vector. The similarity is perfect — and the representation is perfectly useless. This failure mode is called representation collapse, and every self-supervised method before SimSiam needed a specific mechanism to prevent it.
Before SimSiam, the field had developed four distinct anti-collapse strategies, each adding complexity. SimCLR pushed apart representations of different images (negative pairs), requiring enormous batches of 4096 or more to provide enough negatives. MoCo maintained a queue of negative representations using a momentum encoder — a slowly-updating copy of the main network. SwAV replaced negatives with online clustering using the Sinkhorn-Knopp transform. BYOL eliminated negatives entirely but relied on a momentum encoder to provide stable targets. Each method worked, but each came with significant overhead. The question SimSiam asked was radical: what if we remove all of these?
SimSiam architecture: radical simplicity
SimSiam's architecture is strikingly minimal. Start with an image x. Create two augmented views x₁ and x₂ using random cropping, color jittering, horizontal flipping, and Gaussian blurring. Both views pass through the same encoder f — a ResNet-50 followed by a 3-layer projection MLP. This produces projection vectors z₁ = f(x₁) and z₂ = f(x₂).
Here is where the asymmetry enters. One branch passes through an additional prediction MLP h, producing p₁ = h(z₁). The other branch's output z₂ is treated as a fixed target via Stop Gradient. The model maximizes the between p₁ and stopgrad(z₂).
The full loss symmetrizes this: both branches take turns being the predictor and the target. But crucially, the Stop Gradient is always applied to the target side.
Simplified to show the idea — not the real implementation.
# f: backbone + projection MLP
# h: prediction MLP
for x in loader: # load a minibatch
x1, x2 = aug(x), aug(x) # two random augmentations
z1, z2 = f(x1), f(x2) # projections (n × d)
p1, p2 = h(z1), h(z2) # predictions (n × d)
L = D(p1, z2)/2 + D(p2, z1)/2 # symmetrized loss
L.backward() # backprop
update(f, h) # SGD update
def D(p, z):
z = z.detach() # ⬅ Stop Gradient: treat z as constant
p = normalize(p, dim=1) # L2-normalize
z = normalize(z, dim=1) # L2-normalize
return -(p * z).sum(dim=1).mean()The key discovery: Stop Gradient prevents collapse
The paper's most important empirical finding is clean and dramatic. Take the SimSiam architecture. Keep everything — the encoder, the projection MLP, the prediction MLP, , the optimizer, every hyperparameter — exactly the same. Change one thing: remove the Stop Gradient. The result? Instant collapse. The loss immediately drops to its minimum possible value of −1, the output standard deviation drops to zero (meaning all outputs are identical), and the kNN accuracy stays at chance level (0.1% on ImageNet).
Put the Stop Gradient back? The loss converges normally, the output standard deviation stays near the theoretically expected value of 1/√d (where d = 2048), the kNN accuracy steadily improves, and the final linear evaluation reaches 67.7%.
This single experiment is the paper's strongest argument. It proves that architecture alone — the predictor, batch normalization, L2 normalization — is not sufficient to prevent collapse. Stop Gradient is the essential ingredient.
Ablation studies: what matters and what does not
The paper systematically tests every component to determine which ones are essential for preventing collapse and which merely affect accuracy.
Prediction MLP. Removing the prediction MLP h entirely causes collapse — but this is a mathematical artifact. With the symmetrized loss and no predictor, the gradient of the Stop Gradient version is identical (up to a scale of ½) to the gradient without Stop Gradient. The predictor breaks this symmetry and makes Stop Gradient meaningful. Interestingly, using a fixed (no decay) for h produces slightly better results, suggesting h should adapt to the latest representations rather than converge prematurely.
. SimSiam works remarkably well across batch sizes from 64 to 4096 — a sharp contrast to SimCLR and SwAV, which need batches of 4096 to work well. Even a batch size of 128 drops accuracy by only 0.8%. This is a major practical advantage: SimSiam is friendly to standard 8-GPU setups with a batch size of 256 or 512.
Batch normalization. Removing all BN from the MLP heads does not cause collapse (the accuracy is low at 34.6% but non-trivial). Adding BN to hidden layers helps optimization, bringing accuracy to 67.4%. Adding BN to the projection MLP's output further boosts to 68.1%. But adding BN to the prediction MLP's output causes training instability. The conclusion: BN helps optimization but does not prevent collapse.
Similarity function. Replacing cosine similarity with cross-entropy similarity still produces reasonable results (63.2% vs 68.1%). This shows collapse prevention is not tied to the specific similarity measure.
The hypothesis: SimSiam as alternating optimization
Why does Stop Gradient work? The paper proposes an elegant hypothesis: SimSiam implicitly solves an Expectation-Maximization (EM) style problem with two sets of variables.
Consider a with two arguments — the network parameters θ and an auxiliary set of variables η, where ηₓ represents the "ideal" representation of image x. The optimization problem minimizes the expected distance between the network's output for augmented views and these ideal representations.
Like , this problem can be solved by alternating: fix η, update θ (a network training step); then fix θ, update η (assigning each image its average representation over augmentations). The Stop Gradient is a natural consequence: when solving for θ, the targets η are constants, so no gradient flows through them.
SimSiam approximates this alternation in a single step: it samples one augmentation to estimate η (instead of computing the full expectation), and takes one step for θ. The predictor MLP h helps fill the gap left by this approximation — it can learn to predict the expected representation that a single sample cannot provide.
SimSiam as a unifying hub
One of SimSiam's deepest contributions is conceptual: it reveals that the leading self-supervised methods are all variations of the same underlying Siamese architecture, each with one extra ingredient. SimSiam is the stripped-down core.
SimCLR = SimSiam + negative pairs. SimCLR prevents collapse by repulsing different images. Adding the predictor or Stop Gradient to SimCLR neither helps nor hurts — the contrastive mechanism already solves a different optimization problem.
BYOL = SimSiam + momentum encoder. BYOL avoids collapse using a slowly-updating copy of the encoder as the target. SimSiam shows this momentum is helpful for accuracy but not essential for preventing collapse.
SwAV = SimSiam + online clustering. SwAV replaces direct similarity with cluster assignment via the Sinkhorn-Knopp transform. Removing Stop Gradient from SwAV causes divergence — consistent with clustering being an inherently alternating formulation.
MoCo = SimSiam + momentum encoder + negative queue. MoCo combines a momentum encoder with a queue of negatives. SimSiam removes both.
Results: competitive despite simplicity
On ImageNet linear evaluation with a ResNet-50 backbone, SimSiam achieves 68.1% with 100- pre-training — the highest among all methods at this duration. With 200 epochs it reaches 70.0%, and with 800 epochs, 71.3%. While BYOL achieves higher accuracy at longer schedules (74.3% at 800 epochs), SimSiam accomplishes this using only a batch size of 256, standard SGD (no LARS optimizer), and no momentum encoder.
Perhaps more importantly, SimSiam's representations transfer well to downstream tasks. On PASCAL VOC detection and COCO detection and , SimSiam matches or exceeds the other methods, and all methods surpass or match ImageNet supervised pre-training. The "optimal" SimSiam recipe (with adjusted learning rate and weight decay) achieves the best transfer results, suggesting the learned representations capture genuinely useful visual features.
The authors emphasize a deeper point: despite the many differences between SimCLR, MoCo, BYOL, SwAV, and SimSiam, all are highly successful for transfer learning. The common factor is the Siamese architecture itself, suggesting it is the core reason for their success.
The bigger picture: Siamese networks as an inductive bias
The paper closes with a thought-provoking observation. Convolutions succeeded because they encode a powerful : via . A filter that detects edges works the same way everywhere in the image.
Siamese networks encode a similarly powerful bias: transformation via weight sharing. By passing two augmented views through the same encoder, the network is structurally biased to produce similar outputs for inputs that differ only by augmentation — which is precisely the definition of invariance. Just as convolutions hardwire the assumption that "a pattern is the same regardless of position," Siamese networks hardwire the assumption that "an image is the same regardless of crop, color shift, or blur."
SimSiam's simplicity makes this insight sharply visible. When you strip away negatives, momentum, and clustering, what remains is a pure invariance-learning machine. The Siamese architecture is not just a convenient scaffold — it is the fundamental mechanism.
Timeline: the road to SimSiam
2006
Contrastive learning foundations
Hadsell, Chopra, and LeCun introduced contrastive loss for learning invariant mappings. The idea of attracting similar pairs and repulsing dissimilar ones laid the groundwork.
2018
Instance discrimination
Wu et al. proposed treating every image as its own class with a memory bank of representations, establishing instance-level contrastive learning.
2020
SimCLR — simple contrastive framework
Chen et al. showed that a simple framework with strong augmentations, a projection head, and large batches of negatives is highly effective.
2020
MoCo v2 — momentum contrastive learning
He et al. used a momentum encoder and negative queue, enabling effective contrastive learning with small batches.
2020
BYOL — no negatives needed
Grill et al. eliminated negative pairs entirely, using only a momentum encoder for stable targets. But what exactly prevented collapse remained unclear.
2020
SwAV — clustering meets Siamese networks
Caron et al. introduced online clustering via Sinkhorn-Knopp into the Siamese framework, avoiding explicit negatives through balanced partition constraints.
2021
SimSiam — the minimal Siamese baseline
Chen and He stripped everything away: no negatives, no momentum encoder, no clustering. Just a shared encoder, a predictor, and Stop Gradient. Proved the Siamese architecture itself is the core mechanism.
2021
Barlow Twins — redundancy reduction
Zbontar et al. continued the non-contrastive direction, preventing collapse by decorrelating embedding dimensions rather than using Stop Gradient — another simplification inspired by SimSiam's philosophy.
CitationChen, He. Exploring Simple Siamese Representation Learning. CVPR, 2021.
Terms in this paper
- Siamese Networkالشبكة السيامية
- Stop Gradientإيقاف التدرّج
- Representation Collapseانهيار التمثيلات
- Contrastive Learningالتعلم التبايُني
- Cosine Similarityتشابه جيب التمام
- Self-Supervised Learningالتعلم ذاتي الإشراف
- Data Augmentationتعزيز البيانات
- Momentum Encoderمُرمِّز الزخم
- Negative Sampleالعيّنة السلبية
- Positive Pairالزوج الإيجابي
- Projection Headرأس الإسقاط
- Backboneالبنية الأساسية
- Batch Normalizationتسوية الدفعات الحسابية
- Transfer Learningنقل التعلم
- Invarianceالثبات