Representation Learning2021intermediate12 min read

Exploring Simple Siamese Representation Learning

استكشاف تعلُّم التمثيلات بشبكات سيامية بسيطة

Chen, X. · He, K. — CVPR

The problem

By 2020, self-supervised visual representation learning had made remarkable progress through methods like SimCLR, MoCo, SwAV, and BYOL. But each relied on a specific trick to avoid — the failure mode where a network outputs the same constant vector for every input. SimCLR needed negative pairs and large batches. MoCo needed a and a queue. SwAV needed online . BYOL needed a momentum . The question remained: which of these components is truly essential, and which are incidental?

The contribution

SimSiam: a minimalist that learns meaningful representations without negative pairs, without large batches, and without a momentum encoder. The architecture is simply a shared encoder with a prediction on one branch and a on the other. The paper demonstrates that the Stop Gradient operation alone is the essential ingredient for preventing collapse, and provides a hypothesis that SimSiam implicitly solves an Expectation-Maximization-like alternating optimization problem. SimSiam achieves 71.3% top-1 accuracy on ImageNet linear evaluation (800 epochs) and competitive results on detection and segmentation.

The impact

SimSiam stripped to its conceptual bones, proving that the Siamese architecture itself — not the bells and whistles around it — is the core reason these methods work. It served as a unifying hub connecting SimCLR, BYOL, SwAV, and MoCo: each can be described as "SimSiam plus one extra component." This reframing shifted the field's understanding and influenced subsequent work like Barlow Twins, VICReg, and other non-contrastive methods that continued simplifying the recipe.

Imagine a sculptor and a photographer both studying the same statue. The sculptor makes a clay copy; the photographer takes a snapshot. A judge compares the two and says "make them more alike." But here's the twist: only the sculptor gets to hear the feedback. The photographer's snapshot is frozen — treated as a fixed target.

Without this one-way feedback rule, both artists would take the laziest shortcut: they'd both produce a featureless grey blob, identical for every statue. That blob is "collapse."

SimSiam discovered that this asymmetric feedback — one side learns, the other stays still — is all you need to prevent collapse. No need for negative examples, no need for thousands of different statues in each batch, no need for a slowly-updating "shadow photographer." Just the one-way rule.

The enemy: representation collapse

The dream of self-supervised learning is simple: teach a network to understand images without any human labels. The strategy? Show the network two augmented views of the same image — cropped differently, color-jittered, maybe blurred — and train it to produce similar representations for both. If it can match views of the same image, it must be learning something meaningful about the image's content.

But there is a devastating shortcut. The network could learn to output the exact same vector for every input. Two views of a cat? Same vector. Two views of a car? Same vector. The similarity is perfect — and the representation is perfectly useless. This failure mode is called representation collapse, and every self-supervised method before SimSiam needed a specific mechanism to prevent it.

Open in Lab
Drag the slider to see: without Stop Gradient, all outputs collapse to a constant. With Stop Gradient, representations spread across the space.
The demo wakes as you arrive…

Before SimSiam, the field had developed four distinct anti-collapse strategies, each adding complexity. SimCLR pushed apart representations of different images (negative pairs), requiring enormous batches of 4096 or more to provide enough negatives. MoCo maintained a queue of negative representations using a momentum encoder — a slowly-updating copy of the main network. SwAV replaced negatives with online clustering using the Sinkhorn-Knopp transform. BYOL eliminated negatives entirely but relied on a momentum encoder to provide stable targets. Each method worked, but each came with significant overhead. The question SimSiam asked was radical: what if we remove all of these?

SimSiam architecture: radical simplicity

SimSiam's architecture is strikingly minimal. Start with an image x. Create two augmented views x₁ and x₂ using random cropping, color jittering, horizontal flipping, and Gaussian blurring. Both views pass through the same encoder f — a ResNet-50 followed by a 3-layer projection MLP. This produces projection vectors z₁ = f(x₁) and z₂ = f(x₂).

Here is where the asymmetry enters. One branch passes through an additional prediction MLP h, producing p₁ = h(z₁). The other branch's output z₂ is treated as a fixed target via Stop Gradient. The model maximizes the between p₁ and stopgrad(z₂).

The full loss symmetrizes this: both branches take turns being the predictor and the target. But crucially, the Stop Gradient is always applied to the target side.

Open in Lab
Explore the SimSiam architecture. Click components to see their role. Note the asymmetric gradient flow.
The demo wakes as you arrive…
D(p1,z2)=−p1∥p1∥2⋅z2∥z2∥2\mathcal{D}(p_1, z_2) = -\frac{p_1}{\|p_1\|_2} \cdot \frac{z_2}{\|z_2\|_2}
Negative cosine similarity — The loss for one direction: normalize both vectors to unit length, compute their dot product, and negate it. Minimizing this loss pushes the prediction p₁ to align with the target z₂. The normalization to unit vectors means the loss only cares about direction, not magnitude.
L=12 D ⁣(p1,  stopgrad(z2))+12 D ⁣(p2,  stopgrad(z1))\mathcal{L} = \frac{1}{2}\,\mathcal{D}\!\bigl(p_1,\;\text{stopgrad}(z_2)\bigr) + \frac{1}{2}\,\mathcal{D}\!\bigl(p_2,\;\text{stopgrad}(z_1)\bigr)
Symmetrized SimSiam loss — The full loss averages two terms: branch 1 predicts branch 2's frozen target, and vice versa. Symmetrization doubles the predictions per image but is not required for collapse prevention — it simply improves accuracy.
SimSiam pseudocode (PyTorch-like)python

Simplified to show the idea — not the real implementation.

# f: backbone + projection MLP
# h: prediction MLP
for x in loader:                # load a minibatch
    x1, x2 = aug(x), aug(x)    # two random augmentations
    z1, z2 = f(x1), f(x2)      # projections (n × d)
    p1, p2 = h(z1), h(z2)      # predictions (n × d)
    L = D(p1, z2)/2 + D(p2, z1)/2  # symmetrized loss
    L.backward()               # backprop
    update(f, h)               # SGD update

def D(p, z):
    z = z.detach()             # ⬅ Stop Gradient: treat z as constant
    p = normalize(p, dim=1)    # L2-normalize
    z = normalize(z, dim=1)    # L2-normalize
    return -(p * z).sum(dim=1).mean()

The key discovery: Stop Gradient prevents collapse

The paper's most important empirical finding is clean and dramatic. Take the SimSiam architecture. Keep everything — the encoder, the projection MLP, the prediction MLP, , the optimizer, every hyperparameter — exactly the same. Change one thing: remove the Stop Gradient. The result? Instant collapse. The loss immediately drops to its minimum possible value of −1, the output standard deviation drops to zero (meaning all outputs are identical), and the kNN accuracy stays at chance level (0.1% on ImageNet).

Put the Stop Gradient back? The loss converges normally, the output standard deviation stays near the theoretically expected value of 1/√d (where d = 2048), the kNN accuracy steadily improves, and the final linear evaluation reaches 67.7%.

This single experiment is the paper's strongest argument. It proves that architecture alone — the predictor, batch normalization, L2 normalization — is not sufficient to prevent collapse. Stop Gradient is the essential ingredient.

Open in Lab
Toggle Stop Gradient on/off to see how gradient flow changes. Without it, both branches update together and collapse to a trivial solution.
The demo wakes as you arrive…

Ablation studies: what matters and what does not

The paper systematically tests every component to determine which ones are essential for preventing collapse and which merely affect accuracy.

Prediction MLP. Removing the prediction MLP h entirely causes collapse — but this is a mathematical artifact. With the symmetrized loss and no predictor, the gradient of the Stop Gradient version is identical (up to a scale of ½) to the gradient without Stop Gradient. The predictor breaks this symmetry and makes Stop Gradient meaningful. Interestingly, using a fixed (no decay) for h produces slightly better results, suggesting h should adapt to the latest representations rather than converge prematurely.

. SimSiam works remarkably well across batch sizes from 64 to 4096 — a sharp contrast to SimCLR and SwAV, which need batches of 4096 to work well. Even a batch size of 128 drops accuracy by only 0.8%. This is a major practical advantage: SimSiam is friendly to standard 8-GPU setups with a batch size of 256 or 512.

Batch normalization. Removing all BN from the MLP heads does not cause collapse (the accuracy is low at 34.6% but non-trivial). Adding BN to hidden layers helps optimization, bringing accuracy to 67.4%. Adding BN to the projection MLP's output further boosts to 68.1%. But adding BN to the prediction MLP's output causes training instability. The conclusion: BN helps optimization but does not prevent collapse.

Similarity function. Replacing cosine similarity with cross-entropy similarity still produces reasonable results (63.2% vs 68.1%). This shows collapse prevention is not tied to the specific similarity measure.

Open in Lab
Explore ablation results. Toggle components to see their effect on accuracy and collapse behavior.
The demo wakes as you arrive…

The hypothesis: SimSiam as alternating optimization

Why does Stop Gradient work? The paper proposes an elegant hypothesis: SimSiam implicitly solves an Expectation-Maximization (EM) style problem with two sets of variables.

Consider a with two arguments — the network parameters θ and an auxiliary set of variables η, where ηₓ represents the "ideal" representation of image x. The optimization problem minimizes the expected distance between the network's output for augmented views and these ideal representations.

Like , this problem can be solved by alternating: fix η, update θ (a network training step); then fix θ, update η (assigning each image its average representation over augmentations). The Stop Gradient is a natural consequence: when solving for θ, the targets η are constants, so no gradient flows through them.

SimSiam approximates this alternation in a single step: it samples one augmentation to estimate η (instead of computing the full expectation), and takes one step for θ. The predictor MLP h helps fill the gap left by this approximation — it can learn to predict the expected representation that a single sample cannot provide.

L(θ,η)=Ex,T[∥Fθ(T(x))−ηx∥22]\mathcal{L}(\theta, \eta) = \mathbb{E}_{x, \mathcal{T}} \bigl[\|F_\theta(\mathcal{T}(x)) - \eta_x\|_2^2\bigr]
Hypothesized implicit objective — The hypothesized loss has two sets of variables: θ (network parameters) and η (per-image target representations). F is the encoder, T is a random augmentation, x is an image. Minimizing alternately over θ and η yields an EM-like algorithm where Stop Gradient emerges naturally.
ηxt  ←  ET ⁣[Fθt ⁣(T(x))]\eta_x^t \;\leftarrow\; \mathbb{E}_{\mathcal{T}}\!\bigl[F_{\theta^t}\!(\mathcal{T}(x))\bigr]
Solving for η — the target update — When solving for η with θ fixed, the optimal ηₓ is simply the average representation of image x over all possible augmentations. In practice, SimSiam approximates this with a single augmentation sample, and the predictor h helps estimate this expectation.

SimSiam as a unifying hub

One of SimSiam's deepest contributions is conceptual: it reveals that the leading self-supervised methods are all variations of the same underlying Siamese architecture, each with one extra ingredient. SimSiam is the stripped-down core.

SimCLR = SimSiam + negative pairs. SimCLR prevents collapse by repulsing different images. Adding the predictor or Stop Gradient to SimCLR neither helps nor hurts — the contrastive mechanism already solves a different optimization problem.

BYOL = SimSiam + momentum encoder. BYOL avoids collapse using a slowly-updating copy of the encoder as the target. SimSiam shows this momentum is helpful for accuracy but not essential for preventing collapse.

SwAV = SimSiam + online clustering. SwAV replaces direct similarity with cluster assignment via the Sinkhorn-Knopp transform. Removing Stop Gradient from SwAV causes divergence — consistent with clustering being an inherently alternating formulation.

MoCo = SimSiam + momentum encoder + negative queue. MoCo combines a momentum encoder with a queue of negatives. SimSiam removes both.

Open in Lab
Click each method to see what SimSiam adds or removes. SimSiam is the minimal core.
The demo wakes as you arrive…

Results: competitive despite simplicity

On ImageNet linear evaluation with a ResNet-50 backbone, SimSiam achieves 68.1% with 100- pre-training — the highest among all methods at this duration. With 200 epochs it reaches 70.0%, and with 800 epochs, 71.3%. While BYOL achieves higher accuracy at longer schedules (74.3% at 800 epochs), SimSiam accomplishes this using only a batch size of 256, standard SGD (no LARS optimizer), and no momentum encoder.

Perhaps more importantly, SimSiam's representations transfer well to downstream tasks. On PASCAL VOC detection and COCO detection and , SimSiam matches or exceeds the other methods, and all methods surpass or match ImageNet supervised pre-training. The "optimal" SimSiam recipe (with adjusted learning rate and weight decay) achieves the best transfer results, suggesting the learned representations capture genuinely useful visual features.

The authors emphasize a deeper point: despite the many differences between SimCLR, MoCo, BYOL, SwAV, and SimSiam, all are highly successful for transfer learning. The common factor is the Siamese architecture itself, suggesting it is the core reason for their success.

The bigger picture: Siamese networks as an inductive bias

The paper closes with a thought-provoking observation. Convolutions succeeded because they encode a powerful : via . A filter that detects edges works the same way everywhere in the image.

Siamese networks encode a similarly powerful bias: transformation via weight sharing. By passing two augmented views through the same encoder, the network is structurally biased to produce similar outputs for inputs that differ only by augmentation — which is precisely the definition of invariance. Just as convolutions hardwire the assumption that "a pattern is the same regardless of position," Siamese networks hardwire the assumption that "an image is the same regardless of crop, color shift, or blur."

SimSiam's simplicity makes this insight sharply visible. When you strip away negatives, momentum, and clustering, what remains is a pure invariance-learning machine. The Siamese architecture is not just a convenient scaffold — it is the fundamental mechanism.

Timeline: the road to SimSiam

  1. 2006

    Contrastive learning foundations

    Hadsell, Chopra, and LeCun introduced contrastive loss for learning invariant mappings. The idea of attracting similar pairs and repulsing dissimilar ones laid the groundwork.

  2. 2018

    Instance discrimination

    Wu et al. proposed treating every image as its own class with a memory bank of representations, establishing instance-level contrastive learning.

  3. 2020

    SimCLR — simple contrastive framework

    Chen et al. showed that a simple framework with strong augmentations, a projection head, and large batches of negatives is highly effective.

  4. 2020

    MoCo v2 — momentum contrastive learning

    He et al. used a momentum encoder and negative queue, enabling effective contrastive learning with small batches.

  5. 2020

    BYOL — no negatives needed

    Grill et al. eliminated negative pairs entirely, using only a momentum encoder for stable targets. But what exactly prevented collapse remained unclear.

  6. 2020

    SwAV — clustering meets Siamese networks

    Caron et al. introduced online clustering via Sinkhorn-Knopp into the Siamese framework, avoiding explicit negatives through balanced partition constraints.

  7. 2021

    SimSiam — the minimal Siamese baseline

    Chen and He stripped everything away: no negatives, no momentum encoder, no clustering. Just a shared encoder, a predictor, and Stop Gradient. Proved the Siamese architecture itself is the core mechanism.

  8. 2021

    Barlow Twins — redundancy reduction

    Zbontar et al. continued the non-contrastive direction, preventing collapse by decorrelating embedding dimensions rather than using Stop Gradient — another simplification inspired by SimSiam's philosophy.

CitationChen, He. Exploring Simple Siamese Representation Learning. CVPR, 2021.

Terms in this paper