Computer Vision2020intermediate11 min read

Momentum Contrast for Unsupervised Visual Representation Learning

تباين الزخم للتعلّم التمثيلي البصري غير الخاضع للإشراف

He, K. · Fan, H. · Wu, Y. · Xie, S. · Girshick, R. — CVPR

The problem

Unsupervised representation learning had triumphed in NLP (word2vec, BERT), but in computer vision, supervised ImageNet pre-training remained king. methods showed promise but faced a fundamental tension: they need a large and consistent set of negative samples to work well. End-to-end approaches (like SimCLR) require enormous batch sizes — often 4096 or 8192 — because the negatives come only from the current mini-batch. Memory-bank approaches store representations from past epochs, but these go stale quickly as the updates, creating inconsistency. Neither scales gracefully.

The contribution

MoCo re-frames contrastive learning as : a query encoder, a -updated key encoder, and a of encoded keys that decouples dictionary size from . The key encoder's weights are an of the query encoder (momentum = 0.999), keeping keys consistent even as training progresses. The queue stores 65,536 keys and discards the oldest ones — far more negatives than any mini-batch can hold. This simple design matches or exceeds supervised pre-training on 7 downstream detection and segmentation tasks, closing the gap between unsupervised and supervised visual representation learning.

The impact

MoCo proved that unsupervised visual features can transfer as well as — or better than — supervised ones for detection and segmentation. It introduced the , an idea adopted by BYOL, SimSiam, and DINO. Its queue mechanism inspired efficient negative sampling across the field. MoCo v2 and v3 extended the idea to ViTs. Together with SimCLR, MoCo launched the self-supervised vision revolution that made labels optional for pre-training.

Imagine a detective learning to identify suspects without ever being told their names. She has one tool: a book of mugshots she carries (the queue). When a new suspect walks in, she takes two photos from different angles. She hands one photo to her sharp-eyed partner (the query encoder) and another to a calmer, more consistent partner (the momentum encoder) who updates his method slowly so his descriptions stay stable. The sharp-eyed partner asks: "which mugshot in the book matches this person?" — and learns from each correct and incorrect match. Every day, the newest mugshot goes into the book and the oldest is thrown out. Over time, both partners become excellent at recognizing people — all without anyone ever telling them a single name.

The gap: vision lacked its BERT moment

By 2019, NLP had crossed a threshold. Models like BERT learned powerful representations from unlabeled text, then transferred them to downstream tasks with minimal . Vision, however, was stuck: the standard recipe was still "pre-train on ImageNet with human-annotated labels, then fine-tune." Labeling millions of images is expensive, slow, and limited to pre-defined categories.

Contrastive learning offered a path out — the idea is simple: pull representations of similar things together and push dissimilar things apart in an space. But two practical problems kept it from matching supervised pre-training:

  • Too few negatives. End-to-end methods like SimCLR use only the current mini-batch as negatives. To get enough, you need thousands of samples per batch — requiring many GPUs and enormous memory.
  • Stale negatives. Memory-bank approaches store past representations, but those were encoded by older versions of the network. As the encoder changes, old keys become inconsistent, and the model trains on contradictory signals.
Open in Lab
Compare three contrastive approaches. Notice how MoCo combines a large dictionary with consistent keys.
The demo wakes as you arrive…

The MoCo idea: a dynamic dictionary with a momentum encoder

MoCo reframes contrastive learning as dictionary look-up. Think of it as building a search engine for images:

You have a query (a new image representation) and a dictionary of keys (representations of other images). You want the query to match the one key that represents the same image (just augmented differently) and to not match any other key. The better the dictionary — larger, more consistent — the better the learned representations.

MoCo achieves this with three components:

  • A query encoder fqf_q — trained normally with .
  • A momentum encoder fkf_k — its weights are not trained by gradients, but are an exponential moving average of the query encoder's weights.
  • A queue of encoded keys — a FIFO buffer that stores keys from recent mini-batches, decoupling dictionary size from batch size.
Open in Lab
Click each component to see its role. Drag the momentum slider to see how it affects key consistency.
The demo wakes as you arrive…

The momentum encoder: slow and steady wins the race

Here is the crucial insight: if you update the key encoder at the same rate as the query encoder (as in end-to-end training), the keys in your dictionary become inconsistent — early keys were made by a very different encoder than recent keys. This is like having a dictionary where half the entries were written in Old English and half in modern English.

MoCo's solution: update the key encoder very slowly using exponential moving average. After each training step, the key encoder's parameters θk\theta_k become a weighted blend of the old key-encoder parameters and the new query-encoder parameters θq\theta_q:

θk←m⋅θk+(1−m)⋅θq\theta_k \leftarrow m \cdot \theta_k + (1 - m) \cdot \theta_q
Momentum update — the heartbeat of MoCo — m = momentum coefficient (0.999). Each step, θ_k keeps 99.9% of its old self and absorbs only 0.1% from the query encoder. The key encoder evolves smoothly, never jumping — so all keys in the queue were made by nearly identical encoders.
Open in Lab
Drag the momentum slider between 0 and 1. Watch how high momentum keeps the key encoder evolving smoothly while low momentum causes wild jumps.
The demo wakes as you arrive…

The queue: a rolling memory of negatives

Contrastive learning needs negatives — lots of them. More negatives mean a harder discrimination task, which forces the encoder to learn better features. But where do you get 65,536 negatives when your batch size is only 256?

MoCo's answer: a queue. Each mini-batch produces encoded keys from the momentum encoder. These keys are enqueued. The oldest keys are dequeued. The queue size is independent of batch size — you can have a dictionary of 65,536 keys while training with a batch of 256 on a single GPU.

This is like a security camera archive: you keep the last few days of footage (the queue), and when new footage arrives, you delete the oldest. At any moment you can search through far more footage than what is currently recording.

Open in Lab
Watch how new keys enter the queue and old keys leave. The dictionary stays large and fresh.
The demo wakes as you arrive…

The InfoNCE loss: pick the matching key

With the architecture set, the model needs a training signal. MoCo uses — a form of contrastive loss that looks like a classification problem. Given a query qq, one positive key k+k^+ (the same image, differently augmented), and KK negative keys {k0−,k1−,…}\{k^-_0, k^-_1, \ldots\} from the queue, the model must identify the positive. The loss is simply a cross-entropy over the similarity scores:

Lq=−log⁡exp⁡(q⋅k+/τ)exp⁡(q⋅k+/τ)+∑i=0K−1exp⁡(q⋅ki−/τ)\mathcal{L}_q = -\log \frac{\exp(q \cdot k^+ / \tau)}{\exp(q \cdot k^+ / \tau) + \sum_{i=0}^{K-1} \exp(q \cdot k^-_i / \tau)}
InfoNCE loss — contrastive learning as classification — q·k = dot product similarity · τ = temperature (0.07) controls sharpness · The numerator rewards matching the positive key · The denominator penalizes matching any negative · This is log-softmax over the true positive's score

Reading the formula intuitively: the loss says "make the between the query and its matching key as large as possible, while making the dot products with all non-matching keys as small as possible." The τ\tau controls how peaked the distribution is — a lower temperature (0.07) makes the model more decisive, demanding a clear winner.

This is exactly a (K+1)(K+1)-way softmax classification, where the correct "class" is the positive key among KK negatives. With K=65,536K = 65,536, the model must pick the right image from 65,537 candidates — a very hard discrimination task that demands rich, informative features.

Open in Lab
Drag the query toward or away from the positive key. Watch how the loss changes as the similarity distribution sharpens.
The demo wakes as you arrive…

The complete training loop

Now let's assemble the full pipeline. For each mini-batch:

Step 1 — Augment. Take each image xx and create two random augmentations: xqx_q for the query encoder and xkx_k for the key encoder. Augmentations include random cropping, color jittering, horizontal flipping, and grayscale conversion.

Step 2 — Encode. Pass xqx_q through the query encoder fqf_q to get query qq. Pass xkx_k through the momentum encoder fkf_k to get key k+k^+. The key encoder runs with torch.no_grad() — no gradients flow through it.

Step 3 — Contrast. Compute the InfoNCE loss: qq should match k+k^+ and not match any of the KK keys from the queue.

Step 4 — Update query encoder. Backpropagate through the loss and update θq\theta_q with SGD.

Step 5 — Update momentum encoder. Apply the momentum update: θk←0.999⋅θk+0.001⋅θq\theta_k \leftarrow 0.999 \cdot \theta_k + 0.001 \cdot \theta_q.

Step 6 — Update queue. Enqueue the new keys k+k^+ from this batch, dequeue the oldest keys.

Open in Lab
Step through one complete MoCo training iteration.
The demo wakes as you arrive…

The same idea in code

MoCo pseudocode — the full algorithm in ~30 linespython

Simplified to show the idea — not the real implementation.

# f_q: query encoder    f_k: momentum encoder (same arch)
# queue: tensor of shape [K, C]   K = 65536, C = 128
# m: momentum = 0.999   tau: temperature = 0.07

for x in loader:                        # each image
    x_q = augment(x)                    # random crop + color jitter + ...
    x_k = augment(x)                    # different random augmentation

    q = f_q(x_q)                        # query: [N, C], gradient flows
    q = normalize(q, dim=1)

    with torch.no_grad():
        k = f_k(x_k)                    # key: [N, C], no gradient
        k = normalize(k, dim=1)

    # Positive logits: Nx1
    l_pos = torch.bmm(q.unsqueeze(1), k.unsqueeze(2)).squeeze(-1)
    # Negative logits: NxK  (queue holds K keys from past batches)
    l_neg = q @ queue.clone().T

    logits = torch.cat([l_pos, l_neg], dim=1) / tau   # Nx(1+K)
    labels = torch.zeros(N, dtype=torch.long)         # positive is index 0
    loss = F.cross_entropy(logits, labels)

    loss.backward()                     # update f_q only
    optimizer.step()

    # Momentum update f_k
    for p_q, p_k in zip(f_q.parameters(), f_k.parameters()):
        p_k.data = m * p_k.data + (1 - m) * p_q.data

    # Update queue (FIFO)
    queue = torch.cat([k, queue[:K - N]], dim=0)

Results: closing the gap

MoCo was evaluated in two ways: linear evaluation (freeze the encoder, train only a linear classifier on top) and (fine-tune on downstream tasks like and on PASCAL VOC and COCO).

On ImageNet linear evaluation with a ResNet-50 backbone, MoCo achieves 60.6% top-1 accuracy — competitive with SimCLR's result (which required 8× the batch size).

The real surprise came from transfer learning. On PASCAL VOC object detection, MoCo pre-training surpassed ImageNet supervised pre-training by +0.5 AP. On COCO detection and segmentation, MoCo matched or exceeded supervised pre-training across all metrics. This was the first convincing evidence that unsupervised visual pre-training can be better than supervised for real-world tasks.

Open in Lab
Compare MoCo against supervised pre-training across detection and segmentation benchmarks.
The demo wakes as you arrive…

Ablations: what actually matters?

The authors systematically ablated three key choices:

  • Queue size (K). Increasing the queue from 1024 to 65,536 steadily improves accuracy. More negatives = harder task = better features. Beyond 65K, gains plateau.
  • Momentum (m). The sweet spot is m=0.999m = 0.999. Too low (0.9): inconsistent keys destroy training. Too high (0.9999): the key encoder barely moves, losing adaptability.
  • Training schedule. 200 epochs work much better than 100. The slow momentum update means the key encoder needs longer to converge.

Legacy: what MoCo sparked

MoCo's ideas rippled through the field in ways its authors could not have foreseen:

BYOL asked: what if we don't need negatives at all? It kept the momentum encoder but removed the queue, showing that the asymmetry between query and key encoders alone prevents collapse.

SimCLR took the opposite path — no momentum encoder, no queue, just massive batches. The tension between MoCo's efficiency and SimCLR's simplicity drove rapid progress in the field.

MoCo v2 (2020) adopted SimCLR's stronger augmentations and MLP projection head, boosting ImageNet accuracy to 71.1%. MoCo v3 (2021) extended the framework to Vision Transformers. DINO, SwAV, VICReg, and Barlow Twins all built on insights from MoCo.

  1. 2018

    InstDisc (Instance Discrimination)

    Wu et al. proposed treating each image as its own class — the first memory bank approach. Showed the idea works but keys became stale as the encoder updated.

  2. 2018

    CPC (Contrastive Predictive Coding)

    Van den Oord et al. introduced InfoNCE loss and the idea of predicting future representations contrastively. Worked across audio, text, and images.

  3. 2020

    MoCo (this paper)

    Momentum encoder + queue = large consistent dictionary. First to match supervised pre-training on detection and segmentation transfers.

  4. 2020

    SimCLR

    Chen et al. showed that simple contrastive learning works amazingly well — but needs batch sizes of 4096+. Introduced stronger augmentations and MLP projection head.

  5. 2020

    BYOL

    Grill et al. removed negatives entirely. Kept the momentum encoder and added a predictor head to the query branch. Proved that asymmetry prevents collapse.

  6. 2021

    DINO

    Caron et al. combined self-distillation with a momentum teacher on Vision Transformers. Discovered that self-supervised ViTs learn semantic segmentation spontaneously.

Why it matters

CitationHe, Fan, Wu, Xie, Girshick. Momentum Contrast for Unsupervised Visual Representation Learning. CVPR, 2020.

Terms in this paper