Computer Vision2020intermediate11 min read
Momentum Contrast for Unsupervised Visual Representation Learning
تباين الزخم للتعلّم التمثيلي البصري غير الخاضع للإشراف
He, K. · Fan, H. · Wu, Y. · Xie, S. · Girshick, R. — CVPR
The problem
Unsupervised representation learning had triumphed in NLP (word2vec, BERT), but in computer vision, supervised ImageNet pre-training remained king. methods showed promise but faced a fundamental tension: they need a large and consistent set of negative samples to work well. End-to-end approaches (like SimCLR) require enormous batch sizes — often 4096 or 8192 — because the negatives come only from the current mini-batch. Memory-bank approaches store representations from past epochs, but these go stale quickly as the updates, creating inconsistency. Neither scales gracefully.
The contribution
MoCo re-frames contrastive learning as : a query encoder, a -updated key encoder, and a of encoded keys that decouples dictionary size from . The key encoder's weights are an of the query encoder (momentum = 0.999), keeping keys consistent even as training progresses. The queue stores 65,536 keys and discards the oldest ones — far more negatives than any mini-batch can hold. This simple design matches or exceeds supervised pre-training on 7 downstream detection and segmentation tasks, closing the gap between unsupervised and supervised visual representation learning.
The impact
MoCo proved that unsupervised visual features can transfer as well as — or better than — supervised ones for detection and segmentation. It introduced the , an idea adopted by BYOL, SimSiam, and DINO. Its queue mechanism inspired efficient negative sampling across the field. MoCo v2 and v3 extended the idea to ViTs. Together with SimCLR, MoCo launched the self-supervised vision revolution that made labels optional for pre-training.
Imagine a detective learning to identify suspects without ever being told their names. She has one tool: a book of mugshots she carries (the queue). When a new suspect walks in, she takes two photos from different angles. She hands one photo to her sharp-eyed partner (the query encoder) and another to a calmer, more consistent partner (the momentum encoder) who updates his method slowly so his descriptions stay stable. The sharp-eyed partner asks: "which mugshot in the book matches this person?" — and learns from each correct and incorrect match. Every day, the newest mugshot goes into the book and the oldest is thrown out. Over time, both partners become excellent at recognizing people — all without anyone ever telling them a single name.
The gap: vision lacked its BERT moment
By 2019, NLP had crossed a threshold. Models like BERT learned powerful representations from unlabeled text, then transferred them to downstream tasks with minimal . Vision, however, was stuck: the standard recipe was still "pre-train on ImageNet with human-annotated labels, then fine-tune." Labeling millions of images is expensive, slow, and limited to pre-defined categories.
Contrastive learning offered a path out — the idea is simple: pull representations of similar things together and push dissimilar things apart in an space. But two practical problems kept it from matching supervised pre-training:
- Too few negatives. End-to-end methods like SimCLR use only the current mini-batch as negatives. To get enough, you need thousands of samples per batch — requiring many GPUs and enormous memory.
- Stale negatives. Memory-bank approaches store past representations, but those were encoded by older versions of the network. As the encoder changes, old keys become inconsistent, and the model trains on contradictory signals.
The MoCo idea: a dynamic dictionary with a momentum encoder
MoCo reframes contrastive learning as dictionary look-up. Think of it as building a search engine for images:
You have a query (a new image representation) and a dictionary of keys (representations of other images). You want the query to match the one key that represents the same image (just augmented differently) and to not match any other key. The better the dictionary — larger, more consistent — the better the learned representations.
MoCo achieves this with three components:
- A query encoder — trained normally with .
- A momentum encoder — its weights are not trained by gradients, but are an exponential moving average of the query encoder's weights.
- A queue of encoded keys — a FIFO buffer that stores keys from recent mini-batches, decoupling dictionary size from batch size.
The momentum encoder: slow and steady wins the race
Here is the crucial insight: if you update the key encoder at the same rate as the query encoder (as in end-to-end training), the keys in your dictionary become inconsistent — early keys were made by a very different encoder than recent keys. This is like having a dictionary where half the entries were written in Old English and half in modern English.
MoCo's solution: update the key encoder very slowly using exponential moving average. After each training step, the key encoder's parameters become a weighted blend of the old key-encoder parameters and the new query-encoder parameters :
The queue: a rolling memory of negatives
Contrastive learning needs negatives — lots of them. More negatives mean a harder discrimination task, which forces the encoder to learn better features. But where do you get 65,536 negatives when your batch size is only 256?
MoCo's answer: a queue. Each mini-batch produces encoded keys from the momentum encoder. These keys are enqueued. The oldest keys are dequeued. The queue size is independent of batch size — you can have a dictionary of 65,536 keys while training with a batch of 256 on a single GPU.
This is like a security camera archive: you keep the last few days of footage (the queue), and when new footage arrives, you delete the oldest. At any moment you can search through far more footage than what is currently recording.
The InfoNCE loss: pick the matching key
With the architecture set, the model needs a training signal. MoCo uses — a form of contrastive loss that looks like a classification problem. Given a query , one positive key (the same image, differently augmented), and negative keys from the queue, the model must identify the positive. The loss is simply a cross-entropy over the similarity scores:
Reading the formula intuitively: the loss says "make the between the query and its matching key as large as possible, while making the dot products with all non-matching keys as small as possible." The controls how peaked the distribution is — a lower temperature (0.07) makes the model more decisive, demanding a clear winner.
This is exactly a -way softmax classification, where the correct "class" is the positive key among negatives. With , the model must pick the right image from 65,537 candidates — a very hard discrimination task that demands rich, informative features.
The complete training loop
Now let's assemble the full pipeline. For each mini-batch:
Step 1 — Augment. Take each image and create two random augmentations: for the query encoder and for the key encoder. Augmentations include random cropping, color jittering, horizontal flipping, and grayscale conversion.
Step 2 — Encode. Pass through the query encoder to get query . Pass through the momentum encoder to get key . The key encoder runs with torch.no_grad() — no gradients flow through it.
Step 3 — Contrast. Compute the InfoNCE loss: should match and not match any of the keys from the queue.
Step 4 — Update query encoder. Backpropagate through the loss and update with SGD.
Step 5 — Update momentum encoder. Apply the momentum update: .
Step 6 — Update queue. Enqueue the new keys from this batch, dequeue the oldest keys.
The same idea in code
Simplified to show the idea — not the real implementation.
# f_q: query encoder f_k: momentum encoder (same arch)
# queue: tensor of shape [K, C] K = 65536, C = 128
# m: momentum = 0.999 tau: temperature = 0.07
for x in loader: # each image
x_q = augment(x) # random crop + color jitter + ...
x_k = augment(x) # different random augmentation
q = f_q(x_q) # query: [N, C], gradient flows
q = normalize(q, dim=1)
with torch.no_grad():
k = f_k(x_k) # key: [N, C], no gradient
k = normalize(k, dim=1)
# Positive logits: Nx1
l_pos = torch.bmm(q.unsqueeze(1), k.unsqueeze(2)).squeeze(-1)
# Negative logits: NxK (queue holds K keys from past batches)
l_neg = q @ queue.clone().T
logits = torch.cat([l_pos, l_neg], dim=1) / tau # Nx(1+K)
labels = torch.zeros(N, dtype=torch.long) # positive is index 0
loss = F.cross_entropy(logits, labels)
loss.backward() # update f_q only
optimizer.step()
# Momentum update f_k
for p_q, p_k in zip(f_q.parameters(), f_k.parameters()):
p_k.data = m * p_k.data + (1 - m) * p_q.data
# Update queue (FIFO)
queue = torch.cat([k, queue[:K - N]], dim=0)Results: closing the gap
MoCo was evaluated in two ways: linear evaluation (freeze the encoder, train only a linear classifier on top) and (fine-tune on downstream tasks like and on PASCAL VOC and COCO).
On ImageNet linear evaluation with a ResNet-50 backbone, MoCo achieves 60.6% top-1 accuracy — competitive with SimCLR's result (which required 8× the batch size).
The real surprise came from transfer learning. On PASCAL VOC object detection, MoCo pre-training surpassed ImageNet supervised pre-training by +0.5 AP. On COCO detection and segmentation, MoCo matched or exceeded supervised pre-training across all metrics. This was the first convincing evidence that unsupervised visual pre-training can be better than supervised for real-world tasks.
Ablations: what actually matters?
The authors systematically ablated three key choices:
- Queue size (K). Increasing the queue from 1024 to 65,536 steadily improves accuracy. More negatives = harder task = better features. Beyond 65K, gains plateau.
- Momentum (m). The sweet spot is . Too low (0.9): inconsistent keys destroy training. Too high (0.9999): the key encoder barely moves, losing adaptability.
- Training schedule. 200 epochs work much better than 100. The slow momentum update means the key encoder needs longer to converge.
Legacy: what MoCo sparked
MoCo's ideas rippled through the field in ways its authors could not have foreseen:
BYOL asked: what if we don't need negatives at all? It kept the momentum encoder but removed the queue, showing that the asymmetry between query and key encoders alone prevents collapse.
SimCLR took the opposite path — no momentum encoder, no queue, just massive batches. The tension between MoCo's efficiency and SimCLR's simplicity drove rapid progress in the field.
MoCo v2 (2020) adopted SimCLR's stronger augmentations and MLP projection head, boosting ImageNet accuracy to 71.1%. MoCo v3 (2021) extended the framework to Vision Transformers. DINO, SwAV, VICReg, and Barlow Twins all built on insights from MoCo.
2018
InstDisc (Instance Discrimination)
Wu et al. proposed treating each image as its own class — the first memory bank approach. Showed the idea works but keys became stale as the encoder updated.
2018
CPC (Contrastive Predictive Coding)
Van den Oord et al. introduced InfoNCE loss and the idea of predicting future representations contrastively. Worked across audio, text, and images.
2020
MoCo (this paper)
Momentum encoder + queue = large consistent dictionary. First to match supervised pre-training on detection and segmentation transfers.
2020
SimCLR
Chen et al. showed that simple contrastive learning works amazingly well — but needs batch sizes of 4096+. Introduced stronger augmentations and MLP projection head.
2020
BYOL
Grill et al. removed negatives entirely. Kept the momentum encoder and added a predictor head to the query branch. Proved that asymmetry prevents collapse.
2021
DINO
Caron et al. combined self-distillation with a momentum teacher on Vision Transformers. Discovered that self-supervised ViTs learn semantic segmentation spontaneously.
Why it matters
CitationHe, Fan, Wu, Xie, Girshick. Momentum Contrast for Unsupervised Visual Representation Learning. CVPR, 2020.
Terms in this paper
- Contrastive Learningالتعلم التبايُني
- Momentum Encoderمُرمِّز الزخم
- Queueالطابور
- InfoNCE Lossخسارة InfoNCE
- Dictionary Look-upالبحث في القاموس
- Exponential Moving Averageالمتوسط المتحرك الأُسّي
- FIFOالوارد أولاً يخرج أولاً
- Linear Probingالاختبار الخطي