Self-Supervised Learning2018intermediate11 min read
Representation Learning with Contrastive Predictive Coding
تعلُّم التمثيلات عبر الترميز التنبُّئي التبايُني
van den Oord, A. · Li, Y. · Vinyals, O. — arXiv
The problem
Unsupervised learning in 2018 relied on either generative models that reconstruct raw inputs pixel-by-pixel (expensive and focused on low-level detail) or purely discriminative approaches limited to single modalities. There was no universal framework that could learn high-level features from raw data — audio, images, text, or observations — using a single, principled objective without any labels.
The contribution
: a self-supervised framework that learns representations by predicting the future in . An compresses raw observations into embeddings; an summarizes past embeddings into a ; a linear predictor uses the context to identify the true future among negatives, trained with a new called InfoNCE. InfoNCE is proven to be a lower bound on between context and future — so maximizing it forces the representation to capture the slow, high-level structure that makes the future predictable. The same framework works across audio, vision, text, and RL.
The impact
CPC introduced InfoNCE — the loss function that became the backbone of modern . SimCLR, MoCo, wav2vec 2.0, and CLIP all descend from this idea. CPC proved that predicting the future in latent space, rather than reconstructing raw inputs, produces representations that transfer across tasks and modalities. It catalyzed the self-supervised revolution that made large-scale pre- without labels not just possible, but dominant.
Imagine you're at a concert with your eyes closed. After hearing the first few bars of a melody, you can guess the next notes — not perfectly, but well enough to tell the real continuation from a random snippet of a different song.
You don't need to reconstruct the exact waveform of the next note. You don't need to be a musician. You just need to tell "what actually follows" apart from "what doesn't belong." That contrast — real future vs. random impostor — is the entire training signal.
CPC works the same way: it listens to the past, builds a compressed summary, and learns to pick the true future out of a lineup — and the skill it acquires doing this turns out to be a deep understanding of the data's structure.
The problem: generative models waste capacity on noise
Before CPC, the dominant approach to unsupervised learning was generative modeling: train a model to reconstruct its input (like a deep autoencoder) or predict every pixel or sample. The problem is that most of the information in raw signals — the exact brightness of each pixel, the precise waveform of audio — is high-frequency noise that says nothing about the underlying meaning.
A model that tries to perfectly reconstruct a face image spends most of its capacity modeling skin texture and lighting, not learning that it's a face. A speech model that reconstructs waveforms spends capacity on room acoustics, not phonemes.
What we actually want is the high-level, slow-changing structure: the melody across an audio clip, the object identity across image patches, the topic across a paragraph. That structure is what makes downstream tasks — , retrieval, reasoning — possible.
The idea: predict the future in latent space
CPC's insight is a shift in what you predict and how you evaluate the prediction:
What: Don't predict raw pixels or waveforms. Instead, predict a compact representation of the future — an embedding that captures the high-level content while discarding noise.
How: Don't compare your prediction to the truth pixel-by-pixel. Instead, treat it as a classification problem: can you pick the true future embedding out of a lineup of random negatives?
This combination is powerful. The model is forced to capture the shared, slow-varying structure between past and future (maximizing mutual information) while ignoring unpredictable noise. And since the prediction happens in embedding space, it's fast and works for any modality — audio, images, text, RL.
The architecture: three simple pieces
CPC chains three components end-to-end — think of them as a microphone, a narrator, and a detective:
1. Encoder — the microphone. It compresses each raw observation (an audio frame, an image patch, a word) into a latent embedding . This strips away high-frequency noise and keeps the useful content. For speech, the encoder is a strided convolutional network running directly on raw 16 kHz audio; for images, it's a ResNet; for text, a lookup.
2. Autoregressive model — the narrator. It reads the sequence of embeddings up to the current time step and produces a single context vector that summarizes everything seen so far. This is a , , or masked CNN — any model that captures temporal order. The context is the model's "running story" of what has happened.
3. Linear predictors — the detectives. For each future step (1, 2, 3, …), a separate linear transform predicts what the future embedding should look like. Each predictor specializes in a different "look-ahead distance."
InfoNCE: the contrastive loss that started it all
The key question is: how do you train this pipeline without labels? The answer is a cleverly designed loss function called InfoNCE (Information Noise-Contrastive Estimation).
The setup is like a police lineup: at every time step and for each future step , assemble a set of embeddings. Exactly one is the real future (the positive); the other are random embeddings drawn from the data (the negatives). The model must identify which one is real, using only its context vector .
The scoring function is a simple that measures compatibility: . Higher score means "this looks like the real future given my context."
Read it as a classifier over candidates: the loss is categorical cross-entropy where the "correct class" is the true future. If the model perfectly distinguishes the positive from the negatives, the loss reaches its minimum . If it guesses randomly, the loss is .
The beautiful part is the theoretical guarantee: minimizing InfoNCE is equivalent to maximizing a lower bound on mutual information between and : . As you add more negatives, the bound gets tighter — the model is forced to capture more shared information.
Why mutual information matters
Mutual information measures how much knowing the context reduces your uncertainty about the future . Maximizing it forces the model to encode exactly the information that is shared between past and future — the slow, structural patterns — while ignoring the information that is private to each timestep — the noise.
Think of it as a filter for relevance: the melody is shared between the first and second bar of music (high mutual information), but the exact recording noise at each moment is independent (low mutual information). The model learns to represent the melody and discard the noise — not because we told it to, but because that's the only way to solve the contrastive task.
This is why CPC captures general-purpose features: the shared structure across time is exactly what downstream tasks need — phoneme identity in speech, object identity in images, topic coherence in text.
The density ratio trick: why CPC avoids generative modeling
A critical design choice: CPC does not model the full conditional distribution or the marginal . It only models their ratio:
This is the : it tells you how much more likely a future observation is given the context versus in general. CPC learns this ratio directly via the contrastive loss, sidestepping the need for any .
Why is this clever? Modeling in high-dimensional spaces (pixels, audio) is extremely hard and wastes capacity on noise. But the ratio captures only the dependency between context and future — exactly the mutual information we want. It's the shortcut from raw data to high-level understanding.
CPC in code
Simplified to show the idea — not the real implementation.
import numpy as np
def info_nce_loss(context, positive, negatives):
"""
context: (d,) — the summary of the past
positive: (d,) — the true future embedding
negatives: (N-1, d) — random embeddings from other timesteps
Returns: scalar loss (lower = better at finding the real future)
"""
# Score every candidate: dot product measures compatibility
pos_score = context @ positive # scalar
neg_scores = negatives @ context # (N-1,)
# Stack all scores: positive first, then negatives
all_scores = np.concatenate([[pos_score], neg_scores])
# Softmax → probability of each candidate being "the real future"
logits = all_scores - all_scores.max() # numerical stability
log_softmax = logits - np.log(np.exp(logits).sum())
# Cross-entropy: we want the first (positive) to get all the probability
loss = -log_softmax[0]
return loss
# This is categorical cross-entropy with the positive as the correct class.
# Minimizing it = maximizing mutual information between context and future.
# SimCLR, MoCo, and wav2vec 2.0 all use variants of this exact loss.One framework, four domains
The paper's strongest claim is universality: the same contrastive-predictive objective works across fundamentally different data types. The encoder changes, but the loss and the architecture pattern remain identical:
Audio (LibriSpeech): The encoder is a 5-layer strided CNN on raw 16 kHz waveforms, producing a vector every 10ms. The autoregressive model is a GRU. Despite never seeing phone labels, the representations achieve 64.6% accuracy on phoneme classification with a simple — matching supervised baselines.
Vision (ImageNet): Images are split into a 7×7 grid of overlapping patches. Each patch is encoded by a ResNet-v2-101. A masked CNN (like PixelCNN) serves as the autoregressive model, predicting downward column by column. CPC achieves 48.7% top-1 accuracy with a linear classifier on the frozen representations.
Text (BookCorpus): Each sentence is the unit. A 1D CNN encodes words into a sentence embedding; a GRU summarizes past sentences into context. The model predicts future sentence embeddings.
Reinforcement Learning (DeepMind Lab): CPC is added as an auxiliary loss to an A3C agent in 3D navigation tasks. It improves sample efficiency and speeds up learning, especially in visually complex environments.
Results: representations that transfer
The gold standard for evaluating unsupervised representations is the linear probe: freeze the learned encoder, then train only a linear classifier on top. If a simple linear layer can classify well, the representations have already organized the data into linearly separable clusters — proving the encoder learned meaningful structure, not just memorized patterns.
On speech phoneme classification with LibriSpeech, CPC representations achieve 64.6% phone accuracy — competitive with fully supervised approaches and far above random or MFCC baselines. On ImageNet with a linear classifier on frozen ResNet-v2 features, CPC reaches 48.7% top-1 accuracy. On the text domain, sentence-level representations show strong performance on sentiment classification and question-type detection.
Perhaps most revealing: CPC representations encode multiple levels of structure simultaneously. In the speech domain, the same representation that predicts future audio also captures both phoneme identity and speaker identity — two orthogonal tasks — in linearly separable subspaces. The model wasn't trained on either task, yet both are readable from the same embedding.
Why CPC changed everything
2013
Word2Vec
Showed that predicting context words (a contrastive objective via negative sampling) produces powerful word embeddings. The spiritual ancestor of InfoNCE.
2018
CPC — Contrastive Predictive Coding
InfoNCE loss + latent prediction + universality across audio, vision, text, and RL. The seed paper for modern contrastive learning.
2020
SimCLR — Simple Contrastive Learning
Replaced temporal prediction with augmentation-based pairs. Used InfoNCE with a temperature parameter. Showed that bigger batches (= more negatives) and strong augmentations are key.
2020
MoCo — Momentum Contrast
Solved the "need huge batches" problem by maintaining a momentum-updated queue of negatives. Made contrastive learning scalable without giant GPU memory.
2020
wav2vec 2.0
CPC's direct descendant for speech. Combined contrastive loss with masked prediction and a quantization module. Achieved near-human speech recognition with just 10 minutes of labeled data.
2021
CLIP
Extended contrastive learning to vision-language pairs. InfoNCE on image-text batches produced the most transferable visual representations ever built.
CPC is the bridge between Word2Vec's intuition — "predict your neighbors to learn meaning" — and the modern self-supervised paradigm where a single pre-trained encoder, trained without labels, powers everything from speech recognition to image search to multimodal understanding.
Citationvan den Oord, Li, Vinyals. Representation Learning with Contrastive Predictive Coding. arXiv, 2018.
Terms in this paper
- InfoNCE Lossخسارة InfoNCE
- Contrastive Predictive Coding (CPC)الترميز التنبُّئي التبايُني (CPC)
- Density Ratioنسبة الكثافة
- Mutual Informationالمعلومات المتبادلة
- Contrastive Learningالتعلم التبايُني
- Negative Samplingالتعيين السلبي
- Linear Probeالمسبار الخطي
- Autoregressive Modelالنموذج التوليدي التراجعي