Self-Supervised Learning2018intermediate11 min read

Representation Learning with Contrastive Predictive Coding

تعلُّم التمثيلات عبر الترميز التنبُّئي التبايُني

van den Oord, A. · Li, Y. · Vinyals, O. — arXiv

The problem

Unsupervised learning in 2018 relied on either generative models that reconstruct raw inputs pixel-by-pixel (expensive and focused on low-level detail) or purely discriminative approaches limited to single modalities. There was no universal framework that could learn high-level features from raw data — audio, images, text, or observations — using a single, principled objective without any labels.

The contribution

: a self-supervised framework that learns representations by predicting the future in . An compresses raw observations into embeddings; an summarizes past embeddings into a ; a linear predictor uses the context to identify the true future among negatives, trained with a new called InfoNCE. InfoNCE is proven to be a lower bound on between context and future — so maximizing it forces the representation to capture the slow, high-level structure that makes the future predictable. The same framework works across audio, vision, text, and RL.

The impact

CPC introduced InfoNCE — the loss function that became the backbone of modern . SimCLR, MoCo, wav2vec 2.0, and CLIP all descend from this idea. CPC proved that predicting the future in latent space, rather than reconstructing raw inputs, produces representations that transfer across tasks and modalities. It catalyzed the self-supervised revolution that made large-scale pre- without labels not just possible, but dominant.

Imagine you're at a concert with your eyes closed. After hearing the first few bars of a melody, you can guess the next notes — not perfectly, but well enough to tell the real continuation from a random snippet of a different song.

You don't need to reconstruct the exact waveform of the next note. You don't need to be a musician. You just need to tell "what actually follows" apart from "what doesn't belong." That contrast — real future vs. random impostor — is the entire training signal.

CPC works the same way: it listens to the past, builds a compressed summary, and learns to pick the true future out of a lineup — and the skill it acquires doing this turns out to be a deep understanding of the data's structure.

The problem: generative models waste capacity on noise

Before CPC, the dominant approach to unsupervised learning was generative modeling: train a model to reconstruct its input (like a deep autoencoder) or predict every pixel or sample. The problem is that most of the information in raw signals — the exact brightness of each pixel, the precise waveform of audio — is high-frequency noise that says nothing about the underlying meaning.

A model that tries to perfectly reconstruct a face image spends most of its capacity modeling skin texture and lighting, not learning that it's a face. A speech model that reconstructs waveforms spends capacity on room acoustics, not phonemes.

What we actually want is the high-level, slow-changing structure: the melody across an audio clip, the object identity across image patches, the topic across a paragraph. That structure is what makes downstream tasks — , retrieval, reasoning — possible.

Open in Lab
Compare: a generative model reconstructs every pixel (wasting effort on noise), while CPC only needs to identify the real future from fakes.
The demo wakes as you arrive…

The idea: predict the future in latent space

CPC's insight is a shift in what you predict and how you evaluate the prediction:

What: Don't predict raw pixels or waveforms. Instead, predict a compact representation of the future — an embedding that captures the high-level content while discarding noise.

How: Don't compare your prediction to the truth pixel-by-pixel. Instead, treat it as a classification problem: can you pick the true future embedding out of a lineup of random negatives?

This combination is powerful. The model is forced to capture the shared, slow-varying structure between past and future (maximizing mutual information) while ignoring unpredictable noise. And since the prediction happens in embedding space, it's fast and works for any modality — audio, images, text, RL.

Open in Lab
The full CPC pipeline: encoder → autoregressive model → context → predict → contrast. Click each stage to see its role.
The demo wakes as you arrive…

The architecture: three simple pieces

CPC chains three components end-to-end — think of them as a microphone, a narrator, and a detective:

1. Encoder gencg_{\text{enc}} — the microphone. It compresses each raw observation xtx_t (an audio frame, an image patch, a word) into a latent embedding zt=genc(xt)z_t = g_{\text{enc}}(x_t). This strips away high-frequency noise and keeps the useful content. For speech, the encoder is a strided convolutional network running directly on raw 16 kHz audio; for images, it's a ResNet; for text, a lookup.

2. Autoregressive model garg_{\text{ar}} — the narrator. It reads the sequence of embeddings z1,z2,…,ztz_1, z_2, \ldots, z_t up to the current time step and produces a single context vector ct=gar(z≤t)c_t = g_{\text{ar}}(z_{\leq t}) that summarizes everything seen so far. This is a , , or masked CNN — any model that captures temporal order. The context is the model's "running story" of what has happened.

3. Linear predictors WkW_k — the detectives. For each future step kk (1, 2, 3, …), a separate linear transform z^t+k=Wkct\hat{z}_{t+k} = W_k c_t predicts what the future embedding should look like. Each predictor specializes in a different "look-ahead distance."

Open in Lab
Click any component to see its inputs, outputs, and role in the pipeline.
The demo wakes as you arrive…

InfoNCE: the contrastive loss that started it all

The key question is: how do you train this pipeline without labels? The answer is a cleverly designed loss function called InfoNCE (Information Noise-Contrastive Estimation).

The setup is like a police lineup: at every time step tt and for each future step kk, assemble a set of NN embeddings. Exactly one is the real future zt+kz_{t+k} (the positive); the other N−1N-1 are random embeddings drawn from the data (the negatives). The model must identify which one is real, using only its context vector ctc_t.

The scoring function is a simple that measures compatibility: fk(xt+k,ct)=exp⁡(zt+k⊤Wkct)f_k(x_{t+k}, c_t) = \exp(z_{t+k}^\top W_k c_t). Higher score means "this looks like the real future given my context."

LN=−EX[log⁡fk(xt+k,ct)∑xj∈Xfk(xj,ct)]\mathcal{L}_N = -\mathbb{E}_X \left[ \log \frac{f_k(x_{t+k}, c_t)}{\sum_{x_j \in X} f_k(x_j, c_t)} \right]
InfoNCE loss — the core training objective — The numerator scores the true future · the denominator scores all candidates (1 positive + N-1 negatives) · minimizing this loss = maximizing the probability of picking the real future = maximizing mutual information

Read it as a classifier over candidates: the loss is categorical cross-entropy where the "correct class" is the true future. If the model perfectly distinguishes the positive from the negatives, the loss reaches its minimum −log⁡(1)=0-\log(1) = 0. If it guesses randomly, the loss is log⁡(N)\log(N).

The beautiful part is the theoretical guarantee: minimizing InfoNCE is equivalent to maximizing a lower bound on mutual information between ctc_t and xt+kx_{t+k}: I(xt+k;ct)≥log⁡(N)−LNI(x_{t+k}; c_t) \geq \log(N) - \mathcal{L}_N. As you add more negatives, the bound gets tighter — the model is forced to capture more shared information.

Open in Lab
Drag to select the "real future" from the lineup. Watch how the loss changes as you add more negatives.
The demo wakes as you arrive…

Why mutual information matters

Mutual information I(xt+k;ct)I(x_{t+k}; c_t) measures how much knowing the context ctc_t reduces your uncertainty about the future xt+kx_{t+k}. Maximizing it forces the model to encode exactly the information that is shared between past and future — the slow, structural patterns — while ignoring the information that is private to each timestep — the noise.

Think of it as a filter for relevance: the melody is shared between the first and second bar of music (high mutual information), but the exact recording noise at each moment is independent (low mutual information). The model learns to represent the melody and discard the noise — not because we told it to, but because that's the only way to solve the contrastive task.

This is why CPC captures general-purpose features: the shared structure across time is exactly what downstream tasks need — phoneme identity in speech, object identity in images, topic coherence in text.

Open in Lab
Slide to adjust the number of negatives (N) and see how the mutual information lower bound tightens.
The demo wakes as you arrive…

The density ratio trick: why CPC avoids generative modeling

A critical design choice: CPC does not model the full conditional distribution p(xt+k∣ct)p(x_{t+k}|c_t) or the marginal p(xt+k)p(x_{t+k}). It only models their ratio:

fk(xt+k,ct)∝p(xt+k∣ct)p(xt+k)f_k(x_{t+k}, c_t) \propto \frac{p(x_{t+k}|c_t)}{p(x_{t+k})}

This is the : it tells you how much more likely a future observation is given the context versus in general. CPC learns this ratio directly via the contrastive loss, sidestepping the need for any .

Why is this clever? Modeling p(x)p(x) in high-dimensional spaces (pixels, audio) is extremely hard and wastes capacity on noise. But the ratio p(x∣c)/p(x)p(x|c)/p(x) captures only the dependency between context and future — exactly the mutual information we want. It's the shortcut from raw data to high-level understanding.

CPC in code

InfoNCE loss — the core of CPCpython

Simplified to show the idea — not the real implementation.

import numpy as np

def info_nce_loss(context, positive, negatives):
    """
    context:   (d,)  — the summary of the past
    positive:  (d,)  — the true future embedding
    negatives: (N-1, d) — random embeddings from other timesteps
    Returns: scalar loss (lower = better at finding the real future)
    """
    # Score every candidate: dot product measures compatibility
    pos_score = context @ positive                       # scalar
    neg_scores = negatives @ context                     # (N-1,)

    # Stack all scores: positive first, then negatives
    all_scores = np.concatenate([[pos_score], neg_scores])

    # Softmax → probability of each candidate being "the real future"
    logits = all_scores - all_scores.max()               # numerical stability
    log_softmax = logits - np.log(np.exp(logits).sum())

    # Cross-entropy: we want the first (positive) to get all the probability
    loss = -log_softmax[0]
    return loss

# This is categorical cross-entropy with the positive as the correct class.
# Minimizing it = maximizing mutual information between context and future.
# SimCLR, MoCo, and wav2vec 2.0 all use variants of this exact loss.

One framework, four domains

The paper's strongest claim is universality: the same contrastive-predictive objective works across fundamentally different data types. The encoder changes, but the loss and the architecture pattern remain identical:

Audio (LibriSpeech): The encoder is a 5-layer strided CNN on raw 16 kHz waveforms, producing a vector every 10ms. The autoregressive model is a GRU. Despite never seeing phone labels, the representations achieve 64.6% accuracy on phoneme classification with a simple — matching supervised baselines.

Vision (ImageNet): Images are split into a 7×7 grid of overlapping patches. Each patch is encoded by a ResNet-v2-101. A masked CNN (like PixelCNN) serves as the autoregressive model, predicting downward column by column. CPC achieves 48.7% top-1 accuracy with a linear classifier on the frozen representations.

Text (BookCorpus): Each sentence is the unit. A 1D CNN encodes words into a sentence embedding; a GRU summarizes past sentences into context. The model predicts future sentence embeddings.

Reinforcement Learning (DeepMind Lab): CPC is added as an auxiliary loss to an A3C agent in 3D navigation tasks. It improves sample efficiency and speeds up learning, especially in visually complex environments.

Open in Lab
Switch between audio, vision, text, and RL to see how CPC adapts its encoder while keeping the same loss.
The demo wakes as you arrive…

Results: representations that transfer

The gold standard for evaluating unsupervised representations is the linear probe: freeze the learned encoder, then train only a linear classifier on top. If a simple linear layer can classify well, the representations have already organized the data into linearly separable clusters — proving the encoder learned meaningful structure, not just memorized patterns.

On speech phoneme classification with LibriSpeech, CPC representations achieve 64.6% phone accuracy — competitive with fully supervised approaches and far above random or MFCC baselines. On ImageNet with a linear classifier on frozen ResNet-v2 features, CPC reaches 48.7% top-1 accuracy. On the text domain, sentence-level representations show strong performance on sentiment classification and question-type detection.

Perhaps most revealing: CPC representations encode multiple levels of structure simultaneously. In the speech domain, the same representation that predicts future audio also captures both phoneme identity and speaker identity — two orthogonal tasks — in linearly separable subspaces. The model wasn't trained on either task, yet both are readable from the same embedding.

Open in Lab
See how a frozen CPC encoder's representations cluster by phoneme. A linear classifier draws boundaries without modifying the encoder.
The demo wakes as you arrive…

Why CPC changed everything

  1. 2013

    Word2Vec

    Showed that predicting context words (a contrastive objective via negative sampling) produces powerful word embeddings. The spiritual ancestor of InfoNCE.

  2. 2018

    CPC — Contrastive Predictive Coding

    InfoNCE loss + latent prediction + universality across audio, vision, text, and RL. The seed paper for modern contrastive learning.

  3. 2020

    SimCLR — Simple Contrastive Learning

    Replaced temporal prediction with augmentation-based pairs. Used InfoNCE with a temperature parameter. Showed that bigger batches (= more negatives) and strong augmentations are key.

  4. 2020

    MoCo — Momentum Contrast

    Solved the "need huge batches" problem by maintaining a momentum-updated queue of negatives. Made contrastive learning scalable without giant GPU memory.

  5. 2020

    wav2vec 2.0

    CPC's direct descendant for speech. Combined contrastive loss with masked prediction and a quantization module. Achieved near-human speech recognition with just 10 minutes of labeled data.

  6. 2021

    CLIP

    Extended contrastive learning to vision-language pairs. InfoNCE on image-text batches produced the most transferable visual representations ever built.

CPC is the bridge between Word2Vec's intuition — "predict your neighbors to learn meaning" — and the modern self-supervised paradigm where a single pre-trained encoder, trained without labels, powers everything from speech recognition to image search to multimodal understanding.

Citationvan den Oord, Li, Vinyals. Representation Learning with Contrastive Predictive Coding. arXiv, 2018.

Terms in this paper