Representation Learning2017intermediate10 min read

Neural Discrete Representation Learning

تعلُّم التمثيلات المنفصلة بالشبكات العصبية

van den Oord, A. · Vinyals, O. · Kavukcuoglu, K. — NeurIPS

The problem

Variational autoencoders (VAEs) use continuous latent variables — smooth, real-valued vectors — to represent data. But many real-world modalities are inherently discrete: language is a sequence of symbols, speech is a sequence of phonemes, and images can be described by words. Continuous latents also suffer from "": when paired with a powerful , the decoder learns to ignore the entirely and generates output from scratch. Previous attempts at discrete latent variables (Gumbel Softmax, NVIL, VIMCO) suffered from high variance gradients and never matched the performance of continuous VAEs.

The contribution

VQ-VAE: a that replaces continuous latent variables with discrete codes from a learned . The maps each spatial region to a continuous vector, then a nearest-neighbour lookup snaps it to the closest entry in the codebook — vector quantisation. A copies gradients from the decoder back to the encoder, bypassing the non-differentiable quantisation step. A three-part loss trains encoder, codebook, and decoder jointly. The result: the first discrete-latent VAE to match continuous VAEs in likelihood, while naturally avoiding posterior collapse and learning meaningful, compressed representations across images, speech, and video.

The impact

VQ-VAE established discrete latent spaces as a practical alternative to continuous ones and became a foundational building block for modern generative AI. DALL·E used a VQ-VAE to tokenise images before generating them with a . Jukebox applied the same idea to music. BEiT used VQ-VAE codes as visual tokens for self-supervised vision Transformers. Latent diffusion models compress images into a inspired by VQ-VAE before running diffusion — the architecture behind Stable Diffusion. The core insight — that you can treat continuous signals as sequences of discrete tokens — bridged the gap between language models and other modalities.

A regular compresses data into a smooth landscape — you can land anywhere on a continuous surface. The problem is that nearby points may mean very different things, and a powerful decoder learns to ignore the landscape entirely.

VQ-VAE replaces the smooth landscape with a peg board: exactly 512 pegs, each at a fixed position. The encoder throws a ring and it snaps to the nearest peg. The decoder only knows how to read pegs, so it must use them — no ignoring allowed. The pegs themselves learn to move to the most useful positions during .

The result? A compact, unambiguous code — like describing a painting by its palette numbers rather than the exact RGB of every pixel.

The problem: continuous latent spaces are too smooth

The Variational Autoencoder (VAE) was a breakthrough: compress data into a low-dimensional latent space, then reconstruct from it. But VAE latent spaces are continuous — every point in the space is valid, and tiny movements can produce wildly different outputs. This creates two problems:

  • Posterior collapse. When the decoder is powerful enough (e.g. a PixelCNN), it learns to generate data without consulting the latent code at all. The encoder's output is ignored, and the latent space becomes meaningless — it "collapses" to the .

  • Mismatch with discrete modalities. Language, phonemes, object categories — many important aspects of the world are fundamentally discrete. Forcing them into a continuous space wastes capacity on encoding smoothness rather than meaning.

Open in Lab
Left: a continuous latent space where the decoder can ignore the code. Right: a discrete codebook that the decoder must consult. Toggle to see the difference.
The demo wakes as you arrive…

The idea: snap to the nearest codebook entry

VQ-VAE keeps the encoder-decoder structure of a VAE but adds a crucial middle step: vector quantisation. Here is the full pipeline:

  1. The encoder takes input xx (an image, audio clip, or video frame) and produces a grid of continuous vectors ze(x)z_e(x) — one vector per spatial position.

  2. Each vector is matched to its nearest neighbour in a learned codebook e∈RK×De \in \mathbb{R}^{K \times D}, where KK is the codebook size (e.g. 512) and DD is the dimension. The index of the nearest entry becomes the discrete code for that position.

  3. The decoder receives the corresponding codebook vectors zq(x)z_q(x) and reconstructs the original input. Because the decoder only ever sees codebook entries, it must use them — eliminating posterior collapse by design.

Open in Lab
Follow the data from input through encoder, codebook lookup, and decoder. Click each stage to highlight it.
The demo wakes as you arrive…

Vector quantisation: from continuous to discrete

The quantisation step is deceptively simple. Given the encoder output ze(x)z_e(x), find the codebook entry eke_k that minimises the Euclidean distance:

zq(x)=ek,where k=arg⁡min⁡j∥ze(x)−ej∥2z_q(x) = e_k, \quad \text{where } k = \arg\min_j \| z_e(x) - e_j \|_2
Nearest-neighbour lookup — the quantisation step — The encoder's continuous output is replaced by the nearest codebook vector. This is deterministic: each position gets exactly one code.

Think of it as a post office sorting machine: each parcel (encoder output) slides down a ramp and drops into the nearest bin (codebook entry). The parcel's identity is now its bin number — a discrete code.

The posterior distribution q(z∣x)q(z|x) becomes one-hot: probability 1 for the nearest codebook entry, 0 for everything else. With a uniform prior p(z)p(z), the is constant (log⁡K\log K) and can be ignored during training — a major simplification over standard VAEs.

Open in Lab
Drag the encoder output (blue dot) across the embedding space. Watch it snap to the nearest codebook entry (orange) and see the selected index update.
The demo wakes as you arrive…

The gradient problem: how to learn through a hard snap

The nearest-neighbour lookup is not differentiable — there is no smooth through arg⁡min⁡\arg\min. If we cannot backpropagate through the quantisation step, the encoder gets no learning signal from the reconstruction loss. VQ-VAE solves this with the straight-through estimator: during the , use the quantised codes zq(x)z_q(x); during the , simply copy the gradient from zq(x)z_q(x) to ze(x)z_e(x), as if the quantisation never happened.

Why does this work? Because ze(x)z_e(x) and zq(x)z_q(x) live in the same DD-dimensional space. The gradient ∇zL\nabla_z L tells the encoder "move your output this way to reduce " — and that direction is just as useful at zez_e as it is at zqz_q, since they are close neighbours in the same space.

Open in Lab
Forward pass (blue arrows) goes through quantisation. Backward pass (red arrows) copies the gradient directly, bypassing the snap.
The demo wakes as you arrive…

The three-part loss: decoder, codebook, and commitment

Three different parts of VQ-VAE need to learn, and the straight-through estimator alone does not train the codebook vectors. The solution is an elegant three-term loss:

L=log⁡p(x∣zq(x))⏟reconstruction+∥sg[ze(x)]−e∥22⏟codebook+β∥ze(x)−sg[e]∥22⏟commitmentL = \underbrace{\log p(x | z_q(x))}_{\text{reconstruction}} + \underbrace{\| \text{sg}[z_e(x)] - e \|_2^2}_{\text{codebook}} + \underbrace{\beta \| z_e(x) - \text{sg}[e] \|_2^2}_{\text{commitment}}
The VQ-VAE loss — three terms, three learners — sg = Stop Gradient (treat the operand as a constant). Term 1 trains the decoder + encoder. Term 2 moves codebook entries toward encoder outputs. Term 3 prevents the encoder from drifting away from the codebook. β ≈ 0.25 in practice.

Think of it as a three-way negotiation:

  • The decoder says: "give me codes that let me reconstruct well" → reconstruction loss.
  • The codebook says: "I'll move my entries to where the encoder is pointing" → codebook loss (like k-means).
  • The encoder says: "I promise not to run too far from the nearest codebook entry" → .

The Stop Gradient operator sg[·] is what keeps each term from interfering with the others: the codebook loss only moves the codebook (not the encoder), and the commitment loss only constrains the encoder (not the codebook).

Open in Lab
Adjust β (commitment weight) and watch how the encoder output, codebook entries, and reconstruction quality change. Higher β keeps the encoder closer to the codebook.
The demo wakes as you arrive…

Learning the prior: from compression to generation

After VQ-VAE is trained, the latent space is a grid of discrete codes — essentially a low-resolution "image" of indices. To generate new data, we need to learn the distribution of these codes. VQ-VAE trains a separate as the prior:

  • For images: a PixelCNN over the 2D grid of discrete latents. Because the grid is small (e.g. 32×32 instead of 128×128 pixels), the PixelCNN can model global structure efficiently.
  • For audio: a WaveNet over the 1D sequence of discrete latents. With 64× compression, the WaveNet captures long-range phoneme-level patterns that are invisible in raw waveforms.

This two-stage approach — first learn a compressed codebook, then model the code distribution — is exactly the recipe later adopted by DALL·E and Jukebox.

The VQ layer in code

Vector quantisation — the core mechanismpython

Simplified to show the idea — not the real implementation.

import numpy as np

class VectorQuantizer:
    """A codebook of K vectors, each of dimension D."""
    def __init__(self, K=512, D=64, beta=0.25):
        self.codebook = np.random.randn(K, D) * 0.1   # K entries
        self.beta = beta                                # commitment weight

    def forward(self, z_e):
        """z_e: (N, D) — one encoder output per spatial position."""
        # 1. Find nearest codebook entry for each encoder output
        dists = np.sum((z_e[:, None] - self.codebook[None]) ** 2, axis=-1)  # (N, K)
        indices = np.argmin(dists, axis=-1)                                  # (N,)
        z_q = self.codebook[indices]                                          # (N, D)

        # 2. Compute losses
        codebook_loss = np.mean((z_e.copy() - z_q) ** 2)  # sg on z_e: move codebook
        commit_loss   = np.mean((z_e - z_q.copy()) ** 2)  # sg on e:  constrain encoder

        loss = codebook_loss + self.beta * commit_loss

        # 3. Straight-through: copy z_q to output, but gradient goes to z_e
        z_q_st = z_e + (z_q - z_e)   # forward: z_q, backward: gradient → z_e

        return z_q_st, loss, indices

What VQ-VAE learned: images, speech, and video

The paper demonstrates VQ-VAE on three modalities, each revealing different strengths of discrete latent spaces:

Images (ImageNet 128×128). The encoder compresses 128×128×3 images to a 32×32×1 grid of discrete codes with K=512K=512 — a 42.6× reduction in bits. Reconstructions are only slightly blurrier than originals. A PixelCNN prior over the 32×32 latent grid generates coherent full images with global structure.

Speech (VCTK, 109 speakers). With 64× temporal compression, the latent codes capture phoneme-level content and discard speaker-specific details like pitch and timbre. When reconstructed with a different speaker identity fed to the decoder, the result is speaker conversion: same words, different voice. Mapping the 512 discrete codes to phonemes gives 49.3% accuracy — versus 7.2% random chance — proving the codes discovered phoneme-like units without any linguistic supervision.

Video (DeepMind Lab). VQ-VAE generates action-conditioned video frames entirely in the latent space, only decoding to pixels at the end. This avoids the compounding errors of pixel-space generation and maintains visual quality over long sequences.

Open in Lab
See how VQ-VAE compresses a 128×128 image to a 32×32 grid of codebook indices. Toggle between original, latent grid, and reconstruction.
The demo wakes as you arrive…

Why discrete latents changed everything

  1. 2013

    VAE (Kingma & Welling)

    Variational autoencoders with continuous Gaussian latent spaces. Enabled generation by sampling, but suffered from blurry outputs and posterior collapse with powerful decoders.

  2. 2016

    PixelCNN & WaveNet

    Powerful autoregressive models for images and audio. Could generate high-quality samples but were slow and operated on raw pixels/samples without learned compression.

  3. 2017

    VQ-VAE (this paper)

    Discrete latent spaces via vector quantisation. First discrete-latent VAE to match continuous performance. Discovered phonemes from raw speech without supervision.

  4. 2019

    VQ-VAE-2

    Hierarchical VQ-VAE with multiple codebook scales. Generated high-fidelity 256×256 images rivalling GANs, without adversarial training.

  5. 2020

    Jukebox (OpenAI)

    Hierarchical VQ-VAE for raw music generation. Generated minutes of coherent music with lyrics and style, applying the discrete-token idea to audio at unprecedented scale.

  6. 2021

    DALL·E (OpenAI)

    Used a discrete VAE (dVAE, a VQ-VAE variant) to tokenise images, then generated image tokens from text captions using a Transformer. Text-to-image generation at scale.

  7. 2021

    BEiT

    Used VQ-VAE codes as visual tokens for masked image modelling — the visual equivalent of BERT's masked language modelling. Showed that discrete image tokens enable powerful self-supervised vision.

  8. 2022

    Latent Diffusion / Stable Diffusion

    Runs diffusion in a compressed latent space (inspired by VQ-VAE's encoder-decoder) rather than pixel space. Drastically reduced compute while maintaining image quality — the architecture behind Stable Diffusion.

VQ-VAE's legacy is a single, powerful idea: continuous signals can be faithfully represented as sequences of discrete tokens from a learned dictionary. This idea unlocked the Transformer-as-universal-generator paradigm — where the same architecture that models language can model images, music, video, and beyond.

Citationvan den Oord, Vinyals, Kavukcuoglu. Neural Discrete Representation Learning. NeurIPS, 2017.

Terms in this paper