Representation Learning2017intermediate10 min read
Neural Discrete Representation Learning
تعلُّم التمثيلات المنفصلة بالشبكات العصبية
van den Oord, A. · Vinyals, O. · Kavukcuoglu, K. — NeurIPS
The problem
Variational autoencoders (VAEs) use continuous latent variables — smooth, real-valued vectors — to represent data. But many real-world modalities are inherently discrete: language is a sequence of symbols, speech is a sequence of phonemes, and images can be described by words. Continuous latents also suffer from "": when paired with a powerful , the decoder learns to ignore the entirely and generates output from scratch. Previous attempts at discrete latent variables (Gumbel Softmax, NVIL, VIMCO) suffered from high variance gradients and never matched the performance of continuous VAEs.
The contribution
VQ-VAE: a that replaces continuous latent variables with discrete codes from a learned . The maps each spatial region to a continuous vector, then a nearest-neighbour lookup snaps it to the closest entry in the codebook — vector quantisation. A copies gradients from the decoder back to the encoder, bypassing the non-differentiable quantisation step. A three-part loss trains encoder, codebook, and decoder jointly. The result: the first discrete-latent VAE to match continuous VAEs in likelihood, while naturally avoiding posterior collapse and learning meaningful, compressed representations across images, speech, and video.
The impact
VQ-VAE established discrete latent spaces as a practical alternative to continuous ones and became a foundational building block for modern generative AI. DALL·E used a VQ-VAE to tokenise images before generating them with a . Jukebox applied the same idea to music. BEiT used VQ-VAE codes as visual tokens for self-supervised vision Transformers. Latent diffusion models compress images into a inspired by VQ-VAE before running diffusion — the architecture behind Stable Diffusion. The core insight — that you can treat continuous signals as sequences of discrete tokens — bridged the gap between language models and other modalities.
A regular compresses data into a smooth landscape — you can land anywhere on a continuous surface. The problem is that nearby points may mean very different things, and a powerful decoder learns to ignore the landscape entirely.
VQ-VAE replaces the smooth landscape with a peg board: exactly 512 pegs, each at a fixed position. The encoder throws a ring and it snaps to the nearest peg. The decoder only knows how to read pegs, so it must use them — no ignoring allowed. The pegs themselves learn to move to the most useful positions during .
The result? A compact, unambiguous code — like describing a painting by its palette numbers rather than the exact RGB of every pixel.
The problem: continuous latent spaces are too smooth
The Variational Autoencoder (VAE) was a breakthrough: compress data into a low-dimensional latent space, then reconstruct from it. But VAE latent spaces are continuous — every point in the space is valid, and tiny movements can produce wildly different outputs. This creates two problems:
-
Posterior collapse. When the decoder is powerful enough (e.g. a PixelCNN), it learns to generate data without consulting the latent code at all. The encoder's output is ignored, and the latent space becomes meaningless — it "collapses" to the .
-
Mismatch with discrete modalities. Language, phonemes, object categories — many important aspects of the world are fundamentally discrete. Forcing them into a continuous space wastes capacity on encoding smoothness rather than meaning.
The idea: snap to the nearest codebook entry
VQ-VAE keeps the encoder-decoder structure of a VAE but adds a crucial middle step: vector quantisation. Here is the full pipeline:
-
The encoder takes input (an image, audio clip, or video frame) and produces a grid of continuous vectors — one vector per spatial position.
-
Each vector is matched to its nearest neighbour in a learned codebook , where is the codebook size (e.g. 512) and is the dimension. The index of the nearest entry becomes the discrete code for that position.
-
The decoder receives the corresponding codebook vectors and reconstructs the original input. Because the decoder only ever sees codebook entries, it must use them — eliminating posterior collapse by design.
Vector quantisation: from continuous to discrete
The quantisation step is deceptively simple. Given the encoder output , find the codebook entry that minimises the Euclidean distance:
Think of it as a post office sorting machine: each parcel (encoder output) slides down a ramp and drops into the nearest bin (codebook entry). The parcel's identity is now its bin number — a discrete code.
The posterior distribution becomes one-hot: probability 1 for the nearest codebook entry, 0 for everything else. With a uniform prior , the is constant () and can be ignored during training — a major simplification over standard VAEs.
The gradient problem: how to learn through a hard snap
The nearest-neighbour lookup is not differentiable — there is no smooth through . If we cannot backpropagate through the quantisation step, the encoder gets no learning signal from the reconstruction loss. VQ-VAE solves this with the straight-through estimator: during the , use the quantised codes ; during the , simply copy the gradient from to , as if the quantisation never happened.
Why does this work? Because and live in the same -dimensional space. The gradient tells the encoder "move your output this way to reduce " — and that direction is just as useful at as it is at , since they are close neighbours in the same space.
The three-part loss: decoder, codebook, and commitment
Three different parts of VQ-VAE need to learn, and the straight-through estimator alone does not train the codebook vectors. The solution is an elegant three-term loss:
Think of it as a three-way negotiation:
- The decoder says: "give me codes that let me reconstruct well" → reconstruction loss.
- The codebook says: "I'll move my entries to where the encoder is pointing" → codebook loss (like k-means).
- The encoder says: "I promise not to run too far from the nearest codebook entry" → .
The Stop Gradient operator sg[·] is what keeps each term from interfering with the others: the codebook loss only moves the codebook (not the encoder), and the commitment loss only constrains the encoder (not the codebook).
Learning the prior: from compression to generation
After VQ-VAE is trained, the latent space is a grid of discrete codes — essentially a low-resolution "image" of indices. To generate new data, we need to learn the distribution of these codes. VQ-VAE trains a separate as the prior:
- For images: a PixelCNN over the 2D grid of discrete latents. Because the grid is small (e.g. 32×32 instead of 128×128 pixels), the PixelCNN can model global structure efficiently.
- For audio: a WaveNet over the 1D sequence of discrete latents. With 64× compression, the WaveNet captures long-range phoneme-level patterns that are invisible in raw waveforms.
This two-stage approach — first learn a compressed codebook, then model the code distribution — is exactly the recipe later adopted by DALL·E and Jukebox.
The VQ layer in code
Simplified to show the idea — not the real implementation.
import numpy as np
class VectorQuantizer:
"""A codebook of K vectors, each of dimension D."""
def __init__(self, K=512, D=64, beta=0.25):
self.codebook = np.random.randn(K, D) * 0.1 # K entries
self.beta = beta # commitment weight
def forward(self, z_e):
"""z_e: (N, D) — one encoder output per spatial position."""
# 1. Find nearest codebook entry for each encoder output
dists = np.sum((z_e[:, None] - self.codebook[None]) ** 2, axis=-1) # (N, K)
indices = np.argmin(dists, axis=-1) # (N,)
z_q = self.codebook[indices] # (N, D)
# 2. Compute losses
codebook_loss = np.mean((z_e.copy() - z_q) ** 2) # sg on z_e: move codebook
commit_loss = np.mean((z_e - z_q.copy()) ** 2) # sg on e: constrain encoder
loss = codebook_loss + self.beta * commit_loss
# 3. Straight-through: copy z_q to output, but gradient goes to z_e
z_q_st = z_e + (z_q - z_e) # forward: z_q, backward: gradient → z_e
return z_q_st, loss, indicesWhat VQ-VAE learned: images, speech, and video
The paper demonstrates VQ-VAE on three modalities, each revealing different strengths of discrete latent spaces:
Images (ImageNet 128×128). The encoder compresses 128×128×3 images to a 32×32×1 grid of discrete codes with — a 42.6× reduction in bits. Reconstructions are only slightly blurrier than originals. A PixelCNN prior over the 32×32 latent grid generates coherent full images with global structure.
Speech (VCTK, 109 speakers). With 64× temporal compression, the latent codes capture phoneme-level content and discard speaker-specific details like pitch and timbre. When reconstructed with a different speaker identity fed to the decoder, the result is speaker conversion: same words, different voice. Mapping the 512 discrete codes to phonemes gives 49.3% accuracy — versus 7.2% random chance — proving the codes discovered phoneme-like units without any linguistic supervision.
Video (DeepMind Lab). VQ-VAE generates action-conditioned video frames entirely in the latent space, only decoding to pixels at the end. This avoids the compounding errors of pixel-space generation and maintains visual quality over long sequences.
Why discrete latents changed everything
2013
VAE (Kingma & Welling)
Variational autoencoders with continuous Gaussian latent spaces. Enabled generation by sampling, but suffered from blurry outputs and posterior collapse with powerful decoders.
2016
PixelCNN & WaveNet
Powerful autoregressive models for images and audio. Could generate high-quality samples but were slow and operated on raw pixels/samples without learned compression.
2017
VQ-VAE (this paper)
Discrete latent spaces via vector quantisation. First discrete-latent VAE to match continuous performance. Discovered phonemes from raw speech without supervision.
2019
VQ-VAE-2
Hierarchical VQ-VAE with multiple codebook scales. Generated high-fidelity 256×256 images rivalling GANs, without adversarial training.
2020
Jukebox (OpenAI)
Hierarchical VQ-VAE for raw music generation. Generated minutes of coherent music with lyrics and style, applying the discrete-token idea to audio at unprecedented scale.
2021
DALL·E (OpenAI)
Used a discrete VAE (dVAE, a VQ-VAE variant) to tokenise images, then generated image tokens from text captions using a Transformer. Text-to-image generation at scale.
2021
BEiT
Used VQ-VAE codes as visual tokens for masked image modelling — the visual equivalent of BERT's masked language modelling. Showed that discrete image tokens enable powerful self-supervised vision.
2022
Latent Diffusion / Stable Diffusion
Runs diffusion in a compressed latent space (inspired by VQ-VAE's encoder-decoder) rather than pixel space. Drastically reduced compute while maintaining image quality — the architecture behind Stable Diffusion.
VQ-VAE's legacy is a single, powerful idea: continuous signals can be faithfully represented as sequences of discrete tokens from a learned dictionary. This idea unlocked the Transformer-as-universal-generator paradigm — where the same architecture that models language can model images, music, video, and beyond.
Citationvan den Oord, Vinyals, Kavukcuoglu. Neural Discrete Representation Learning. NeurIPS, 2017.
Terms in this paper
- Autoencoderالمُرمِّز الذاتي
- Variational Autoencoder (VAE)المرمّز التلقائي المتغير الاحتمالي
- Vector Quantizationالتكميم المتجهي
- Codebookجدول الرموز
- Latent Spaceالفضاء الكامن
- Latent Representationالتمثيل الكامن
- Posterior Collapseانهيار الاحتمال البعدي
- Commitment Lossخسارة الالتزام
- Straight-Through Estimatorمُقدِّر المرور المباشر
- Reconstruction Errorخطأ إعادة البناء
- Embeddingالتضمين