Generative Models2013intermediate10 min read

Auto-Encoding Variational Bayes

الترميز الذاتي بنهج بايز التبايُني

Kingma, D.P. · Welling, M. — ICLR

The problem

Classical autoencoders compress data into a deterministic code, but the resulting is full of gaps — meaningless regions that produce garbage when decoded. Meanwhile, traditional requires model-specific derivations that don't scale to large datasets or deep neural networks.

The contribution

The : a principled that wraps an inside a probabilistic framework. The outputs a distribution (mean + variance) instead of a fixed vector, and the makes sampling differentiable so the whole system trains with standard . The function — the — naturally balances reconstruction quality against latent-space regularity.

The impact

The VAE established the modern framework for deep generative models. It spawned β-VAE for disentangled representations, inspired the latent-space idea behind diffusion models (DDPM), and remains the backbone of world models in reinforcement learning. Every time an AI generates, interpolates, or imagines — the VAE's probabilistic latent space is the ancestor of that capability.

A regular autoencoder is a photocopy machine: it compresses each document into a filing code and retrieves it perfectly — but ask for a code between two documents and you get a smeared mess.

The VAE replaces the filing cabinet with a cloud of fireflies: each document scatters a soft glow around its location, and the overlapping glows fill the gaps. Reach into any lit region and you pull out a coherent, new document that blends the traits of its neighbors.

The problem: autoencoders memorize, they don't imagine

A standard autoencoder squeezes input xx through a into a latent code zz, then reconstructs x^\hat{x} from zz. It learns to minimize — and it does so brilliantly. But the latent space it builds has a fatal flaw: the codes are scattered islands with dead ocean between them.

Pick a point between two codes and decode it — you get noise, not a meaningful blend. The autoencoder never learned what "between" means because it was never forced to fill in the gaps. It memorized a lookup table, not a generative model.

What we need is a latent space where every point decodes to something meaningful — a continuous, navigable landscape. The VAE achieves this by making one radical change: instead of encoding each input as a point, encode it as a probability distribution.

Open in Lab
Left: a standard autoencoder's latent space — isolated islands with dead zones. Right: a VAE's latent space — overlapping distributions fill every gap.
The demo wakes as you arrive…

The Bayesian foundation: why distributions?

The VAE is rooted in Bayesian thinking. We assume there is a hidden generative process: nature first picks a zz from a simple p(z)p(z) (a standard Gaussian), then generates the data xx from pθ(x∣z)p_\theta(x|z) (the ). Given an observed xx, we want the p(z∣x)p(z|x) — "which zz could have generated this xx?" — but computing it requires integrating over all possible zz values, which is .

The solution: approximate the posterior with a learned encoder qϕ(z∣x)q_\phi(z|x) that outputs a N(μ,σ2)\mathcal{N}(\mu, \sigma^2) for each input. This is called variational inference — we replace an impossible exact computation with a trainable approximation.

VAE architecture: encode a cloud, decode a sample

The VAE has two neural networks working in tandem:

The encoder qϕ(z∣x)q_\phi(z|x) takes an input xx and outputs two vectors: a mean μ\mu and a log-variance log⁡σ2\log \sigma^2. Together, they define a Gaussian cloud in latent space — the model's belief about where this input lives.

The decoder pθ(x∣z)p_\theta(x|z) takes a single point zz sampled from that cloud and reconstructs the original input. If xx is an image, the decoder outputs pixel intensities; if it's continuous data, it may output Gaussian parameters.

The magic: because we encode distributions instead of points, the clouds for similar inputs overlap, filling the latent space with meaningful content everywhere.

Open in Lab
Click any layer to see its role in the encoder–decoder pipeline.
The demo wakes as you arrive…

The ELBO: one loss, two forces

We cannot maximize the true data likelihood log⁡p(x)\log p(x) directly — it requires integrating over all zz. Instead, we maximize a tractable lower bound called the Evidence Lower BOund (ELBO). Think of it as tightening a floor beneath the true objective: pushing the floor up always pushes the objective up too.

The ELBO decomposes into two terms that pull in opposite directions, creating a productive tension:

L(θ,ϕ;x)=Eqϕ(z∣x)[log⁡pθ(x∣z)]⏟Reconstruction−DKL(qϕ(z∣x)∥p(z))⏟Regularization\mathcal{L}(\theta,\phi;x) = \underbrace{\mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)]}_{\text{Reconstruction}} - \underbrace{D_{KL}(q_\phi(z|x) \| p(z))}_{\text{Regularization}}
The Evidence Lower Bound (ELBO) — the VAE's entire objective — Term 1 (reconstruction): how well does the decoder rebuild x from z? Maximize it. · Term 2 (KL divergence): how far is the encoder's distribution from the standard Gaussian prior? Minimize it. · Together they force the model to reconstruct well AND keep a smooth, regular latent space.

Imagine two coaches training an athlete. The reconstruction coach demands perfect memory: "reproduce the input exactly!" The KL coach demands simplicity: "keep your representations close to a standard Gaussian — don't over-specialize!" The athlete (the model) must satisfy both, and the sweet spot is a latent space that's both expressive and well-organized.

Open in Lab
Drag the KL weight to see how reconstruction quality and latent regularity trade off.
The demo wakes as you arrive…

The reparameterization trick: making randomness differentiable

There's a critical obstacle: the encoder outputs a distribution, and we sample zz from it. But sampling is not differentiable — you can't backpropagate through a random number generator. Without gradients, the encoder can't learn.

The reparameterization trick solves this elegantly. Instead of sampling zz directly from N(μ,σ2)\mathcal{N}(\mu, \sigma^2), we:

  1. Sample noise ϵ∼N(0,1)\epsilon \sim \mathcal{N}(0, 1) — this is a fixed, external random source

  2. Compute z=μ+σ⋅ϵz = \mu + \sigma \cdot \epsilon

Now zz is a deterministic function of μ\mu and σ\sigma (which we can differentiate) plus an external noise source (which doesn't need gradients). The randomness is still there, but it's been moved outside the . Think of it as a river: we can't steer the rain, but we can build channels that direct where the water flows.

z=μϕ(x)+σϕ(x)⋅ϵ,ϵ∼N(0,I)z = \mu_\phi(x) + \sigma_\phi(x) \cdot \epsilon, \quad \epsilon \sim \mathcal{N}(0, I)
The reparameterization trick — the key enabler — μ and σ are outputs of the encoder (learnable) · ε is external noise (not learnable) · Gradients flow through μ and σ back to the encoder · The sampling randomness comes only from ε
Open in Lab
Step through the trick: see how gradients bypass the sampling step via ε.
The demo wakes as you arrive…

The KL term in closed form

When both the qϕ(z∣x)q_\phi(z|x) and the prior p(z)p(z) are Gaussians, the has a beautiful closed-form solution — no sampling required. For a latent space of dimension JJ:

DKL=−12∑j=1J(1+log⁡σj2−μj2−σj2)D_{KL} = -\frac{1}{2} \sum_{j=1}^{J} \left(1 + \log \sigma_j^2 - \mu_j^2 - \sigma_j^2\right)
KL divergence between two Gaussians — no sampling needed — Each latent dimension j contributes independently · If μ_j = 0 and σ_j = 1, that dimension's KL = 0 (it matches the prior perfectly) · The encoder is penalized for pushing means away from 0 or shrinking/expanding variances away from 1

The payoff: a navigable latent space

The KL regularizer's pressure toward the prior doesn't just keep training stable — it gives the latent space two extraordinary properties:

Continuity — nearby points in latent space decode to similar outputs. Walk smoothly from a "3" to an "8" and you see one digit gradually morphing into the other, not a sudden jump.

Completeness — every region of latent space decodes to something plausible. There are no "dead zones" producing garbage. The overlapping Gaussian clouds ensure coverage.

These properties make the VAE's latent space a creative tool: interpolate between data points, generate novel examples, or explore the space to discover what the model has learned.

Open in Lab
Drag the slider to interpolate between two digits. Notice the smooth transitions — every intermediate point is a valid image.
The demo wakes as you arrive…

The same idea in code

VAE forward pass — the complete corepython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn
import torch.nn.functional as F

class VAE(nn.Module):
    def __init__(self, input_dim=784, latent_dim=20):
        super().__init__()
        # Encoder: input → hidden → (mu, log_var)
        self.fc1 = nn.Linear(input_dim, 400)
        self.fc_mu = nn.Linear(400, latent_dim)
        self.fc_logvar = nn.Linear(400, latent_dim)
        # Decoder: z → hidden → reconstructed input
        self.fc3 = nn.Linear(latent_dim, 400)
        self.fc4 = nn.Linear(400, input_dim)

    def encode(self, x):
        h = F.relu(self.fc1(x))
        return self.fc_mu(h), self.fc_logvar(h)      # two outputs!

    def reparameterize(self, mu, log_var):
        std = torch.exp(0.5 * log_var)               # σ = exp(log_var / 2)
        eps = torch.randn_like(std)                   # ε ~ N(0, I)
        return mu + std * eps                          # z = μ + σ·ε

    def decode(self, z):
        h = F.relu(self.fc3(z))
        return torch.sigmoid(self.fc4(h))             # pixel probabilities

    def forward(self, x):
        mu, log_var = self.encode(x)
        z = self.reparameterize(mu, log_var)
        return self.decode(z), mu, log_var

# ELBO loss = reconstruction + KL divergence
def vae_loss(recon_x, x, mu, log_var):
    recon = F.binary_cross_entropy(recon_x, x, reduction='sum')
    kl = -0.5 * torch.sum(1 + log_var - mu.pow(2) - log_var.exp())
    return recon + kl

Training: the AEVB algorithm

The Auto-Encoding Variational Bayes (AEVB) algorithm is remarkably simple. For each :

  1. Encode: pass each data point xx through the encoder to get μ\mu and log⁡σ2\log \sigma^2

  2. Sample: draw zz using the reparameterization trick

  3. Decode: reconstruct x^\hat{x} from zz

  4. Compute the ELBO loss: reconstruction error + KL divergence

  5. Backpropagate and update θ\theta (decoder) and ϕ\phi (encoder) simultaneously

That's it. Standard (or ) handles the rest. No EM steps, no MCMC chains, no model-specific derivations. The algorithm scales to millions of data points because it works with mini-batches — exactly like training any other neural network.

Generation: sampling new data

Once trained, generation is trivially simple: sample z∼N(0,I)z \sim \mathcal{N}(0, I) from the prior and pass it through the decoder. Because the KL term trained the encoder to match the prior, the decoder has seen samples from that distribution during training and knows what to do with them.

This is fundamentally different from a , which needs an adversarial training loop. The VAE generates by : draw from the prior, decode. No , no , no training instabilities.

Open in Lab
Click "Generate" to sample from the prior and decode new images.
The demo wakes as you arrive…

Connection to EM and variational inference

The VAE's framework connects directly to two classical ideas:

The alternates between inferring latent variables () and updating model parameters (). The VAE does both simultaneously via — the encoder performs a "neural E-step" and the decoder a "neural M-step."

Variational inference traditionally requires hand-derived update equations for each model. The VAE's key insight: let neural networks parameterize both the approximate posterior and the generative model, then optimize everything with backpropagation. This is "" — the encoder network learns to perform inference for any input in a single , rather than running an optimization loop per data point.

Why it mattered

  1. 2013

    VAE (this paper)

    The reparameterization trick enables end-to-end training of deep generative models with latent variables. Published alongside Rezende et al.'s similar approach.

  2. 2014

    GAN

    Goodfellow introduced an adversarial alternative: a generator and discriminator compete. Sharper images but no encoder, no latent inference, and training instabilities.

  3. 2017

    β-VAE

    Scaling the KL weight beyond 1 disentangles latent dimensions — each captures a single factor of variation (pose, color, size). Enabled controllable generation.

  4. 2017

    VQ-VAE

    Replaced Gaussian latent codes with discrete codebook entries. Combined the VAE framework with vector quantization for high-quality image and audio generation.

  5. 2020

    DDPM (Diffusion Models)

    Denoising diffusion works in a latent-space-like framework. The VAE's autoencoder stage later became essential for Latent Diffusion (Stable Diffusion).

  6. 2018

    World Models

    Ha & Schmidhuber used a VAE to compress game frames into a latent representation, then trained an RL agent entirely inside the latent dream world.

  7. 2018

    Normalizing Flows

    Extended the VAE idea with invertible transformations that allow exact likelihood computation — overcoming the "blurry output" limitation.

The VAE's legacy isn't just the model itself — it's the idea that you can train deep generative models with a simple, principled loss function. Every latent-variable model that followed, from diffusion to VQ-VAE to world models, owes a debt to this 2013 paper that showed how to make randomness differentiable.

CitationKingma, Welling. Auto-Encoding Variational Bayes. ICLR, 2014.

Terms in this paper