Generative Models2013intermediate10 min read
Auto-Encoding Variational Bayes
الترميز الذاتي بنهج بايز التبايُني
Kingma, D.P. · Welling, M. — ICLR
The problem
Classical autoencoders compress data into a deterministic code, but the resulting is full of gaps — meaningless regions that produce garbage when decoded. Meanwhile, traditional requires model-specific derivations that don't scale to large datasets or deep neural networks.
The contribution
The : a principled that wraps an inside a probabilistic framework. The outputs a distribution (mean + variance) instead of a fixed vector, and the makes sampling differentiable so the whole system trains with standard . The function — the — naturally balances reconstruction quality against latent-space regularity.
The impact
The VAE established the modern framework for deep generative models. It spawned β-VAE for disentangled representations, inspired the latent-space idea behind diffusion models (DDPM), and remains the backbone of world models in reinforcement learning. Every time an AI generates, interpolates, or imagines — the VAE's probabilistic latent space is the ancestor of that capability.
A regular autoencoder is a photocopy machine: it compresses each document into a filing code and retrieves it perfectly — but ask for a code between two documents and you get a smeared mess.
The VAE replaces the filing cabinet with a cloud of fireflies: each document scatters a soft glow around its location, and the overlapping glows fill the gaps. Reach into any lit region and you pull out a coherent, new document that blends the traits of its neighbors.
The problem: autoencoders memorize, they don't imagine
A standard autoencoder squeezes input through a into a latent code , then reconstructs from . It learns to minimize — and it does so brilliantly. But the latent space it builds has a fatal flaw: the codes are scattered islands with dead ocean between them.
Pick a point between two codes and decode it — you get noise, not a meaningful blend. The autoencoder never learned what "between" means because it was never forced to fill in the gaps. It memorized a lookup table, not a generative model.
What we need is a latent space where every point decodes to something meaningful — a continuous, navigable landscape. The VAE achieves this by making one radical change: instead of encoding each input as a point, encode it as a probability distribution.
The Bayesian foundation: why distributions?
The VAE is rooted in Bayesian thinking. We assume there is a hidden generative process: nature first picks a from a simple (a standard Gaussian), then generates the data from (the ). Given an observed , we want the — "which could have generated this ?" — but computing it requires integrating over all possible values, which is .
The solution: approximate the posterior with a learned encoder that outputs a for each input. This is called variational inference — we replace an impossible exact computation with a trainable approximation.
VAE architecture: encode a cloud, decode a sample
The VAE has two neural networks working in tandem:
The encoder takes an input and outputs two vectors: a mean and a log-variance . Together, they define a Gaussian cloud in latent space — the model's belief about where this input lives.
The decoder takes a single point sampled from that cloud and reconstructs the original input. If is an image, the decoder outputs pixel intensities; if it's continuous data, it may output Gaussian parameters.
The magic: because we encode distributions instead of points, the clouds for similar inputs overlap, filling the latent space with meaningful content everywhere.
The ELBO: one loss, two forces
We cannot maximize the true data likelihood directly — it requires integrating over all . Instead, we maximize a tractable lower bound called the Evidence Lower BOund (ELBO). Think of it as tightening a floor beneath the true objective: pushing the floor up always pushes the objective up too.
The ELBO decomposes into two terms that pull in opposite directions, creating a productive tension:
Imagine two coaches training an athlete. The reconstruction coach demands perfect memory: "reproduce the input exactly!" The KL coach demands simplicity: "keep your representations close to a standard Gaussian — don't over-specialize!" The athlete (the model) must satisfy both, and the sweet spot is a latent space that's both expressive and well-organized.
The reparameterization trick: making randomness differentiable
There's a critical obstacle: the encoder outputs a distribution, and we sample from it. But sampling is not differentiable — you can't backpropagate through a random number generator. Without gradients, the encoder can't learn.
The reparameterization trick solves this elegantly. Instead of sampling directly from , we:
-
Sample noise — this is a fixed, external random source
-
Compute
Now is a deterministic function of and (which we can differentiate) plus an external noise source (which doesn't need gradients). The randomness is still there, but it's been moved outside the . Think of it as a river: we can't steer the rain, but we can build channels that direct where the water flows.
The KL term in closed form
When both the and the prior are Gaussians, the has a beautiful closed-form solution — no sampling required. For a latent space of dimension :
The payoff: a navigable latent space
The KL regularizer's pressure toward the prior doesn't just keep training stable — it gives the latent space two extraordinary properties:
Continuity — nearby points in latent space decode to similar outputs. Walk smoothly from a "3" to an "8" and you see one digit gradually morphing into the other, not a sudden jump.
Completeness — every region of latent space decodes to something plausible. There are no "dead zones" producing garbage. The overlapping Gaussian clouds ensure coverage.
These properties make the VAE's latent space a creative tool: interpolate between data points, generate novel examples, or explore the space to discover what the model has learned.
The same idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
import torch.nn.functional as F
class VAE(nn.Module):
def __init__(self, input_dim=784, latent_dim=20):
super().__init__()
# Encoder: input → hidden → (mu, log_var)
self.fc1 = nn.Linear(input_dim, 400)
self.fc_mu = nn.Linear(400, latent_dim)
self.fc_logvar = nn.Linear(400, latent_dim)
# Decoder: z → hidden → reconstructed input
self.fc3 = nn.Linear(latent_dim, 400)
self.fc4 = nn.Linear(400, input_dim)
def encode(self, x):
h = F.relu(self.fc1(x))
return self.fc_mu(h), self.fc_logvar(h) # two outputs!
def reparameterize(self, mu, log_var):
std = torch.exp(0.5 * log_var) # σ = exp(log_var / 2)
eps = torch.randn_like(std) # ε ~ N(0, I)
return mu + std * eps # z = μ + σ·ε
def decode(self, z):
h = F.relu(self.fc3(z))
return torch.sigmoid(self.fc4(h)) # pixel probabilities
def forward(self, x):
mu, log_var = self.encode(x)
z = self.reparameterize(mu, log_var)
return self.decode(z), mu, log_var
# ELBO loss = reconstruction + KL divergence
def vae_loss(recon_x, x, mu, log_var):
recon = F.binary_cross_entropy(recon_x, x, reduction='sum')
kl = -0.5 * torch.sum(1 + log_var - mu.pow(2) - log_var.exp())
return recon + klTraining: the AEVB algorithm
The Auto-Encoding Variational Bayes (AEVB) algorithm is remarkably simple. For each :
-
Encode: pass each data point through the encoder to get and
-
Sample: draw using the reparameterization trick
-
Decode: reconstruct from
-
Compute the ELBO loss: reconstruction error + KL divergence
-
Backpropagate and update (decoder) and (encoder) simultaneously
That's it. Standard (or ) handles the rest. No EM steps, no MCMC chains, no model-specific derivations. The algorithm scales to millions of data points because it works with mini-batches — exactly like training any other neural network.
Generation: sampling new data
Once trained, generation is trivially simple: sample from the prior and pass it through the decoder. Because the KL term trained the encoder to match the prior, the decoder has seen samples from that distribution during training and knows what to do with them.
This is fundamentally different from a , which needs an adversarial training loop. The VAE generates by : draw from the prior, decode. No , no , no training instabilities.
Connection to EM and variational inference
The VAE's framework connects directly to two classical ideas:
The alternates between inferring latent variables () and updating model parameters (). The VAE does both simultaneously via — the encoder performs a "neural E-step" and the decoder a "neural M-step."
Variational inference traditionally requires hand-derived update equations for each model. The VAE's key insight: let neural networks parameterize both the approximate posterior and the generative model, then optimize everything with backpropagation. This is "" — the encoder network learns to perform inference for any input in a single , rather than running an optimization loop per data point.
Why it mattered
2013
VAE (this paper)
The reparameterization trick enables end-to-end training of deep generative models with latent variables. Published alongside Rezende et al.'s similar approach.
2014
GAN
Goodfellow introduced an adversarial alternative: a generator and discriminator compete. Sharper images but no encoder, no latent inference, and training instabilities.
2017
β-VAE
Scaling the KL weight beyond 1 disentangles latent dimensions — each captures a single factor of variation (pose, color, size). Enabled controllable generation.
2017
VQ-VAE
Replaced Gaussian latent codes with discrete codebook entries. Combined the VAE framework with vector quantization for high-quality image and audio generation.
2020
DDPM (Diffusion Models)
Denoising diffusion works in a latent-space-like framework. The VAE's autoencoder stage later became essential for Latent Diffusion (Stable Diffusion).
2018
World Models
Ha & Schmidhuber used a VAE to compress game frames into a latent representation, then trained an RL agent entirely inside the latent dream world.
2018
Normalizing Flows
Extended the VAE idea with invertible transformations that allow exact likelihood computation — overcoming the "blurry output" limitation.
The VAE's legacy isn't just the model itself — it's the idea that you can train deep generative models with a simple, principled loss function. Every latent-variable model that followed, from diffusion to VQ-VAE to world models, owes a debt to this 2013 paper that showed how to make randomness differentiable.
CitationKingma, Welling. Auto-Encoding Variational Bayes. ICLR, 2014.
Terms in this paper
- Variational Autoencoder (VAE)المرمّز التلقائي المتغير الاحتمالي
- Reparameterization Trickحيلة إعادة البَرمَتَة
- Evidence Lower Bound (ELBO)الحد الأدنى للدليل
- Variational Inferenceالاستدلال الاحتمالي المتغيِّر
- Amortized Inferenceالاستدلال المُعمَّم
- Reconstruction Errorخطأ إعادة البناء
- KL Divergenceتباعد KL
- Approximate Posteriorالاحتمال البعدي التقريبي
- Ancestral Samplingالسحب الأصلي
- Posterior Collapseانهيار الاحتمال البعدي