Generative Models2015advanced10 min read
Variational Inference with Normalizing Flows
الاستدلال المتغيّري بالتدفقات التسوِيَّة
Rezende, D. J. · Mohamed, S. — ICML
The problem
requires choosing an q(z|x). Most methods use factored Gaussians (mean-field), which can only capture unimodal, axis-aligned distributions. When the true is multimodal, skewed, or has complex correlations, the mean-field approximation is fundamentally limited — it cannot match the shape, leading to a loose bound and poor latent representations. Making the approximate posterior richer usually makes inference intractable.
The contribution
Normalizing flows: start from a simple distribution q₀ (e.g. a Gaussian), then apply a chain of K , differentiable transformations f₁, f₂, …, fₖ. Each transformation warps the density, and the change-of-variables formula — involving the determinant of each step — gives the exact density of the transformed variable. The paper introduces two families with O(D) Jacobian cost: planar flows (slice-and-warp along hyperplanes) and radial flows (expand/contract around a point). These are plugged into the 's ELBO, making the bound tighter without sacrificing scalability. The paper also connects finite flows to infinitesimal flows (Langevin and Hamiltonian dynamics), providing a unified theoretical view.
The impact
Normalizing flows became a foundational tool in generative modeling and probabilistic inference. They inspired a family of powerful descendants: NICE, RealNVP, Glow (invertible 1×1 convolutions for image generation), Inverse Autoregressive Flows (IAF), Masked Autoregressive Flows (MAF), and continuous normalizing flows via Neural ODEs (FFJORD). The change-of-variables perspective also deeply influenced score-based generative models and , which power state-of-the-art image and video generation today.
Imagine a sculptor who can only work with a single ball of clay. A standard VAE insists that ball must stay perfectly spherical — you can move it and resize it, but never reshape it. If the true shape you need is a crescent or a star, tough luck: the sphere is all you get.
Normalizing flows hand the sculptor a sequence of molds. Each mold applies one simple deformation — a stretch here, a squeeze there — and because each mold is reversible (you can always undo it), you can track exactly how the volume changed at every step. After enough molds, the original sphere can become virtually any shape: a crescent, a star, a pretzel. The key insight: you never lose track of the density because each deformation records its own volume change via the Jacobian determinant.
The problem: Gaussian posteriors cannot capture complex shapes
In a Variational Autoencoder, we want to infer the latent variables z that generated an observation x. The true posterior p(z|x) — the distribution of latent codes that could have produced x — is almost always intractable. So we approximate it with a simpler distribution q(z|x), typically a diagonal Gaussian parameterized by an network.
The trouble: a diagonal Gaussian is an axis-aligned ellipsoid. It cannot represent correlations between dimensions, multiple modes, or curved manifolds. When the true posterior bends or splits, the Gaussian just sits in the middle and averages — like trying to cover a banana with a circular sticker.
This mismatch has real consequences: the ELBO bound becomes loose, the model under-uses its latent space (), and generated samples lack diversity. What we need is a way to make q(z|x) flexible enough to match the true posterior, without making the density intractable.
The key tool: the change of variables formula
Normalizing flows rest on one theorem from probability: if you transform a random variable z with an invertible, differentiable function f, the density of the result z' = f(z) is given by the formula. The idea is physically intuitive: if a transformation stretches a region of space (making it bigger), the probability density in that region must decrease proportionally — because the same probability mass is now spread over a larger volume. The Jacobian determinant measures exactly this local volume change.
Think of it like a balloon animal: as you twist one section thinner, the air (probability) concentrates there, making it denser. The Jacobian determinant is the mathematical record of every twist and squeeze, so you always know exactly how dense the air is at any point.
Chaining transformations: from simple to complex
A single invertible transformation can only do so much. The power comes from chaining: apply f₁, then f₂, …, then fₖ. Each step warps the density a little more, and the total density change is the product of all the individual Jacobian determinants — or equivalently, the sum of their logs.
This is the "flow" in normalizing flows: the initial density flows through a pipeline of invertible maps, gradually morphing from a simple Gaussian into a rich, expressive distribution. The "normalizing" part refers to the fact that the process can go in reverse: any complex distribution can be normalized back to a Gaussian by inverting the chain.
Two flow families: planar and radial
The central engineering challenge: computing a D×D Jacobian determinant normally costs O(D³). To make flows practical, we need transformations whose Jacobians have special structure. The paper introduces two families:
Planar flows slice the space with a hyperplane and warp the density on one side. Each transformation has the form f(z) = z + u·h(wᵀz + b), where u and w are D-dimensional vectors, b is a scalar, and h is a smooth nonlinearity like tanh. Thanks to the matrix determinant lemma, the Jacobian determinant reduces to a scalar: |1 + uᵀ·h'(wᵀz + b)·w| — computable in O(D).
Radial flows contract or expand the density around a reference point z₀. The form is f(z) = z + β/(α + r)·(z − z₀), where r = ‖z − z₀‖. This creates localized bumps or dips in the density, also with O(D) Jacobian cost.
Neither family alone is very expressive, but stacking many of them — mixing planar and radial — can approximate surprisingly complex distributions.
Flows inside the VAE: a tighter ELBO
The practical payoff: replace the VAE's simple Gaussian q(z|x) with a . The encoder still outputs a Gaussian z₀ ~ N(μ, σ²), but then z₀ is passed through K flow steps to produce zₖ. The ELBO becomes:
The beauty of this formulation: we only need to sample z₀ from the simple Gaussian (using the ), run it forward through the flow, and sum up the log-Jacobian terms. No need to invert the flow during . The flow parameters can be made data-dependent — the encoder network outputs not just μ and σ but also the flow parameters for each input x — achieving with a rich posterior.
Beyond finite steps: infinitesimal flows
The paper makes a deep theoretical contribution by connecting finite flows to continuous-time dynamics. Instead of K discrete steps, imagine an infinite number of infinitesimally small steps — a smooth path from z₀ to zₖ governed by a differential equation dz/dt = f(z, t).
Under this view, the log-density evolves according to the instantaneous change of variables formula, which replaces the discrete log-det-Jacobian sum with an integral of the divergence of the velocity field. This connects normalizing flows to two powerful families of dynamics:
Langevin flow — gradient of the log-density plus Gaussian noise: follows the energy landscape toward regions of high posterior probability, like a ball rolling downhill in a noisy landscape.
Hamiltonian flow — introduces momentum variables that let the flow traverse complex posteriors efficiently, like a frictionless puck sliding along the posterior surface, conserving energy and covering more ground than pure gradient descent.
This continuous view later inspired Neural ODEs (Chen et al., 2018) and FFJORD (Grathwohl et al., 2018), where the flow is parameterized by a neural network and solved with an ODE integrator.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def planar_flow(z, u, w, b):
"""One planar flow step: f(z) = z + u * h(w^T z + b)"""
h = np.tanh(np.dot(w, z) + b) # scalar activation
return z + u * h # warp z along direction u
def log_det_jacobian_planar(z, u, w, b):
"""Log |det df/dz| for a planar flow — O(D) cost."""
h_prime = 1 - np.tanh(np.dot(w, z) + b)**2 # tanh derivative
psi = h_prime * w # ψ(z) = h'(w^T z + b) * w
return np.log(np.abs(1 + np.dot(u, psi))) # matrix determinant lemma
def apply_flow_chain(z0, flow_params):
"""Apply K planar flows and accumulate log-density change."""
z = z0
log_det_sum = 0.0
for (u, w, b) in flow_params:
log_det_sum += log_det_jacobian_planar(z, u, w, b)
z = planar_flow(z, u, w, b)
return z, log_det_sum
# In a VAE with flows:
# 1. Encoder outputs μ, σ, and flow params {u_k, w_k, b_k}
# 2. Sample z₀ ~ N(μ, σ²I) via reparameterization
# 3. zₖ, Σ log|det| = apply_flow_chain(z₀, flow_params)
# 4. ELBO = E[log p(x|zₖ)] - KL₀ + Σ log|det|
# (the log-det terms make the bound tighter)Why it mattered
2014
NICE (Dinh et al.)
Non-linear Independent Components Estimation — volume-preserving coupling layers with trivial Jacobian determinant (= 1). First practical deep invertible architecture.
2015
This paper (Rezende & Mohamed)
Introduced the normalizing flows framework for variational inference with planar and radial flows. Connected finite flows to Langevin and Hamiltonian dynamics.
2016
RealNVP (Dinh et al.)
Affine coupling layers with learnable scale — invertible by design, tractable Jacobians, and the first flow model to generate convincing images.
2016
IAF (Kingma et al.)
Inverse Autoregressive Flow — used autoregressive structure for fast sampling with triangular Jacobians. Directly improved VAE posterior quality.
2018
Glow (Kingma & Dhariwal)
Invertible 1×1 convolutions + multi-scale architecture. Generated high-resolution faces with exact likelihood computation. Showed flows can compete with GANs on image quality.
2018
Neural ODEs & FFJORD
Took the infinitesimal flow idea to its conclusion: parameterize the velocity field with a neural network, solve with an ODE integrator. FFJORD used Hutchinson's trace estimator for O(D) cost.
2020
Score-based generative models
Connected diffusion processes with score matching and continuous normalizing flows. Showed that learning the score function enables generation without explicit density.
2022
Flow Matching (Lipman et al.)
Unified view connecting continuous normalizing flows with optimal transport. Simpler training than score-based models, now powering state-of-the-art image generators.
CitationRezende, D. J. & Mohamed, S.. Variational Inference with Normalizing Flows. ICML, 2015.
Terms in this paper
- Normalizing Flowالتدفق التسوِيّ
- Variational Inferenceالاستدلال الاحتمالي المتغيِّر
- Change of Variablesتغيير المتغيرات
- Invertibleعكوس
- Jacobianمصفوفة جاكوبي
- Planar Flowالتدفق المستوي
- Radial Flowالتدفق الشعاعي
- Evidence Lower Bound (ELBO)الحد الأدنى للدليل
- Approximate Posteriorالاحتمال البعدي التقريبي
- Latent Variableالمتغير الكامن
- Density Estimationتقدير الكثافة
- Reparameterization Trickحيلة إعادة البَرمَتَة
- Amortized Inferenceالاستدلال المُعمَّم
- KL Divergenceتباعد KL