Generative Models2015advanced10 min read

Variational Inference with Normalizing Flows

الاستدلال المتغيّري بالتدفقات التسوِيَّة

Rezende, D. J. · Mohamed, S. — ICML

The problem

requires choosing an q(z|x). Most methods use factored Gaussians (mean-field), which can only capture unimodal, axis-aligned distributions. When the true is multimodal, skewed, or has complex correlations, the mean-field approximation is fundamentally limited — it cannot match the shape, leading to a loose bound and poor latent representations. Making the approximate posterior richer usually makes inference intractable.

The contribution

Normalizing flows: start from a simple distribution q₀ (e.g. a Gaussian), then apply a chain of K , differentiable transformations f₁, f₂, …, fₖ. Each transformation warps the density, and the change-of-variables formula — involving the determinant of each step — gives the exact density of the transformed variable. The paper introduces two families with O(D) Jacobian cost: planar flows (slice-and-warp along hyperplanes) and radial flows (expand/contract around a point). These are plugged into the 's ELBO, making the bound tighter without sacrificing scalability. The paper also connects finite flows to infinitesimal flows (Langevin and Hamiltonian dynamics), providing a unified theoretical view.

The impact

Normalizing flows became a foundational tool in generative modeling and probabilistic inference. They inspired a family of powerful descendants: NICE, RealNVP, Glow (invertible 1×1 convolutions for image generation), Inverse Autoregressive Flows (IAF), Masked Autoregressive Flows (MAF), and continuous normalizing flows via Neural ODEs (FFJORD). The change-of-variables perspective also deeply influenced score-based generative models and , which power state-of-the-art image and video generation today.

Imagine a sculptor who can only work with a single ball of clay. A standard VAE insists that ball must stay perfectly spherical — you can move it and resize it, but never reshape it. If the true shape you need is a crescent or a star, tough luck: the sphere is all you get.

Normalizing flows hand the sculptor a sequence of molds. Each mold applies one simple deformation — a stretch here, a squeeze there — and because each mold is reversible (you can always undo it), you can track exactly how the volume changed at every step. After enough molds, the original sphere can become virtually any shape: a crescent, a star, a pretzel. The key insight: you never lose track of the density because each deformation records its own volume change via the Jacobian determinant.

The problem: Gaussian posteriors cannot capture complex shapes

In a Variational Autoencoder, we want to infer the latent variables z that generated an observation x. The true posterior p(z|x) — the distribution of latent codes that could have produced x — is almost always intractable. So we approximate it with a simpler distribution q(z|x), typically a diagonal Gaussian parameterized by an network.

The trouble: a diagonal Gaussian is an axis-aligned ellipsoid. It cannot represent correlations between dimensions, multiple modes, or curved manifolds. When the true posterior bends or splits, the Gaussian just sits in the middle and averages — like trying to cover a banana with a circular sticker.

This mismatch has real consequences: the ELBO bound becomes loose, the model under-uses its latent space (), and generated samples lack diversity. What we need is a way to make q(z|x) flexible enough to match the true posterior, without making the density intractable.

Open in Lab
Left: the true posterior is banana-shaped. Right: a Gaussian approximation can only cover part of it. Toggle the flow to see how successive transformations close the gap.
The demo wakes as you arrive…

The key tool: the change of variables formula

Normalizing flows rest on one theorem from probability: if you transform a random variable z with an invertible, differentiable function f, the density of the result z' = f(z) is given by the formula. The idea is physically intuitive: if a transformation stretches a region of space (making it bigger), the probability density in that region must decrease proportionally — because the same probability mass is now spread over a larger volume. The Jacobian determinant measures exactly this local volume change.

q(z′)=q(z)∣det⁡∂f∂z∣−1q(\mathbf{z}') = q(\mathbf{z}) \left| \det \frac{\partial f}{\partial \mathbf{z}} \right|^{-1}
Change of variables — the engine of normalizing flows — q(z) = density before the transform · f = invertible function · the Jacobian determinant measures how much f stretches or compresses volume · the inverse (⁻¹) adjusts density downward where space expands, upward where it contracts

Think of it like a balloon animal: as you twist one section thinner, the air (probability) concentrates there, making it denser. The Jacobian determinant is the mathematical record of every twist and squeeze, so you always know exactly how dense the air is at any point.

Open in Lab
Drag the transformation slider to see how an invertible map warps a 2D Gaussian. The grid shows volume changes; color intensity tracks density.
The demo wakes as you arrive…

Chaining transformations: from simple to complex

A single invertible transformation can only do so much. The power comes from chaining: apply f₁, then f₂, …, then fₖ. Each step warps the density a little more, and the total density change is the product of all the individual Jacobian determinants — or equivalently, the sum of their logs.

This is the "flow" in normalizing flows: the initial density flows through a pipeline of invertible maps, gradually morphing from a simple Gaussian into a rich, expressive distribution. The "normalizing" part refers to the fact that the process can go in reverse: any complex distribution can be normalized back to a Gaussian by inverting the chain.

ln⁡qK(zK)=ln⁡q0(z0)−∑k=1Kln⁡∣det⁡∂fk∂zk−1∣\ln q_K(\mathbf{z}_K) = \ln q_0(\mathbf{z}_0) - \sum_{k=1}^{K} \ln \left| \det \frac{\partial f_k}{\partial \mathbf{z}_{k-1}} \right|
Log-density after K flow steps — Start with ln q₀ (the simple base density), then subtract the sum of log-Jacobian determinants — one per transformation. Each term tracks how much that step expanded or compressed the local volume.
Open in Lab
Watch a Gaussian flow through 1, 2, 4, 8, and 16 transformations. Each step adds a small warp; together they build rich structure.
The demo wakes as you arrive…

Two flow families: planar and radial

The central engineering challenge: computing a D×D Jacobian determinant normally costs O(D³). To make flows practical, we need transformations whose Jacobians have special structure. The paper introduces two families:

Planar flows slice the space with a hyperplane and warp the density on one side. Each transformation has the form f(z) = z + u·h(wᵀz + b), where u and w are D-dimensional vectors, b is a scalar, and h is a smooth nonlinearity like tanh. Thanks to the matrix determinant lemma, the Jacobian determinant reduces to a scalar: |1 + uᵀ·h'(wᵀz + b)·w| — computable in O(D).

Radial flows contract or expand the density around a reference point z₀. The form is f(z) = z + β/(α + r)·(z − z₀), where r = ‖z − z₀‖. This creates localized bumps or dips in the density, also with O(D) Jacobian cost.

Neither family alone is very expressive, but stacking many of them — mixing planar and radial — can approximate surprisingly complex distributions.

f(z)=z+u h(w⊤z+b),∣det⁡∂f∂z∣=∣1+u⊤ψ(z)∣f(\mathbf{z}) = \mathbf{z} + \mathbf{u}\, h(\mathbf{w}^\top \mathbf{z} + b), \qquad \left| \det \frac{\partial f}{\partial \mathbf{z}} \right| = \left| 1 + \mathbf{u}^\top \psi(\mathbf{z}) \right|
Planar flow and its O(D) Jacobian determinant — ψ(z) = h'(wᵀz + b)·w · The hyperplane wᵀz + b = 0 splits the space · u controls the direction and magnitude of the warp · h (e.g. tanh) makes it smooth and bounded
f(z)=z+βα+∥z−z0∥(z−z0)f(\mathbf{z}) = \mathbf{z} + \frac{\beta}{\alpha + \|\mathbf{z} - \mathbf{z}_0\|} (\mathbf{z} - \mathbf{z}_0)
Radial flow — expand or contract around a center point — z₀ = center of the transformation · α > 0 controls spread · β controls direction: positive β pushes density outward, negative pulls inward
Open in Lab
Switch between planar and radial flows. Adjust the parameters to see how each type warps a Gaussian density.
The demo wakes as you arrive…

Flows inside the VAE: a tighter ELBO

The practical payoff: replace the VAE's simple Gaussian q(z|x) with a . The encoder still outputs a Gaussian z₀ ~ N(μ, σ²), but then z₀ is passed through K flow steps to produce zₖ. The ELBO becomes:

F(x)=Eq0(z0)[ln⁡q0(z0)−ln⁡p(x,zK)−∑k=1Kln⁡∣det⁡∂fk∂zk−1∣]\mathcal{F}(\mathbf{x}) = \mathbb{E}_{q_0(\mathbf{z}_0)} \Big[ \ln q_0(\mathbf{z}_0) - \ln p(\mathbf{x}, \mathbf{z}_K) - \sum_{k=1}^{K} \ln \left| \det \frac{\partial f_k}{\partial \mathbf{z}_{k-1}} \right| \Big]
Free energy bound with normalizing flows — q₀(z₀) = initial Gaussian density · p(x, zₖ) = joint model probability evaluated at the final flow output · the sum of log-det-Jacobians adjusts for density changes through the flow · minimizing F tightens the ELBO

The beauty of this formulation: we only need to sample z₀ from the simple Gaussian (using the ), run it forward through the flow, and sum up the log-Jacobian terms. No need to invert the flow during . The flow parameters can be made data-dependent — the encoder network outputs not just μ and σ but also the flow parameters for each input x — achieving with a rich posterior.

Open in Lab
Click each stage of the VAE+Flow pipeline to see what happens to z and the density.
The demo wakes as you arrive…

Beyond finite steps: infinitesimal flows

The paper makes a deep theoretical contribution by connecting finite flows to continuous-time dynamics. Instead of K discrete steps, imagine an infinite number of infinitesimally small steps — a smooth path from z₀ to zₖ governed by a differential equation dz/dt = f(z, t).

Under this view, the log-density evolves according to the instantaneous change of variables formula, which replaces the discrete log-det-Jacobian sum with an integral of the divergence of the velocity field. This connects normalizing flows to two powerful families of dynamics:

Langevin flow — gradient of the log-density plus Gaussian noise: follows the energy landscape toward regions of high posterior probability, like a ball rolling downhill in a noisy landscape.

Hamiltonian flow — introduces momentum variables that let the flow traverse complex posteriors efficiently, like a frictionless puck sliding along the posterior surface, conserving energy and covering more ground than pure gradient descent.

This continuous view later inspired Neural ODEs (Chen et al., 2018) and FFJORD (Grathwohl et al., 2018), where the flow is parameterized by a neural network and solved with an ODE integrator.

∂ln⁡q(z(t))∂t=−tr(∂f∂z(t))\frac{\partial \ln q(\mathbf{z}(t))}{\partial t} = -\text{tr}\left(\frac{\partial f}{\partial \mathbf{z}(t)}\right)
Instantaneous change of variables (continuous-time flows) — The log-density evolves as the negative trace (not determinant!) of the Jacobian — much cheaper to compute than a full determinant, scaling as O(D) instead of O(D³)

The idea in code

Planar flow and log-density computationpython

Simplified to show the idea — not the real implementation.

import numpy as np

def planar_flow(z, u, w, b):
    """One planar flow step: f(z) = z + u * h(w^T z + b)"""
    h = np.tanh(np.dot(w, z) + b)           # scalar activation
    return z + u * h                          # warp z along direction u

def log_det_jacobian_planar(z, u, w, b):
    """Log |det df/dz| for a planar flow — O(D) cost."""
    h_prime = 1 - np.tanh(np.dot(w, z) + b)**2   # tanh derivative
    psi = h_prime * w                              # ψ(z) = h'(w^T z + b) * w
    return np.log(np.abs(1 + np.dot(u, psi)))      # matrix determinant lemma

def apply_flow_chain(z0, flow_params):
    """Apply K planar flows and accumulate log-density change."""
    z = z0
    log_det_sum = 0.0
    for (u, w, b) in flow_params:
        log_det_sum += log_det_jacobian_planar(z, u, w, b)
        z = planar_flow(z, u, w, b)
    return z, log_det_sum

# In a VAE with flows:
# 1. Encoder outputs μ, σ, and flow params {u_k, w_k, b_k}
# 2. Sample z₀ ~ N(μ, σ²I) via reparameterization
# 3. zₖ, Σ log|det| = apply_flow_chain(z₀, flow_params)
# 4. ELBO = E[log p(x|zₖ)] - KL₀ + Σ log|det|
#    (the log-det terms make the bound tighter)

Why it mattered

  1. 2014

    NICE (Dinh et al.)

    Non-linear Independent Components Estimation — volume-preserving coupling layers with trivial Jacobian determinant (= 1). First practical deep invertible architecture.

  2. 2015

    This paper (Rezende & Mohamed)

    Introduced the normalizing flows framework for variational inference with planar and radial flows. Connected finite flows to Langevin and Hamiltonian dynamics.

  3. 2016

    RealNVP (Dinh et al.)

    Affine coupling layers with learnable scale — invertible by design, tractable Jacobians, and the first flow model to generate convincing images.

  4. 2016

    IAF (Kingma et al.)

    Inverse Autoregressive Flow — used autoregressive structure for fast sampling with triangular Jacobians. Directly improved VAE posterior quality.

  5. 2018

    Glow (Kingma & Dhariwal)

    Invertible 1×1 convolutions + multi-scale architecture. Generated high-resolution faces with exact likelihood computation. Showed flows can compete with GANs on image quality.

  6. 2018

    Neural ODEs & FFJORD

    Took the infinitesimal flow idea to its conclusion: parameterize the velocity field with a neural network, solve with an ODE integrator. FFJORD used Hutchinson's trace estimator for O(D) cost.

  7. 2020

    Score-based generative models

    Connected diffusion processes with score matching and continuous normalizing flows. Showed that learning the score function enables generation without explicit density.

  8. 2022

    Flow Matching (Lipman et al.)

    Unified view connecting continuous normalizing flows with optimal transport. Simpler training than score-based models, now powering state-of-the-art image generators.

CitationRezende, D. J. & Mohamed, S.. Variational Inference with Normalizing Flows. ICML, 2015.

Terms in this paper