Generative Models2023advanced11 min read

Flow Matching for Generative Modeling

مُطابقة التدفق للنمذجة التوليدية

Lipman, Y. · Chen, R. T. Q. · Ben-Hamu, H. · Nickel, M. · Le, M. — ICLR

The problem

Continuous Normalizing Flows (CNFs) are a powerful class of generative models that transform a simple noise distribution into a data distribution via a learned . However, training CNFs by maximum requires simulating the ODE at every training step — an expensive operation that scales poorly to large datasets and high-resolution images. Meanwhile, diffusion models scale well but are locked into specific curved probability paths (VP, VE schedules) that require many steps, and their training objectives are tied to rather than direct vector field . There was no efficient, -free method to train CNFs with arbitrary probability paths.

The contribution

The paper introduces (FM), a simulation-free training paradigm for Continuous Normalizing Flows. The key theoretical insight is (CFM): instead of regressing the intractable marginal vector field, the model regresses conditional vector fields anchored to individual data points — and the authors prove this yields identical gradients. FM works with any Gaussian probability path, unifying diffusion paths as special cases. Crucially, it enables paths — straight-line trajectories that are faster to train, faster to sample, and produce better results than curved diffusion paths.

The impact

Flow Matching has become the default training paradigm for continuous generative models, replacing both score matching and simulation-based CNF training. It is the foundation of Stable Diffusion 3, OpenAI's Sora video generation model, and Meta's Voicebox speech synthesis. In robotics, Physical Intelligence's π₀ policy uses Flow Matching for action generation. The paper unified previously disjoint communities — CNFs, diffusion models, and optimal transport — under a single framework, and its OT paths set the standard for efficient sampling in modern generative AI.

Imagine a choreographer arranging a flash mob. A thousand dancers start scattered randomly across a plaza. The choreographer's job is to give each dancer a direction and speed — a tiny arrow — at every moment, so they all flow smoothly into a perfect formation. The choreographer never simulates the whole dance in advance; she just watches where each dancer needs to end up and draws the arrow pointing straight there.

Previous methods (diffusion models) told dancers to wander in spirals before reaching their spot. This paper says: why not walk in straight lines? The dancers arrive faster, the choreography is simpler, and the formation is sharper.

That's Flow Matching — learning the arrows (velocity fields) that transport noise into data along the most direct path.

The problem: simulation is expensive, diffusion is curved

A learns a that pushes a simple distribution (like a Gaussian) into a complex data distribution by solving an ordinary differential equation. Think of it as a river current: every point in space has an arrow telling it where to flow, and following those arrows from t=0t{=}0 to t=1t{=}1 transforms noise into data.

The problem with training CNFs before this paper had two faces. First, maximum likelihood training required simulating the entire ODE at each training step to compute the change in log-density — this involves the trace of the velocity field, which is computationally brutal for high-dimensional data. Second, while diffusion models found simulation-free shortcuts via score matching, those shortcuts locked the probability path into specific curved trajectories (VP, VE schedules) that needed hundreds of ODE steps to sample accurately.

The field needed a training method that was both simulation-free (like score matching) and path-flexible (unlike score matching). Flow Matching is that method.

Open in Lab
Compare diffusion paths (curved, slow) with Optimal Transport paths (straight, fast). Toggle between them to see how particles travel from noise to data.
The demo wakes as you arrive…

The core idea: regress the velocity, skip the simulation

The central idea of Flow Matching is disarmingly simple. Instead of training a CNF by simulating it (which is expensive), or by matching scores (which constrains the path), we directly regress the velocity field. We define a target probability path ptp_t from noise (t=0t{=}0) to data (t=1t{=}1) and a target vector field utu_t that generates this path. Then we train a neural network vθv_\theta to predict utu_t everywhere.

The ideal training loss is the Flow Matching objective — a simple between the predicted and target velocity fields:

LFM(θ)=Et∼U[0,1], x∼pt(x)∥vθ(t,x)−ut(x)∥2\mathcal{L}_{\text{FM}}(\theta) = \mathbb{E}_{t \sim \mathcal{U}[0,1],\, x \sim p_t(x)} \left\| v_\theta(t, x) - u_t(x) \right\|^2
Flow Matching loss — direct vector field regression — At each training step: sample a random time tt, sample a point xx from the probability path at that time, and minimize the squared error between the network's predicted velocity vθ(t,x)v_\theta(t, x) and the true target velocity ut(x)u_t(x). No ODE simulation, no Jacobian computation — just a regression loss.

There is one catch: the marginal velocity field ut(x)u_t(x) is intractable. Computing it requires integrating over the entire data distribution, which is the very thing we are trying to learn. This is where the paper's theoretical breakthrough comes in.

The trick: conditional flow matching

The insight that makes everything work: instead of regressing the intractable marginal vector field ut(x)u_t(x), regress the conditional vector field ut(x∣x1)u_t(x | x_1) — the velocity that transports noise to a single data point x1x_1. The Conditional Flow Matching (CFM) loss is:

LCFM(θ)=Et, q(x1), pt(x∣x1)∥vθ(t,x)−ut(x∣x1)∥2\mathcal{L}_{\text{CFM}}(\theta) = \mathbb{E}_{t,\, q(x_1),\, p_t(x|x_1)} \left\| v_\theta(t, x) - u_t(x | x_1) \right\|^2
Conditional Flow Matching loss — the tractable surrogate — Sample a data point x1x_1 from the data distribution qq, sample a noisy point xx from the conditional path pt(x∣x1)p_t(x | x_1), and regress the network's prediction onto the *conditional* target velocity ut(x∣x1)u_t(x | x_1). Each training sample only needs one data point — no integration over the full dataset.

Why does this work? The paper proves (Theorem 2) that the FM loss and the CFM loss have identical gradients with respect to the network parameters θ\theta. This is the same type of insight behind denoising score matching: individual conditional targets may point in different directions (each one pulls toward a single data point), but their average over all data points equals the true marginal direction.

Think of it like surveying a crowd: asking every person "which way is the exit?" gives noisy individual answers, but averaging all fingers yields the correct direction. Each conditional velocity field is one finger; the marginal field is the average.

Open in Lab
Each arrow shows a conditional velocity field pointing to one data point. Toggle to see how they average into the smooth marginal field the network actually learns.
The demo wakes as you arrive…

Gaussian probability paths: a unified recipe

Flow Matching works with any probability path, but the paper focuses on Gaussian conditional paths — because they yield closed-form velocity targets that make training trivial. A Gaussian conditional path takes the form:

pt(x∣x1)=N(x∣μt(x1), σt(x1)2I)p_t(x | x_1) = \mathcal{N}(x \mid \mu_t(x_1),\, \sigma_t(x_1)^2 I)

where μt\mu_t and σt\sigma_t are time-dependent schedules controlling the mean and standard deviation. At t=0t{=}0, the path should be pure noise (μ0=0\mu_0 = 0, σ0=1\sigma_0 = 1). At t=1t{=}1, it should concentrate on the data point (μ1=x1\mu_1 = x_1, σ1≈0\sigma_1 \approx 0). Everything in between is a smooth .

The paper proves (Theorem 3) that for any such Gaussian path, the conditional velocity field has a simple closed form:

ut(x∣x1)=σt′σt(x−μt)+μt′u_t(x | x_1) = \frac{\sigma'_t}{\sigma_t}(x - \mu_t) + \mu'_t
Closed-form conditional velocity for Gaussian paths — The velocity at any point xx along the Gaussian path has two components: the first term scales the deviation from the mean (expanding or contracting the distribution), and the second term shifts the center toward the data point. The primes denote time derivatives: μt′=dμt/dt\mu'_t = d\mu_t / dt and σt′=dσt/dt\sigma'_t = d\sigma_t / dt.

The power move: Optimal Transport paths

With the general framework in place, the authors propose a specific path choice that is dramatically better than diffusion paths: Optimal Transport (OT) displacement interpolation. The OT path uses the simplest possible schedules:

μt(x1)=t⋅x1,σt=1−(1−σmin⁡)t\mu_t(x_1) = t \cdot x_1, \qquad \sigma_t = 1 - (1 - \sigma_{\min})t

This means each sample follows a straight line from its noise origin to its data destination. No curves, no spirals, no detours. Plugging these schedules into the Gaussian velocity formula gives the OT conditional velocity:

ut(x∣x1)=x1−(1−σmin⁡)x1−(1−σmin⁡)tu_t(x | x_1) = \frac{x_1 - (1 - \sigma_{\min})x}{1 - (1 - \sigma_{\min})t}
OT conditional velocity — the straight-line target — At every time tt, the velocity simply points from the current position xx toward the data point x1x_1, scaled by a time-dependent factor. As t→1t \to 1, the factor grows, pushing harder toward the destination — like a GPS recalculation that becomes more insistent as you approach the target.
Open in Lab
Watch how OT paths (straight lines) transport particles from noise to data with constant speed. Drag the time slider to follow the flow.
The demo wakes as you arrive…

Why are straight paths better? Three reasons:

Faster sampling. A straight trajectory is easier for numerical ODE solvers to follow — even a basic Euler integrator can produce high-quality samples in ∼\sim100 function evaluations, compared to ∼\sim250+ for diffusion paths.

Easier learning. When the velocity field is smooth and nearly constant along each trajectory, the neural network has an easier regression target. Diffusion paths create rapidly varying velocity fields near t=0t{=}0 and t=1t{=}1 that are harder to fit.

Better generalization. Empirically, OT paths consistently produce lower scores and better log-likelihoods than diffusion paths under the same training budget.

The training recipe: five lines of pseudocode

One of the most striking features of Flow Matching is how simple the becomes. The entire algorithm reduces to:

  1. Sample a data point x1∼q(x1)x_1 \sim q(x_1) and noise x0∼N(0,I)x_0 \sim \mathcal{N}(0, I)

  2. Sample a random time t∼U[0,1]t \sim \mathcal{U}[0, 1]

  3. Interpolate: xt=(1−(1−σmin⁡)t) x0+t x1x_t = (1 - (1 - \sigma_{\min})t)\, x_0 + t\, x_1

  4. Compute the target velocity: ut=x1−(1−σmin⁡) x0u_t = x_1 - (1 - \sigma_{\min})\, x_0

  5. Minimize ∥vθ(t,xt)−ut∥2\| v_\theta(t, x_t) - u_t \|^2

That is the entire training loop. No score functions, no lookups, no ODE simulation. Just sample, interpolate, regress.

Flow Matching training step (OT path)python

Simplified to show the idea — not the real implementation.

# Flow Matching with OT paths — one training step
sigma_min = 1e-5
t = torch.rand(batch_size, 1)                    # random time
x0 = torch.randn_like(x1)                        # noise sample
# Interpolate along straight line from noise to data
x_t = (1 - (1 - sigma_min) * t) * x0 + t * x1
# Target velocity: direction of the straight line
u_t = x1 - (1 - sigma_min) * x0
# Train the network to predict this velocity
loss = ((model(t, x_t) - u_t) ** 2).mean()
loss.backward()

Sampling: just follow the arrows

Once the velocity field vθv_\theta is trained, generating samples is pure ODE integration. Start from random noise x0∼N(0,I)x_0 \sim \mathcal{N}(0, I) and solve:

ddtϕt(x)=vθ(t,ϕt(x)),ϕ0(x)=x0\frac{d}{dt} \phi_t(x) = v_\theta(t, \phi_t(x)), \quad \phi_0(x) = x_0

Any standard works — dopri5 (adaptive Runge-Kutta) for maximum quality, or a simple Euler method for fast generation with fewer steps. Because OT paths produce nearly straight trajectories, even crude solvers give good results.

For computing exact likelihoods, the model augments the ODE with the formula, tracking the log-determinant of the Jacobian as the flow progresses:

log⁡p1(x1)=log⁡p0(x0)−∫01∇⋅vθ(t,ϕt(x)) dt\log p_1(x_1) = \log p_0(x_0) - \int_0^1 \nabla \cdot v_\theta(t, \phi_t(x))\, dt
Exact log-likelihood via continuous change of variables — The log-probability of a data point under the model equals the log-probability of its noise origin (known, since it is Gaussian) minus the accumulated divergence of the velocity field along the flow trajectory. The divergence measures how much the flow expands or compresses volume at each point.
Open in Lab
Click "Sample" to generate new data points by integrating the learned velocity field from noise. Watch how particles flow from random positions to form the target distribution.
The demo wakes as you arrive…

Results: better images, faster training, fewer steps

The authors trained Flow Matching on CIFAR-10 and ImageNet (32×3232{\times}32, 64×6464{\times}64, 128×128128{\times}128) using the same architecture as prior diffusion models. The results are consistent across every setting:

FM with OT paths beats diffusion baselines on both FID and likelihood. On CIFAR-10, FM-OT achieves FID 6.35 vs DDPM's 7.48, and NLL 2.99 vs 3.12 bits per dimension. On ImageNet 64×6464{\times}64, FM-OT reaches FID 14.45 vs ScoreFlow's 24.95.

FM with diffusion paths is already more stable than standard diffusion training. Even without switching to OT, the FM training objective provides smoother loss curves and faster convergence than DDPM or score matching on the same diffusion paths.

OT paths require far fewer sampling steps. FM-OT achieves comparable quality with 142 NFE (number of function evaluations) on CIFAR-10, versus 274 for DDPM. On ImageNet 64×6464{\times}64, the gap widens to 138 vs 601.

Open in Lab
Compare FID scores and sampling steps (NFE) across methods. FM with OT paths consistently wins on both axes.
The demo wakes as you arrive…

Legacy: the engine behind modern generative AI

Flow Matching did not just offer a better training method — it changed how the field thinks about generative modeling. By proving that diffusion is one instance of a broader family of probability paths, it freed researchers to explore new geometries, new interpolation strategies, and new domains far beyond images.

The impact has been sweeping and rapid. Sora (OpenAI, 2024) uses flow matching for video generation. Stable Diffusion 3 (Stability AI, 2024) replaced its diffusion backbone with flow matching. Meta's Voicebox and AudioBox generate speech and sound effects via flow matching. In robotics, π₀ (Physical Intelligence, 2024) uses flow matching to generate continuous robotic actions.

Concurrent work at ICLR 2023 — Rectified Flow (Liu et al.) and Stochastic Interpolants (Albergo & Vanden-Eijnden) — independently arrived at structurally similar frameworks, confirming that this was an idea whose time had come. Together, these papers marked the transition from "diffusion models" to "flow-based generative models" as the dominant paradigm.

  1. 2018

    Neural ODE (Chen et al.)

    Introduced Continuous Normalizing Flows trained via the adjoint sensitivity method. Powerful but required expensive ODE simulation during training.

  2. 2020

    Score-based diffusion (Song et al.)

    Unified score matching and diffusion via SDEs, enabling scalable training but constraining the probability path to VP/VE schedules.

  3. 2023

    Flow Matching (this paper, Lipman et al.)

    Simulation-free CNF training via conditional vector field regression. Introduced OT paths for straight-line transport, outperforming diffusion baselines.

  4. 2023

    Concurrent — Rectified Flow & Stochastic Interpolants

    Liu et al. and Albergo & Vanden-Eijnden independently proposed similar simulation-free flow training, confirming the paradigm from different angles.

  5. 2024

    Stable Diffusion 3, Sora, π₀

    Flow matching became the backbone of state-of-the-art image generation (SD3), video generation (Sora), and robotic policy learning (π₀).

CitationLipman, Chen, Ben-Hamu, Nickel, Le. Flow Matching for Generative Modeling. ICLR, 2023.

Terms in this paper