Generative Models2023advanced11 min read
Flow Matching for Generative Modeling
مُطابقة التدفق للنمذجة التوليدية
Lipman, Y. · Chen, R. T. Q. · Ben-Hamu, H. · Nickel, M. · Le, M. — ICLR
The problem
Continuous Normalizing Flows (CNFs) are a powerful class of generative models that transform a simple noise distribution into a data distribution via a learned . However, training CNFs by maximum requires simulating the ODE at every training step — an expensive operation that scales poorly to large datasets and high-resolution images. Meanwhile, diffusion models scale well but are locked into specific curved probability paths (VP, VE schedules) that require many steps, and their training objectives are tied to rather than direct vector field . There was no efficient, -free method to train CNFs with arbitrary probability paths.
The contribution
The paper introduces (FM), a simulation-free training paradigm for Continuous Normalizing Flows. The key theoretical insight is (CFM): instead of regressing the intractable marginal vector field, the model regresses conditional vector fields anchored to individual data points — and the authors prove this yields identical gradients. FM works with any Gaussian probability path, unifying diffusion paths as special cases. Crucially, it enables paths — straight-line trajectories that are faster to train, faster to sample, and produce better results than curved diffusion paths.
The impact
Flow Matching has become the default training paradigm for continuous generative models, replacing both score matching and simulation-based CNF training. It is the foundation of Stable Diffusion 3, OpenAI's Sora video generation model, and Meta's Voicebox speech synthesis. In robotics, Physical Intelligence's π₀ policy uses Flow Matching for action generation. The paper unified previously disjoint communities — CNFs, diffusion models, and optimal transport — under a single framework, and its OT paths set the standard for efficient sampling in modern generative AI.
Imagine a choreographer arranging a flash mob. A thousand dancers start scattered randomly across a plaza. The choreographer's job is to give each dancer a direction and speed — a tiny arrow — at every moment, so they all flow smoothly into a perfect formation. The choreographer never simulates the whole dance in advance; she just watches where each dancer needs to end up and draws the arrow pointing straight there.
Previous methods (diffusion models) told dancers to wander in spirals before reaching their spot. This paper says: why not walk in straight lines? The dancers arrive faster, the choreography is simpler, and the formation is sharper.
That's Flow Matching — learning the arrows (velocity fields) that transport noise into data along the most direct path.
The problem: simulation is expensive, diffusion is curved
A learns a that pushes a simple distribution (like a Gaussian) into a complex data distribution by solving an ordinary differential equation. Think of it as a river current: every point in space has an arrow telling it where to flow, and following those arrows from to transforms noise into data.
The problem with training CNFs before this paper had two faces. First, maximum likelihood training required simulating the entire ODE at each training step to compute the change in log-density — this involves the trace of the velocity field, which is computationally brutal for high-dimensional data. Second, while diffusion models found simulation-free shortcuts via score matching, those shortcuts locked the probability path into specific curved trajectories (VP, VE schedules) that needed hundreds of ODE steps to sample accurately.
The field needed a training method that was both simulation-free (like score matching) and path-flexible (unlike score matching). Flow Matching is that method.
The core idea: regress the velocity, skip the simulation
The central idea of Flow Matching is disarmingly simple. Instead of training a CNF by simulating it (which is expensive), or by matching scores (which constrains the path), we directly regress the velocity field. We define a target probability path from noise () to data () and a target vector field that generates this path. Then we train a neural network to predict everywhere.
The ideal training loss is the Flow Matching objective — a simple between the predicted and target velocity fields:
There is one catch: the marginal velocity field is intractable. Computing it requires integrating over the entire data distribution, which is the very thing we are trying to learn. This is where the paper's theoretical breakthrough comes in.
The trick: conditional flow matching
The insight that makes everything work: instead of regressing the intractable marginal vector field , regress the conditional vector field — the velocity that transports noise to a single data point . The Conditional Flow Matching (CFM) loss is:
Why does this work? The paper proves (Theorem 2) that the FM loss and the CFM loss have identical gradients with respect to the network parameters . This is the same type of insight behind denoising score matching: individual conditional targets may point in different directions (each one pulls toward a single data point), but their average over all data points equals the true marginal direction.
Think of it like surveying a crowd: asking every person "which way is the exit?" gives noisy individual answers, but averaging all fingers yields the correct direction. Each conditional velocity field is one finger; the marginal field is the average.
Gaussian probability paths: a unified recipe
Flow Matching works with any probability path, but the paper focuses on Gaussian conditional paths — because they yield closed-form velocity targets that make training trivial. A Gaussian conditional path takes the form:
where and are time-dependent schedules controlling the mean and standard deviation. At , the path should be pure noise (, ). At , it should concentrate on the data point (, ). Everything in between is a smooth .
The paper proves (Theorem 3) that for any such Gaussian path, the conditional velocity field has a simple closed form:
The power move: Optimal Transport paths
With the general framework in place, the authors propose a specific path choice that is dramatically better than diffusion paths: Optimal Transport (OT) displacement interpolation. The OT path uses the simplest possible schedules:
This means each sample follows a straight line from its noise origin to its data destination. No curves, no spirals, no detours. Plugging these schedules into the Gaussian velocity formula gives the OT conditional velocity:
Why are straight paths better? Three reasons:
Faster sampling. A straight trajectory is easier for numerical ODE solvers to follow — even a basic Euler integrator can produce high-quality samples in 100 function evaluations, compared to 250+ for diffusion paths.
Easier learning. When the velocity field is smooth and nearly constant along each trajectory, the neural network has an easier regression target. Diffusion paths create rapidly varying velocity fields near and that are harder to fit.
Better generalization. Empirically, OT paths consistently produce lower scores and better log-likelihoods than diffusion paths under the same training budget.
The training recipe: five lines of pseudocode
One of the most striking features of Flow Matching is how simple the becomes. The entire algorithm reduces to:
-
Sample a data point and noise
-
Sample a random time
-
Interpolate:
-
Compute the target velocity:
-
Minimize
That is the entire training loop. No score functions, no lookups, no ODE simulation. Just sample, interpolate, regress.
Simplified to show the idea — not the real implementation.
# Flow Matching with OT paths — one training step
sigma_min = 1e-5
t = torch.rand(batch_size, 1) # random time
x0 = torch.randn_like(x1) # noise sample
# Interpolate along straight line from noise to data
x_t = (1 - (1 - sigma_min) * t) * x0 + t * x1
# Target velocity: direction of the straight line
u_t = x1 - (1 - sigma_min) * x0
# Train the network to predict this velocity
loss = ((model(t, x_t) - u_t) ** 2).mean()
loss.backward()Sampling: just follow the arrows
Once the velocity field is trained, generating samples is pure ODE integration. Start from random noise and solve:
Any standard works — dopri5 (adaptive Runge-Kutta) for maximum quality, or a simple Euler method for fast generation with fewer steps. Because OT paths produce nearly straight trajectories, even crude solvers give good results.
For computing exact likelihoods, the model augments the ODE with the formula, tracking the log-determinant of the Jacobian as the flow progresses:
Results: better images, faster training, fewer steps
The authors trained Flow Matching on CIFAR-10 and ImageNet (, , ) using the same architecture as prior diffusion models. The results are consistent across every setting:
FM with OT paths beats diffusion baselines on both FID and likelihood. On CIFAR-10, FM-OT achieves FID 6.35 vs DDPM's 7.48, and NLL 2.99 vs 3.12 bits per dimension. On ImageNet , FM-OT reaches FID 14.45 vs ScoreFlow's 24.95.
FM with diffusion paths is already more stable than standard diffusion training. Even without switching to OT, the FM training objective provides smoother loss curves and faster convergence than DDPM or score matching on the same diffusion paths.
OT paths require far fewer sampling steps. FM-OT achieves comparable quality with 142 NFE (number of function evaluations) on CIFAR-10, versus 274 for DDPM. On ImageNet , the gap widens to 138 vs 601.
Legacy: the engine behind modern generative AI
Flow Matching did not just offer a better training method — it changed how the field thinks about generative modeling. By proving that diffusion is one instance of a broader family of probability paths, it freed researchers to explore new geometries, new interpolation strategies, and new domains far beyond images.
The impact has been sweeping and rapid. Sora (OpenAI, 2024) uses flow matching for video generation. Stable Diffusion 3 (Stability AI, 2024) replaced its diffusion backbone with flow matching. Meta's Voicebox and AudioBox generate speech and sound effects via flow matching. In robotics, π₀ (Physical Intelligence, 2024) uses flow matching to generate continuous robotic actions.
Concurrent work at ICLR 2023 — Rectified Flow (Liu et al.) and Stochastic Interpolants (Albergo & Vanden-Eijnden) — independently arrived at structurally similar frameworks, confirming that this was an idea whose time had come. Together, these papers marked the transition from "diffusion models" to "flow-based generative models" as the dominant paradigm.
2018
Neural ODE (Chen et al.)
Introduced Continuous Normalizing Flows trained via the adjoint sensitivity method. Powerful but required expensive ODE simulation during training.
2020
Score-based diffusion (Song et al.)
Unified score matching and diffusion via SDEs, enabling scalable training but constraining the probability path to VP/VE schedules.
2023
Flow Matching (this paper, Lipman et al.)
Simulation-free CNF training via conditional vector field regression. Introduced OT paths for straight-line transport, outperforming diffusion baselines.
2023
Concurrent — Rectified Flow & Stochastic Interpolants
Liu et al. and Albergo & Vanden-Eijnden independently proposed similar simulation-free flow training, confirming the paradigm from different angles.
2024
Stable Diffusion 3, Sora, π₀
Flow matching became the backbone of state-of-the-art image generation (SD3), video generation (Sora), and robotic policy learning (π₀).
CitationLipman, Chen, Ben-Hamu, Nickel, Le. Flow Matching for Generative Modeling. ICLR, 2023.
Terms in this paper
- Flow Matchingمطابقة التدفقات
- Continuous Normalizing Flowتدفق التسوية المستمر
- Velocity Fieldحقل السرعة
- Optimal Transportالنقل الأمثل
- Probability Density Functionدالة الكثافة الاحتمالية
- Neural ODEالمعادلة التفاضلية العصبية
- Ordinary Differential Equationمعادلة تفاضلية عادية
- ODE Solverحالّ المعادلات التفاضلية
- Gaussian Distributionالتوزيع الغاوسي
- Diffusion Modelنموذج الانتشار
- Score-Based Generative Modelالنماذج التوليدية القائمة على التدرج
- Normalizing Flowالتدفق التسوِيّ
- Density Estimationتقدير الكثافة
- FIDمسافة فريشيه للبداية
- Regressionالانحدار الإحصائي