Generative Models2023advanced11 min read
Consistency Models
نماذج الاتّساق
Song, Y. · Dhariwal, P. · Chen, M. · Sutskever, I. — ICML
The problem
Diffusion models produce stunning images, audio, and video, but they require tens to thousands of iterative steps at generation time. Each step runs the full neural network once, making real-time applications impractical. Prior acceleration methods like DDIM and DPM-Solver reduce steps but still need 10–50 passes. Progressive halves the step count each round but requires multiple training stages and never reaches true . The field needed a principled way to map noise directly to data in a single forward pass — without adversarial training and without sacrificing quality.
The contribution
A new family of generative models built on a simple principle: every point on the same trajectory should map to the same clean image. This "self-consistency" property is enforced during training so the model learns to jump from any noise level directly to the data in one step. Two training modes are introduced — Consistency Distillation (CD), which learns from a pre-trained 's ODE trajectories, and Consistency Training (CT), which trains from scratch using an unbiased score estimator. The approach achieves state-of-the-art of 3.55 on CIFAR-10 for one-step generation and supports zero-shot image editing tasks like inpainting, colorization, and .
The impact
Consistency models broke the assumption that diffusion-quality generation requires iterative . They directly inspired Latent Consistency Models (LCMs) that brought real-time image generation to latent diffusion pipelines like Stable Diffusion. The improved Consistency Training (iCT) follow-up and continuous-time formulations pushed standalone training quality even further. The consistency framework became a foundation for fast generative models across images, audio, video, and 3D, bridging the gap between diffusion model quality and -level speed.
Imagine a mountain with many hiking trails, all leading down to the same village. A diffusion model is like a hiker who carefully walks every switchback of the trail — safe but slow. A is like a paraglider who, from any point on any trail, can see the village below and glide straight to it.
The key insight: you don't need to memorize every trail. You just need to learn one rule — "every point on the same trail leads to the same village." That's the self-consistency property, and it lets you skip the walk entirely.
The bottleneck: why diffusion models are slow
Score-based diffusion models generate data by reversing a noise-injection process. Starting from pure Gaussian noise, they iteratively remove noise via a learned — the gradient of the log-probability density. This reverse process follows a Stochastic Differential Equation (SDE), or its deterministic counterpart, the Probability Flow ODE (PF-ODE).
The PF-ODE is especially important: it defines a deterministic trajectory from any noisy sample back to a clean data point. If you could solve this ODE exactly, you would get a perfect sample. But solving it requires many small steps, each one calling the neural network. A typical diffusion model needs 50–1000 forward passes to generate a single image.
Faster ODE solvers like DDIM and DPM-Solver reduce this to 10–50 steps, but going below 10 steps degrades quality sharply. Progressive distillation trains a student model to match two teacher steps in one, halving the step count each round. But it requires multiple rounds of training and inherits errors from each stage.
What if a model could learn to jump from any point on the ODE trajectory directly to the endpoint — in a single step?
Core idea: the self-consistency property
The PF-ODE defines a unique trajectory for every data point. Given a clean image x₀, adding noise at level t produces x_t. The ODE trajectory connects every noisy version x_t back to the same x₀. This means that for any two noise levels t and t' on the same trajectory, the destination is identical.
A consistency function f captures this: f(x_t, t) = f(x_t', t') for all t and t' along the same trajectory. In other words, the function's output is consistent — it does not depend on where along the trajectory you evaluate it. This is the self-consistency property.
If we can train a neural network to satisfy this property, then at generation time we can sample noise x_T, evaluate the network once to get f(x_T, T), and obtain a clean sample directly. One step. No iteration.
Parameterization: satisfying the boundary condition
A valid consistency function must satisfy one hard constraint: at the minimum noise level ε (close to zero), the function must return the input itself — f(x, ε) = x. This is the boundary condition. It ensures the function is anchored at the clean-data end of the trajectory.
To enforce this by construction, the model is parameterized as a weighted combination of the input and a neural network output. Two differentiable scalar functions control the mix: c_skip(t) and c_out(t). At t = ε, the skip connection dominates (c_skip(ε) = 1, c_out(ε) = 0), so the output equals the input. As t grows, the network's output takes over.
Training mode 1: Consistency Distillation (CD)
The first training approach leverages a pre-trained diffusion model as a teacher. The idea is straightforward: take a noisy sample at time step t₍ₙ₊₁₎, use the teacher's to estimate where it would land one step earlier at tₙ, then train the consistency model so its output at both points matches.
Concretely, the training loss enforces that the student's consistency output at the noisier point approximately equals the EMA target network's output at the cleaner point. The teacher provides the ODE step that connects them. The EMA target network θ⁻ — an of the student's weights — stabilizes training, similar to how target networks work in deep reinforcement learning.
As the time discretization gets finer (N → ∞), the distillation loss converges to zero for a perfect consistency function. In practice, N is increased gradually during training — a curriculum that starts with coarse time steps and progressively refines.
Training mode 2: Consistency Training (CT) — no teacher needed
The most surprising result of this paper is that consistency models can be trained entirely from scratch, without any pre-trained diffusion model. The trick is replacing the teacher's ODE step with an unbiased estimate of the score function.
Recall that the score function ∇ log p_t(x_t) can be estimated from a single training sample: it equals −(x_t − x₀) / t². This means we can construct adjacent points on the ODE trajectory using only the ground-truth data and random noise — no teacher required.
The consistency training loss has the same structure as the distillation loss, but instead of relying on the teacher's ODE solver, it uses the analytical relationship between x₀, noise ε, and the forward diffusion process. As the discretization refines (N → ∞), the training objective provably converges to a form whose minimizer is the true consistency function.
This is significant: consistency models are not just a distillation trick — they are a standalone family of generative models with their own training objective.
Beyond one step: multistep sampling and zero-shot editing
While one-step generation is the headline feature, consistency models can also do multistep sampling to improve quality. The idea is simple: generate a rough sample in one step, add noise back to an intermediate level, then denoise again. Each denoise-then-renoise cycle refines the output. With just 2–3 steps the quality improves significantly, offering a smooth compute-quality trade-off.
This same mechanism enables zero-shot image editing without any task-specific training. For inpainting, mask out a region, add noise, and denoise — the consistency model fills in the masked area coherently. For colorization, start with a grayscale image as partial information and let the model hallucinate colors. For super-resolution, begin with a low-resolution image, add noise, and denoise at higher resolution. All of these use the same trained model with no fine-tuning.
Results: one-step quality and benchmarks
Consistency Distillation achieved an FID of 3.55 on CIFAR-10 and 6.20 on ImageNet 64×64 for one-step generation — a new state of the art among distillation methods. With two steps, CIFAR-10 FID dropped further to 2.93, approaching the quality of the teacher model running 100+ steps.
Consistency Training (without any teacher) achieved FID 7.5 on CIFAR-10 for one-step generation, outperforming all existing one-step non-adversarial generators. On LSUN bedrooms and cats at 256×256, consistency models produced compelling samples with 1–3 steps.
Importantly, the same model handles generation and editing. Zero-shot inpainting, colorization, and super-resolution all worked with no task-specific training, using only the multistep sampling procedure.
Architecture and training details
Consistency models use the same U-Net backbone as the teacher diffusion model — no architectural changes are needed. The time-conditioning input simply indicates the noise level. The key training decisions are:
Distance metric d(·,·): The paper explored L2, L1, and LPIPS (Learned Perceptual Image Patch Similarity). LPIPS consistently performed best, as it measures perceptual similarity rather than pixel-by-pixel error — important because the mapping from noise to data is one-to-many.
EMA decay rate μ: The target network θ⁻ is an exponential moving average of θ. A schedule that slowly increases μ toward 1 during training worked best.
Schedule function N(k): The number of discretization steps N is increased during training according to a schedule. Starting coarse and refining avoids early instability while ensuring fine-grained consistency at convergence.
ODE solver choice (for CD): Euler and Heun solvers both work. Heun (a second-order solver) gives slightly better results for the same discretization.
Pseudocode: Consistency Distillation and Training
Simplified to show the idea — not the real implementation.
# Consistency Distillation (CD) — simplified pseudocode
# Given: teacher score model s_phi, EMA rate mu, schedule N(k)
for each training step k:
N = schedule(k) # increase N over training
n = randint(1, N-1) # sample a random time index
x_0 = sample_data() # clean image from dataset
eps = randn_like(x_0) # Gaussian noise
x_tn1 = x_0 + t[n+1] * eps # noisy sample at t_{n+1}
# Teacher ODE step: estimate x at t_n
x_tn_hat = ode_step(x_tn1, t[n+1], t[n], s_phi)
# Student consistency outputs
pred_student = f_theta(x_tn1, t[n+1])
pred_target = f_theta_ema(x_tn_hat, t[n]) # EMA network
loss = d(pred_student, pred_target.detach())
loss.backward()
optimizer.step()
theta_ema = mu * theta_ema + (1 - mu) * thetaSimplified to show the idea — not the real implementation.
# Consistency Training (CT) — simplified pseudocode
# No pre-trained teacher needed — uses unbiased score estimator
for each training step k:
N = schedule(k) # increase N over training
n = randint(1, N-1) # sample a random time index
x_0 = sample_data() # clean image from dataset
eps = randn_like(x_0) # same noise for both levels
x_tn = x_0 + t[n] * eps # noisy sample at t_n
x_tn1 = x_0 + t[n+1] * eps # noisy sample at t_{n+1}
# Both use the SAME noise eps — ensures same trajectory
pred_student = f_theta(x_tn1, t[n+1])
pred_target = f_theta_ema(x_tn, t[n]) # EMA network
loss = d(pred_student, pred_target.detach())
loss.backward()
optimizer.step()
theta_ema = mu * theta_ema + (1 - mu) * thetaTimeline: from diffusion to real-time generation
2020
Score-based SDEs unify diffusion models
Song et al. show that DDPM and score matching are special cases of a general SDE framework, introducing the Probability Flow ODE as a deterministic generation path.
2021
DDIM enables deterministic sampling
Song, Meng, & Ermon propose DDIM — a non-Markovian sampling process that reduces steps from 1000 to 50–100 by following the PF-ODE more efficiently.
2022
Progressive distillation halves steps repeatedly
Salimans & Ho train a student to match two teacher steps in one, iteratively halving from 1024 to 4 steps. Still requires multiple training stages.
2023
Consistency Models achieve one-step generation
Song et al. introduce the self-consistency property and train models that map any noise level directly to data. State-of-the-art FID for one-step generation on CIFAR-10 (3.55) and ImageNet 64×64 (6.20).
2023
Latent Consistency Models (LCMs) bring real-time diffusion
Luo et al. apply consistency distillation in latent space, enabling Stable Diffusion to generate images in 1–4 steps. Real-time generation becomes practical.
2024
Improved and continuous-time consistency training
Song & Dhariwal (iCT) and Lu & Song (sCT) push standalone training quality closer to distillation, with simplified objectives, better schedules, and continuous-time limits.
Consistency models catalyzed a paradigm shift: from "diffusion models are slow but high-quality" to "diffusion-quality generation can be fast." The self-consistency principle proved to be more than a clever trick — it is a fundamental design paradigm for generative models that trade iterative refinement for direct prediction.
CitationSong, Dhariwal, Chen, Sutskever. Consistency Models. ICML, 2023.
Terms in this paper
- Consistency Modelنموذج الاتّساق
- Diffusion Modelنموذج الانتشار
- Probability Flow ODEالمعادلة التفاضلية العادية لتدفّق الاحتمال
- Score Functionدالة الرصيد
- Distillationالتقطير
- One-Step Generationالتوليد بخطوة واحدة
- Noise Scheduleجدول الضوضاء
- Exponential Moving Averageالمتوسط المتحرك الأُسّي
- ODE Solverحالّ المعادلات التفاضلية
- FIDمسافة فريشيه للبداية
- Classifier-Free Guidanceالتوجيه بلا مُصنِّف
- Denoisingإزالة الضوضاء
- Image Inpaintingملء فراغات الصور
- Super-Resolutionتحسين الدقة
- Autoregressive Modelالنموذج التوليدي التراجعي