Generative Models2023advanced11 min read

Consistency Models

نماذج الاتّساق

Song, Y. · Dhariwal, P. · Chen, M. · Sutskever, I. — ICML

The problem

Diffusion models produce stunning images, audio, and video, but they require tens to thousands of iterative steps at generation time. Each step runs the full neural network once, making real-time applications impractical. Prior acceleration methods like DDIM and DPM-Solver reduce steps but still need 10–50 passes. Progressive halves the step count each round but requires multiple training stages and never reaches true . The field needed a principled way to map noise directly to data in a single forward pass — without adversarial training and without sacrificing quality.

The contribution

A new family of generative models built on a simple principle: every point on the same trajectory should map to the same clean image. This "self-consistency" property is enforced during training so the model learns to jump from any noise level directly to the data in one step. Two training modes are introduced — Consistency Distillation (CD), which learns from a pre-trained 's ODE trajectories, and Consistency Training (CT), which trains from scratch using an unbiased score estimator. The approach achieves state-of-the-art of 3.55 on CIFAR-10 for one-step generation and supports zero-shot image editing tasks like inpainting, colorization, and .

The impact

Consistency models broke the assumption that diffusion-quality generation requires iterative . They directly inspired Latent Consistency Models (LCMs) that brought real-time image generation to latent diffusion pipelines like Stable Diffusion. The improved Consistency Training (iCT) follow-up and continuous-time formulations pushed standalone training quality even further. The consistency framework became a foundation for fast generative models across images, audio, video, and 3D, bridging the gap between diffusion model quality and -level speed.

Imagine a mountain with many hiking trails, all leading down to the same village. A diffusion model is like a hiker who carefully walks every switchback of the trail — safe but slow. A is like a paraglider who, from any point on any trail, can see the village below and glide straight to it.

The key insight: you don't need to memorize every trail. You just need to learn one rule — "every point on the same trail leads to the same village." That's the self-consistency property, and it lets you skip the walk entirely.

The bottleneck: why diffusion models are slow

Score-based diffusion models generate data by reversing a noise-injection process. Starting from pure Gaussian noise, they iteratively remove noise via a learned — the gradient of the log-probability density. This reverse process follows a Stochastic Differential Equation (SDE), or its deterministic counterpart, the Probability Flow ODE (PF-ODE).

The PF-ODE is especially important: it defines a deterministic trajectory from any noisy sample back to a clean data point. If you could solve this ODE exactly, you would get a perfect sample. But solving it requires many small steps, each one calling the neural network. A typical diffusion model needs 50–1000 forward passes to generate a single image.

Faster ODE solvers like DDIM and DPM-Solver reduce this to 10–50 steps, but going below 10 steps degrades quality sharply. Progressive distillation trains a student model to match two teacher steps in one, halving the step count each round. But it requires multiple rounds of training and inherits errors from each stage.

What if a model could learn to jump from any point on the ODE trajectory directly to the endpoint — in a single step?

Open in Lab
Left: diffusion models walk the ODE trajectory step by step. Right: a consistency model jumps from any point to the endpoint in one step.
The demo wakes as you arrive…

Core idea: the self-consistency property

The PF-ODE defines a unique trajectory for every data point. Given a clean image x₀, adding noise at level t produces x_t. The ODE trajectory connects every noisy version x_t back to the same x₀. This means that for any two noise levels t and t' on the same trajectory, the destination is identical.

A consistency function f captures this: f(x_t, t) = f(x_t', t') for all t and t' along the same trajectory. In other words, the function's output is consistent — it does not depend on where along the trajectory you evaluate it. This is the self-consistency property.

If we can train a neural network to satisfy this property, then at generation time we can sample noise x_T, evaluate the network once to get f(x_T, T), and obtain a clean sample directly. One step. No iteration.

Open in Lab
Drag the slider to pick any point on the trajectory — the output always maps to the same clean data point.
The demo wakes as you arrive…

Parameterization: satisfying the boundary condition

A valid consistency function must satisfy one hard constraint: at the minimum noise level ε (close to zero), the function must return the input itself — f(x, ε) = x. This is the boundary condition. It ensures the function is anchored at the clean-data end of the trajectory.

To enforce this by construction, the model is parameterized as a weighted combination of the input and a neural network output. Two differentiable scalar functions control the mix: c_skip(t) and c_out(t). At t = ε, the skip connection dominates (c_skip(ε) = 1, c_out(ε) = 0), so the output equals the input. As t grows, the network's output takes over.

fθ(xt,t)=cskip(t) xt+cout(t) Fθ(xt,t)f_\theta(\mathbf{x}_t, t) = c_{\text{skip}}(t)\,\mathbf{x}_t + c_{\text{out}}(t)\,F_\theta(\mathbf{x}_t, t)
Consistency Model Parameterization — The consistency model output is a weighted sum of the raw noisy input x_t (via the skip connection c_skip) and the neural network prediction F_θ (via c_out). At t = ε, c_skip = 1 and c_out = 0, guaranteeing f_θ(x, ε) = x.
Open in Lab
Adjust t to see how c_skip and c_out blend the input with the network output. At t = ε the output equals the input exactly.
The demo wakes as you arrive…

Training mode 1: Consistency Distillation (CD)

The first training approach leverages a pre-trained diffusion model as a teacher. The idea is straightforward: take a noisy sample at time step t₍ₙ₊₁₎, use the teacher's to estimate where it would land one step earlier at tₙ, then train the consistency model so its output at both points matches.

Concretely, the training loss enforces that the student's consistency output at the noisier point approximately equals the EMA target network's output at the cleaner point. The teacher provides the ODE step that connects them. The EMA target network θ⁻ — an of the student's weights — stabilizes training, similar to how target networks work in deep reinforcement learning.

As the time discretization gets finer (N → ∞), the distillation loss converges to zero for a perfect consistency function. In practice, N is increased gradually during training — a curriculum that starts with coarse time steps and progressively refines.

LCDN(θ,θ−)=E[λ(tn) d ⁣(fθ(xtn+1,tn+1),  fθ−(x^tnϕ,tn))]\mathcal{L}_{\text{CD}}^N(\theta, \theta^-) = \mathbb{E}\bigl[\lambda(t_n)\, d\!\bigl(f_\theta(\mathbf{x}_{t_{n+1}}, t_{n+1}),\; f_{\theta^-}(\hat{\mathbf{x}}_{t_n}^{\phi}, t_n)\bigr)\bigr]
Consistency Distillation Loss — The loss measures the distance d between the student output at the noisier time step and the EMA target network output at the adjacent cleaner time step, where the jump between them is computed by the teacher ODE solver. λ is a positive weighting function.
Open in Lab
Visualization of the distillation process: the teacher solves one ODE step, and the student learns to produce the same endpoint from a noisier input.
The demo wakes as you arrive…

Training mode 2: Consistency Training (CT) — no teacher needed

The most surprising result of this paper is that consistency models can be trained entirely from scratch, without any pre-trained diffusion model. The trick is replacing the teacher's ODE step with an unbiased estimate of the score function.

Recall that the score function ∇ log p_t(x_t) can be estimated from a single training sample: it equals −(x_t − x₀) / t². This means we can construct adjacent points on the ODE trajectory using only the ground-truth data and random noise — no teacher required.

The consistency training loss has the same structure as the distillation loss, but instead of relying on the teacher's ODE solver, it uses the analytical relationship between x₀, noise ε, and the forward diffusion process. As the discretization refines (N → ∞), the training objective provably converges to a form whose minimizer is the true consistency function.

This is significant: consistency models are not just a distillation trick — they are a standalone family of generative models with their own training objective.

LCTN(θ,θ−)=E[λ(tn) d ⁣(fθ(xtn+1,tn+1),  fθ−(xtn,tn))]\mathcal{L}_{\text{CT}}^N(\theta, \theta^-) = \mathbb{E}\bigl[\lambda(t_n)\, d\!\bigl(f_\theta(\mathbf{x}_{t_{n+1}}, t_{n+1}),\; f_{\theta^-}(\mathbf{x}_{t_n}, t_n)\bigr)\bigr]
Consistency Training Loss — Similar to the distillation loss but both noisy samples are obtained by adding noise at two adjacent levels to the same clean sample — no teacher ODE solver involved. The same noise is shared, ensuring both points lie on the same forward trajectory.

Beyond one step: multistep sampling and zero-shot editing

While one-step generation is the headline feature, consistency models can also do multistep sampling to improve quality. The idea is simple: generate a rough sample in one step, add noise back to an intermediate level, then denoise again. Each denoise-then-renoise cycle refines the output. With just 2–3 steps the quality improves significantly, offering a smooth compute-quality trade-off.

This same mechanism enables zero-shot image editing without any task-specific training. For inpainting, mask out a region, add noise, and denoise — the consistency model fills in the masked area coherently. For colorization, start with a grayscale image as partial information and let the model hallucinate colors. For super-resolution, begin with a low-resolution image, add noise, and denoise at higher resolution. All of these use the same trained model with no fine-tuning.

Open in Lab
Toggle between 1-step, 2-step, and 3-step sampling to see how quality improves with additional denoise-renoise cycles.
The demo wakes as you arrive…

Results: one-step quality and benchmarks

Consistency Distillation achieved an FID of 3.55 on CIFAR-10 and 6.20 on ImageNet 64×64 for one-step generation — a new state of the art among distillation methods. With two steps, CIFAR-10 FID dropped further to 2.93, approaching the quality of the teacher model running 100+ steps.

Consistency Training (without any teacher) achieved FID 7.5 on CIFAR-10 for one-step generation, outperforming all existing one-step non-adversarial generators. On LSUN bedrooms and cats at 256×256, consistency models produced compelling samples with 1–3 steps.

Importantly, the same model handles generation and editing. Zero-shot inpainting, colorization, and super-resolution all worked with no task-specific training, using only the multistep sampling procedure.

Open in Lab
FID comparison: Consistency Distillation (CD) and Consistency Training (CT) vs prior one-step and few-step methods on CIFAR-10.
The demo wakes as you arrive…

Architecture and training details

Consistency models use the same U-Net backbone as the teacher diffusion model — no architectural changes are needed. The time-conditioning input simply indicates the noise level. The key training decisions are:

Distance metric d(·,·): The paper explored L2, L1, and LPIPS (Learned Perceptual Image Patch Similarity). LPIPS consistently performed best, as it measures perceptual similarity rather than pixel-by-pixel error — important because the mapping from noise to data is one-to-many.

EMA decay rate μ: The target network θ⁻ is an exponential moving average of θ. A schedule that slowly increases μ toward 1 during training worked best.

Schedule function N(k): The number of discretization steps N is increased during training according to a schedule. Starting coarse and refining avoids early instability while ensuring fine-grained consistency at convergence.

ODE solver choice (for CD): Euler and Heun solvers both work. Heun (a second-order solver) gives slightly better results for the same discretization.

Pseudocode: Consistency Distillation and Training

Consistency Distillation — training loop pseudocodepython

Simplified to show the idea — not the real implementation.

# Consistency Distillation (CD) — simplified pseudocode
# Given: teacher score model s_phi, EMA rate mu, schedule N(k)
for each training step k:
    N = schedule(k)                    # increase N over training
    n = randint(1, N-1)                # sample a random time index
    x_0 = sample_data()                # clean image from dataset
    eps = randn_like(x_0)              # Gaussian noise
    x_tn1 = x_0 + t[n+1] * eps        # noisy sample at t_{n+1}

    # Teacher ODE step: estimate x at t_n
    x_tn_hat = ode_step(x_tn1, t[n+1], t[n], s_phi)

    # Student consistency outputs
    pred_student = f_theta(x_tn1, t[n+1])
    pred_target  = f_theta_ema(x_tn_hat, t[n])  # EMA network

    loss = d(pred_student, pred_target.detach())
    loss.backward()
    optimizer.step()
    theta_ema = mu * theta_ema + (1 - mu) * theta
Consistency Training — training loop pseudocode (no teacher)python

Simplified to show the idea — not the real implementation.

# Consistency Training (CT) — simplified pseudocode
# No pre-trained teacher needed — uses unbiased score estimator
for each training step k:
    N = schedule(k)                    # increase N over training
    n = randint(1, N-1)                # sample a random time index
    x_0 = sample_data()                # clean image from dataset
    eps = randn_like(x_0)              # same noise for both levels

    x_tn  = x_0 + t[n]   * eps        # noisy sample at t_n
    x_tn1 = x_0 + t[n+1] * eps        # noisy sample at t_{n+1}

    # Both use the SAME noise eps — ensures same trajectory
    pred_student = f_theta(x_tn1, t[n+1])
    pred_target  = f_theta_ema(x_tn, t[n])    # EMA network

    loss = d(pred_student, pred_target.detach())
    loss.backward()
    optimizer.step()
    theta_ema = mu * theta_ema + (1 - mu) * theta

Timeline: from diffusion to real-time generation

  1. 2020

    Score-based SDEs unify diffusion models

    Song et al. show that DDPM and score matching are special cases of a general SDE framework, introducing the Probability Flow ODE as a deterministic generation path.

  2. 2021

    DDIM enables deterministic sampling

    Song, Meng, & Ermon propose DDIM — a non-Markovian sampling process that reduces steps from 1000 to 50–100 by following the PF-ODE more efficiently.

  3. 2022

    Progressive distillation halves steps repeatedly

    Salimans & Ho train a student to match two teacher steps in one, iteratively halving from 1024 to 4 steps. Still requires multiple training stages.

  4. 2023

    Consistency Models achieve one-step generation

    Song et al. introduce the self-consistency property and train models that map any noise level directly to data. State-of-the-art FID for one-step generation on CIFAR-10 (3.55) and ImageNet 64×64 (6.20).

  5. 2023

    Latent Consistency Models (LCMs) bring real-time diffusion

    Luo et al. apply consistency distillation in latent space, enabling Stable Diffusion to generate images in 1–4 steps. Real-time generation becomes practical.

  6. 2024

    Improved and continuous-time consistency training

    Song & Dhariwal (iCT) and Lu & Song (sCT) push standalone training quality closer to distillation, with simplified objectives, better schedules, and continuous-time limits.

Consistency models catalyzed a paradigm shift: from "diffusion models are slow but high-quality" to "diffusion-quality generation can be fast." The self-consistency principle proved to be more than a clever trick — it is a fundamental design paradigm for generative models that trade iterative refinement for direct prediction.

CitationSong, Dhariwal, Chen, Sutskever. Consistency Models. ICML, 2023.

Terms in this paper