Generative Models2021intermediate11 min read

Denoising Diffusion Implicit Models

نماذج الانتشار الضمنية لإزالة الضجيج

Song, J. · Meng, C. · Ermon, S. — ICLR

The problem

By 2020, diffusion probabilistic models (DDPMs) were generating images rivaling GANs in quality — without the instability of adversarial . But they had a crippling weakness: speed. Generating a single image required simulating a 1,000-step , making them orders of magnitude slower than GANs. 50,000 CIFAR-10 images took about 20 hours on a single GPU, compared to under a minute for a GAN. For high-resolution images, the wait stretched to nearly 1,000 hours. This made DDPMs impractical for any latency-sensitive application.

The contribution

generalizes the framework by replacing the Markovian with a family of non-Markovian processes that share the same marginal distributions and therefore the same training objective. When the stochasticity parameter σ is set to zero, the resulting generative process becomes fully deterministic — an that maps a latent noise vector to a unique image. This deterministic shortcut allows sampling in as few as 10–50 steps while maintaining quality comparable to the original 1,000-step DDPM. The approach requires no retraining: any pretrained DDPM works directly. DDIM also enables meaningful and near-perfect reconstruction via an ODE-like encoding–decoding cycle.

The impact

DDIM made diffusion models practical. Its accelerated deterministic sampling became the default sampler in Stable Diffusion, DALL·E 2, and virtually every modern image generator. The deterministic mapping between noise and image opened the door to latent manipulation techniques used in image editing. The connection to neural ODEs inspired a wave of follow-up work on faster samplers, including DPM-Solver and consistency models. By proving that a single pretrained model supports an entire family of generative processes, DDIM reshaped how the community thinks about the relationship between training and in diffusion models.

Imagine restoring a Renaissance painting that has been completely covered in dust. The traditional method (DDPM) insists on wiping the painting with exactly 1,000 gentle cloth strokes, each removing a microscopic layer of dust — and the direction of each stroke is chosen by a coin flip.

DDIM discovers something remarkable: because you already know the painting's structure from training, you can take a direct path to the clean image. Instead of 1,000 random wipes, you take 50 precise, calculated strokes — deterministic, no coin flips — and arrive at the same restored masterpiece. Better yet, you don't need to learn any new technique; the same training that prepared you for 1,000 strokes works perfectly for 50.

The bottleneck: 1,000 mandatory steps

To understand DDIM, we first need to see why DDPM is slow. In DDPM, the forward process gradually adds Gaussian noise to a clean image x0x_0 over T=1,000T = 1{,}000 steps, producing a sequence x1,x2,…,xTx_1, x_2, \ldots, x_T that ends in pure noise. This is a Markov chain: each step xtx_t depends only on the previous step xt−1x_{t-1}.

The generative process reverses this chain: starting from pure noise xT∼N(0,I)x_T \sim \mathcal{N}(0, I), it denoises step by step until reaching a clean image x0x_0. Because the forward process has TT steps, the must also take TT steps — one neural network evaluation per step. There is no shortcut in a Markov chain.

A key property makes this possible: we can jump from x0x_0 to any step xtx_t directly:

xt=αt x0+1−αt ϵ,ϵ∼N(0,I)x_t = \sqrt{\alpha_t}\, x_0 + \sqrt{1 - \alpha_t}\, \epsilon, \quad \epsilon \sim \mathcal{N}(0, I)
Forward marginal — jumping directly from clean image to any noise level — At any time step tt, the noisy image xtx_t is just a weighted mix of the clean image x0x_0 and random noise ϵ\epsilon. The coefficient αt\alpha_t decreases from 1 to near 0, so early steps are mostly signal and late steps are mostly noise. This formula works for *any* tt — no need to go through all intermediate steps.

This is the crucial observation: the marginal distribution q(xt∣x0)q(x_t | x_0) at each step does not depend on the path taken to get there. Whether we arrived through 1,000 Markov steps or jumped directly, the statistics are the same. The DDPM training objective (predicting the noise ϵ\epsilon from xtx_t) only depends on these marginals — not on the joint distribution q(x1,…,xT∣x0)q(x_1, \ldots, x_T | x_0). This is the door that DDIM opens.

Open in Lab
Compare DDPM's 1,000-step chain with DDIM's accelerated trajectory. Both start from the same noise and arrive at comparable quality — but DDIM gets there much faster.
The demo wakes as you arrive…

The key insight: many paths, same destination

The Markov chain in DDPM is just one particular joint distribution q(x1,…,xT∣x0)q(x_1, \ldots, x_T | x_0) that produces the marginals q(xt∣x0)=N(αt x0,(1−αt)I)q(x_t | x_0) = \mathcal{N}(\sqrt{\alpha_t}\, x_0, (1 - \alpha_t) I). But there are infinitely many other joint distributions that produce the exact same marginals.

Think of it like this: imagine you know the weather statistics for every month of the year. The Markov model says "today's weather depends only on yesterday." But you could also have a model where today's weather depends on yesterday and on the overall season — a non-Markovian model. Both can produce the same monthly averages, but the day-to-day transitions look very different.

DDIM exploits this freedom. It defines a family of non-Markovian forward processes, indexed by a vector σ=(σ1,…,σT)\sigma = (\sigma_1, \ldots, \sigma_T), where each σt\sigma_t controls how much randomness enters at step tt. The family is designed so that the marginals qσ(xt∣x0)q_\sigma(x_t | x_0) are identical to DDPM's for every σ\sigma. Since the training objective depends only on these marginals, the same pretrained model works for the entire family — no retraining needed.

Open in Lab
Toggle between Markovian and non-Markovian forward processes. Both produce the same marginal distributions at every step — but the dependencies between steps are different.
The demo wakes as you arrive…

The update rule: three forces in one equation

The heart of DDIM is a single update equation that generates xt−1x_{t-1} from xtx_t. Before we see the math, let's understand what it's doing.

The model ϵθ(t)(xt)\epsilon_\theta^{(t)}(x_t) has been trained to predict the noise that was added to the clean image. Given this prediction, we can estimate what the clean image x0x_0 looks like. Then we "re-noise" this estimate to the correct level for step t−1t{-}1. The equation blends three components: the predicted clean image, a correction term that points in the "noise direction," and an optional random noise term controlled by σt\sigma_t.

xt−1=αt−1(xt−1−αt  ϵθ(t)(xt)αt)⏟predicted x0+1−αt−1−σt2⋅ϵθ(t)(xt)⏟direction pointing to xt+σtϵt⏟random noisex_{t-1} = \sqrt{\alpha_{t-1}} \underbrace{\left(\frac{x_t - \sqrt{1-\alpha_t}\;\epsilon_\theta^{(t)}(x_t)}{\sqrt{\alpha_t}}\right)}_{\text{predicted } x_0} + \underbrace{\sqrt{1-\alpha_{t-1}-\sigma_t^2} \cdot \epsilon_\theta^{(t)}(x_t)}_{\text{direction pointing to } x_t} + \underbrace{\sigma_t \epsilon_t}_{\text{random noise}}
The DDIM generative update — from noise to image — Three terms, three roles. The first estimates the clean image x0x_0 from the current noisy state. The second corrects the direction toward the noise level of step t−1t{-}1. The third injects fresh randomness scaled by σt\sigma_t. When σt=0\sigma_t = 0 (DDIM), the process is fully deterministic — no randomness at all. When σt\sigma_t matches DDPM's formula, we recover the original stochastic process.
Open in Lab
Adjust σ to see how the three terms of the update equation balance. At σ = 0 (DDIM), the random noise term vanishes entirely.
The demo wakes as you arrive…

Think of it as a tug-of-war between "signal" and "noise." The first term pulls toward the clean image — like a compass pointing home. The second term adjusts the noise level so we land at the right position on the denoising schedule. The third term adds controlled randomness that is useful for diversity but slows convergence.

The parameter η\eta provides a convenient way to interpolate. Setting η=0\eta = 0 gives DDIM (deterministic), η=1\eta = 1 gives DDPM (stochastic), and values in between give a smooth continuum. This is achieved by defining σt(η)=η(1−αt−1)/(1−αt)1−αt/αt−1\sigma_t(\eta) = \eta \sqrt{(1-\alpha_{t-1})/(1-\alpha_t)} \sqrt{1 - \alpha_t / \alpha_{t-1}}.

The acceleration trick: skip steps, keep quality

The second breakthrough is that we don't need to reverse all T=1,000T = 1{,}000 steps. Since the training objective only depends on marginals q(xt∣x0)q(x_t | x_0), we can choose a subsequence τ=[τ1,τ2,…,τS]\tau = [\tau_1, \tau_2, \ldots, \tau_S] of the original [1,…,T][1, \ldots, T] and run the generative process only on those steps. Instead of going through x1000→x999→⋯→x0x_{1000} \to x_{999} \to \cdots \to x_0, we might jump x1000→x900→x800→⋯→x0x_{1000} \to x_{900} \to x_{800} \to \cdots \to x_0 — taking only 10 steps.

The update equation works exactly the same, just with ατi\alpha_{\tau_i} and ατi−1\alpha_{\tau_{i-1}} replacing αt\alpha_t and αt−1\alpha_{t-1}. The model was trained on all 1,000 noise levels, so it can denoise at any of them. We're simply choosing to visit fewer waypoints on the journey from noise to image.

The experimental results are striking. On CIFAR-10, DDIM with 50 steps achieves an of 4.67, comparable to DDPM's 4.04 with 1,000 steps — a 20× speedup with minimal quality loss. With 20 steps, DDIM still scores 6.84, while DDPM collapses to 18.36 under the same budget.

Open in Lab
Drag the slider to change the number of sampling steps. Watch how DDIM (η=0) maintains quality far better than DDPM (η=1) as steps decrease.
The demo wakes as you arrive…

The bonus: a meaningful latent space

Because DDIM is deterministic, the same initial noise xTx_T always produces the same image x0x_0 — regardless of how many steps we use. This is called consistency: a 10-step and a 1,000-step generation from identical xTx_T share the same high-level features (pose, color, composition), differing only in fine details.

This transforms xTx_T into a proper latent code, similar to the latent space in GANs or variational autoencoders. We gain two powerful capabilities. First, interpolation: given two noise vectors xT(0)x_T^{(0)} and xT(1)x_T^{(1)}, we can smoothly walk between them using spherical linear interpolation and generate all intermediate images. The transitions are semantically meaningful — faces morph smoothly, scenes blend naturally.

Second, encoding: by running the DDIM update in reverse (from x0x_0 to xTx_T), we can encode any image into the latent space, then decode it back with near-perfect reconstruction. The reconstruction error drops below 10−410^{-4} with 1,000 encoding steps. This is impossible with DDPM because its stochastic process destroys the one-to-one mapping.

Open in Lab
Drag the slider to interpolate between two latent noise vectors. The deterministic DDIM process produces smooth, semantically meaningful transitions.
The demo wakes as you arrive…

The deeper connection: DDIM as an ODE solver

There is a beautiful mathematical connection hidden beneath the surface. If we reparametrize by dividing each xtx_t by αt\sqrt{\alpha_t} and defining σˉ(t)=(1−αt)/αt\bar{\sigma}(t) = \sqrt{(1-\alpha_t)/\alpha_t}, the DDIM update becomes a single Euler step of an :

dxˉ(t)=ϵθ(t) ⁣(xˉ(t)σˉ2(t)+1)dσˉ(t)d\bar{x}(t) = \epsilon_\theta^{(t)}\!\left(\frac{\bar{x}(t)}{\sqrt{\bar{\sigma}^2(t)+1}}\right) d\bar{\sigma}(t)
DDIM as a neural ODE — continuous-time formulation — The DDIM update is an Euler discretization of this ODE. Running it forward (from t=0t=0 to TT) encodes an image into noise; running it backward (from TT to 00) decodes noise into an image. This is equivalent to the probability flow ODE of the "Variance-Exploding" SDE framework proposed concurrently by Song et al. (2020), unifying the score-based and denoising-based perspectives.

This ODE perspective is powerful for two reasons. First, it means we can apply decades of research on numerical ODE solvers (Adams-Bashforth, Runge-Kutta) to improve sampling quality in fewer steps. Second, it gives us a reversible encoding: unlike DDPM, where stochasticity makes the mapping from x0x_0 to xTx_T one-to-many, the DDIM ODE defines a bijection — every image has exactly one latent code, and vice versa. This makes DDIM a normalizing flow by another name, with the noise prediction network playing the role of the .

DDIM sampling loop (simplified)python

Simplified to show the idea — not the real implementation.

import torch

def ddim_sample(model, alphas, tau, eta=0.0):
    """
    model  : trained noise predictor epsilon_theta
    alphas : full alpha schedule [alpha_1, ..., alpha_T]
    tau    : subsequence of step indices, e.g. [100, 200, ..., 1000]
    eta    : 0 = DDIM (deterministic), 1 = DDPM (stochastic)
    """
    # Start from pure noise
    x = torch.randn_like(alphas[0])  # x_T ~ N(0, I)

    for i in reversed(range(len(tau))):
        t = tau[i]
        t_prev = tau[i - 1] if i > 0 else 0
        a_t = alphas[t]
        a_prev = alphas[t_prev] if t_prev > 0 else 1.0

        # Predict noise
        eps = model(x, t)

        # Estimate clean image (predicted x_0)
        x0_pred = (x - (1 - a_t).sqrt() * eps) / a_t.sqrt()

        # Compute sigma for this step
        sigma = eta * ((1 - a_prev) / (1 - a_t)).sqrt() * (1 - a_t / a_prev).sqrt()

        # Direction pointing to x_t
        dir_xt = (1 - a_prev - sigma**2).sqrt() * eps

        # Combine: predicted x_0 + direction + noise
        x = a_prev.sqrt() * x0_pred + dir_xt + sigma * torch.randn_like(x)

    return x

Why this matters: from theory to Stable Diffusion

DDIM solved the practical barrier that stood between diffusion models and real-world deployment. Before DDIM, diffusion models were an academic curiosity — beautiful math, excellent sample quality, but too slow to use. After DDIM, diffusion models became commercially viable.

The impact was immediate and far-reaching. Every modern diffusion-based system — Stable Diffusion, DALL·E 2, Imagen, Midjourney — uses DDIM-style accelerated sampling as a fundamental building block. The deterministic latent space enabled image editing techniques like DDIM inversion, where real photos are encoded into latent space, edited, and decoded back. The ODE connection inspired DPM-Solver, which pushed the boundary further to 10–20 steps with higher-order ODE solvers. And the core idea — that a single trained model supports many inference strategies — directly inspired consistency models, which distill the entire trajectory into a one-step generator.

Legacy: accelerating diffusion

  1. 2020

    DDPM (Ho et al.)

    Demonstrated that diffusion models can match GAN quality on images — but required 1,000 sequential denoising steps, making sampling extremely slow.

  2. 2021

    DDIM (this paper)

    Introduced non-Markovian forward processes and deterministic sampling, cutting steps from 1,000 to 50 with minimal quality loss — and no retraining.

  3. 2021

    Score SDE (Song et al.)

    Unified diffusion models under a continuous-time SDE framework, with a probability flow ODE equivalent to DDIM's formulation.

  4. 2022

    DALL·E 2 & Stable Diffusion

    Commercial image generators adopted DDIM-style sampling as their default, bringing diffusion models to millions of users.

  5. 2022

    DPM-Solver (Lu et al.)

    Used higher-order ODE solvers (building on DDIM's ODE connection) to push quality with 10–20 steps even further.

  6. 2023

    Consistency Models (Song et al.)

    Took DDIM's philosophy to the extreme: distill the entire ODE trajectory so the model generates a sample in a single step.

CitationSong, Meng, Ermon. Denoising Diffusion Implicit Models. ICLR, 2021.

Terms in this paper