Generative Models2021intermediate11 min read
Denoising Diffusion Implicit Models
نماذج الانتشار الضمنية لإزالة الضجيج
Song, J. · Meng, C. · Ermon, S. — ICLR
The problem
By 2020, diffusion probabilistic models (DDPMs) were generating images rivaling GANs in quality — without the instability of adversarial . But they had a crippling weakness: speed. Generating a single image required simulating a 1,000-step , making them orders of magnitude slower than GANs. 50,000 CIFAR-10 images took about 20 hours on a single GPU, compared to under a minute for a GAN. For high-resolution images, the wait stretched to nearly 1,000 hours. This made DDPMs impractical for any latency-sensitive application.
The contribution
generalizes the framework by replacing the Markovian with a family of non-Markovian processes that share the same marginal distributions and therefore the same training objective. When the stochasticity parameter σ is set to zero, the resulting generative process becomes fully deterministic — an that maps a latent noise vector to a unique image. This deterministic shortcut allows sampling in as few as 10–50 steps while maintaining quality comparable to the original 1,000-step DDPM. The approach requires no retraining: any pretrained DDPM works directly. DDIM also enables meaningful and near-perfect reconstruction via an ODE-like encoding–decoding cycle.
The impact
DDIM made diffusion models practical. Its accelerated deterministic sampling became the default sampler in Stable Diffusion, DALL·E 2, and virtually every modern image generator. The deterministic mapping between noise and image opened the door to latent manipulation techniques used in image editing. The connection to neural ODEs inspired a wave of follow-up work on faster samplers, including DPM-Solver and consistency models. By proving that a single pretrained model supports an entire family of generative processes, DDIM reshaped how the community thinks about the relationship between training and in diffusion models.
Imagine restoring a Renaissance painting that has been completely covered in dust. The traditional method (DDPM) insists on wiping the painting with exactly 1,000 gentle cloth strokes, each removing a microscopic layer of dust — and the direction of each stroke is chosen by a coin flip.
DDIM discovers something remarkable: because you already know the painting's structure from training, you can take a direct path to the clean image. Instead of 1,000 random wipes, you take 50 precise, calculated strokes — deterministic, no coin flips — and arrive at the same restored masterpiece. Better yet, you don't need to learn any new technique; the same training that prepared you for 1,000 strokes works perfectly for 50.
The bottleneck: 1,000 mandatory steps
To understand DDIM, we first need to see why DDPM is slow. In DDPM, the forward process gradually adds Gaussian noise to a clean image over steps, producing a sequence that ends in pure noise. This is a Markov chain: each step depends only on the previous step .
The generative process reverses this chain: starting from pure noise , it denoises step by step until reaching a clean image . Because the forward process has steps, the must also take steps — one neural network evaluation per step. There is no shortcut in a Markov chain.
A key property makes this possible: we can jump from to any step directly:
This is the crucial observation: the marginal distribution at each step does not depend on the path taken to get there. Whether we arrived through 1,000 Markov steps or jumped directly, the statistics are the same. The DDPM training objective (predicting the noise from ) only depends on these marginals — not on the joint distribution . This is the door that DDIM opens.
The key insight: many paths, same destination
The Markov chain in DDPM is just one particular joint distribution that produces the marginals . But there are infinitely many other joint distributions that produce the exact same marginals.
Think of it like this: imagine you know the weather statistics for every month of the year. The Markov model says "today's weather depends only on yesterday." But you could also have a model where today's weather depends on yesterday and on the overall season — a non-Markovian model. Both can produce the same monthly averages, but the day-to-day transitions look very different.
DDIM exploits this freedom. It defines a family of non-Markovian forward processes, indexed by a vector , where each controls how much randomness enters at step . The family is designed so that the marginals are identical to DDPM's for every . Since the training objective depends only on these marginals, the same pretrained model works for the entire family — no retraining needed.
The update rule: three forces in one equation
The heart of DDIM is a single update equation that generates from . Before we see the math, let's understand what it's doing.
The model has been trained to predict the noise that was added to the clean image. Given this prediction, we can estimate what the clean image looks like. Then we "re-noise" this estimate to the correct level for step . The equation blends three components: the predicted clean image, a correction term that points in the "noise direction," and an optional random noise term controlled by .
Think of it as a tug-of-war between "signal" and "noise." The first term pulls toward the clean image — like a compass pointing home. The second term adjusts the noise level so we land at the right position on the denoising schedule. The third term adds controlled randomness that is useful for diversity but slows convergence.
The parameter provides a convenient way to interpolate. Setting gives DDIM (deterministic), gives DDPM (stochastic), and values in between give a smooth continuum. This is achieved by defining .
The acceleration trick: skip steps, keep quality
The second breakthrough is that we don't need to reverse all steps. Since the training objective only depends on marginals , we can choose a subsequence of the original and run the generative process only on those steps. Instead of going through , we might jump — taking only 10 steps.
The update equation works exactly the same, just with and replacing and . The model was trained on all 1,000 noise levels, so it can denoise at any of them. We're simply choosing to visit fewer waypoints on the journey from noise to image.
The experimental results are striking. On CIFAR-10, DDIM with 50 steps achieves an of 4.67, comparable to DDPM's 4.04 with 1,000 steps — a 20× speedup with minimal quality loss. With 20 steps, DDIM still scores 6.84, while DDPM collapses to 18.36 under the same budget.
The bonus: a meaningful latent space
Because DDIM is deterministic, the same initial noise always produces the same image — regardless of how many steps we use. This is called consistency: a 10-step and a 1,000-step generation from identical share the same high-level features (pose, color, composition), differing only in fine details.
This transforms into a proper latent code, similar to the latent space in GANs or variational autoencoders. We gain two powerful capabilities. First, interpolation: given two noise vectors and , we can smoothly walk between them using spherical linear interpolation and generate all intermediate images. The transitions are semantically meaningful — faces morph smoothly, scenes blend naturally.
Second, encoding: by running the DDIM update in reverse (from to ), we can encode any image into the latent space, then decode it back with near-perfect reconstruction. The reconstruction error drops below with 1,000 encoding steps. This is impossible with DDPM because its stochastic process destroys the one-to-one mapping.
The deeper connection: DDIM as an ODE solver
There is a beautiful mathematical connection hidden beneath the surface. If we reparametrize by dividing each by and defining , the DDIM update becomes a single Euler step of an :
This ODE perspective is powerful for two reasons. First, it means we can apply decades of research on numerical ODE solvers (Adams-Bashforth, Runge-Kutta) to improve sampling quality in fewer steps. Second, it gives us a reversible encoding: unlike DDPM, where stochasticity makes the mapping from to one-to-many, the DDIM ODE defines a bijection — every image has exactly one latent code, and vice versa. This makes DDIM a normalizing flow by another name, with the noise prediction network playing the role of the .
Simplified to show the idea — not the real implementation.
import torch
def ddim_sample(model, alphas, tau, eta=0.0):
"""
model : trained noise predictor epsilon_theta
alphas : full alpha schedule [alpha_1, ..., alpha_T]
tau : subsequence of step indices, e.g. [100, 200, ..., 1000]
eta : 0 = DDIM (deterministic), 1 = DDPM (stochastic)
"""
# Start from pure noise
x = torch.randn_like(alphas[0]) # x_T ~ N(0, I)
for i in reversed(range(len(tau))):
t = tau[i]
t_prev = tau[i - 1] if i > 0 else 0
a_t = alphas[t]
a_prev = alphas[t_prev] if t_prev > 0 else 1.0
# Predict noise
eps = model(x, t)
# Estimate clean image (predicted x_0)
x0_pred = (x - (1 - a_t).sqrt() * eps) / a_t.sqrt()
# Compute sigma for this step
sigma = eta * ((1 - a_prev) / (1 - a_t)).sqrt() * (1 - a_t / a_prev).sqrt()
# Direction pointing to x_t
dir_xt = (1 - a_prev - sigma**2).sqrt() * eps
# Combine: predicted x_0 + direction + noise
x = a_prev.sqrt() * x0_pred + dir_xt + sigma * torch.randn_like(x)
return x
Why this matters: from theory to Stable Diffusion
DDIM solved the practical barrier that stood between diffusion models and real-world deployment. Before DDIM, diffusion models were an academic curiosity — beautiful math, excellent sample quality, but too slow to use. After DDIM, diffusion models became commercially viable.
The impact was immediate and far-reaching. Every modern diffusion-based system — Stable Diffusion, DALL·E 2, Imagen, Midjourney — uses DDIM-style accelerated sampling as a fundamental building block. The deterministic latent space enabled image editing techniques like DDIM inversion, where real photos are encoded into latent space, edited, and decoded back. The ODE connection inspired DPM-Solver, which pushed the boundary further to 10–20 steps with higher-order ODE solvers. And the core idea — that a single trained model supports many inference strategies — directly inspired consistency models, which distill the entire trajectory into a one-step generator.
Legacy: accelerating diffusion
2020
DDPM (Ho et al.)
Demonstrated that diffusion models can match GAN quality on images — but required 1,000 sequential denoising steps, making sampling extremely slow.
2021
DDIM (this paper)
Introduced non-Markovian forward processes and deterministic sampling, cutting steps from 1,000 to 50 with minimal quality loss — and no retraining.
2021
Score SDE (Song et al.)
Unified diffusion models under a continuous-time SDE framework, with a probability flow ODE equivalent to DDIM's formulation.
2022
DALL·E 2 & Stable Diffusion
Commercial image generators adopted DDIM-style sampling as their default, bringing diffusion models to millions of users.
2022
DPM-Solver (Lu et al.)
Used higher-order ODE solvers (building on DDIM's ODE connection) to push quality with 10–20 steps even further.
2023
Consistency Models (Song et al.)
Took DDIM's philosophy to the extreme: distill the entire ODE trajectory so the model generates a sample in a single step.
CitationSong, Meng, Ermon. Denoising Diffusion Implicit Models. ICLR, 2021.
Terms in this paper
- Diffusion Modelنموذج الانتشار
- DDPMنموذج الانتشار الاحتمالي لإزالة التشويش
- Denoisingإزالة الضوضاء
- Markov Chainسلسلة ماركوف الاحتمالية
- Forward Processالعملية الأمامية
- Reverse Processالعملية العكسية
- Noise Scheduleجدول الضوضاء
- Generative Modelالنموذج التوليدي
- Latent Variableالمتغير الكامن
- Latent Spaceالفضاء الكامن
- Interpolationالاستيفاء
- Samplingاختيار العينات الاحتمالية
- Variational Inferenceالاستدلال الاحتمالي المتغيِّر
- ELBOالحد الأدنى الاختلافي
- Neural ODEالمعادلة التفاضلية العصبية