Generative Models2020intermediate11 min read
Denoising Diffusion Probabilistic Models
نماذج الانتشار الاحتمالية لإزالة الضوضاء
Ho, J. · Jain, A. · Abbeel, P. — NeurIPS
The problem
By 2020, image generation was dominated by GANs — powerful but notoriously unstable. collapses, (generating only a few types of images), and a delicate adversarial balance made GANs hard to scale and reproduce. VAEs offered stability but produced blurry outputs. The field needed a that combined high visual fidelity with stable, principled training.
The contribution
: a generative model built on a of Gaussian steps. The gradually destroys an image by adding noise over T steps until only static remains. A neural network () learns to reverse each step — predicting and subtracting the noise. The key insight: instead of predicting the clean image directly, predict the noise that was added. A simplified training loss — just mean-squared error between true and predicted noise — made training stable and scalable. The result: sample quality rivaling the best GANs, with none of the instability.
The impact
DDPM ignited the diffusion revolution. DALL·E 2, Stable Diffusion, Midjourney, Imagen — all descend from this paper. It proved that iterative refinement could match and then surpass adversarial training, shifting the entire field of generative AI from GANs to diffusion. It also opened the door to controllable generation, video synthesis, and even robot planning via diffusion.
A sculptor working in marble removes material to reveal the statue hiding inside. DDPM works the same way: the "marble block" is pure noise, and the neural network is the sculptor — it chips away the randomness, one thin layer at a time, until a recognizable image emerges.
The genius of the method is that learning to remove noise is far easier than learning to create images from scratch, just as it is easier to teach someone to sand and polish than to carve freehand.
The generative landscape before DDPM
Before DDPM, two families of generative models dominated image synthesis:
GANs (Generative Adversarial Networks) pitted a against a in a game. They produced sharp images but suffered from training instability: mode collapse (the generator learns to produce only a few outputs), vanishing gradients for the generator, and extreme sensitivity to hyperparameters. Scaling GANs required years of engineering tricks.
VAEs (Variational Autoencoders) encoded images into a and decoded them back. Training was stable and principled (maximizing a variational lower bound), but the decoded images were systematically blurry — the model averaged over possible outputs rather than committing to one sharp reconstruction.
Diffusion models had been proposed theoretically by Sohl-Dickstein et al. (2015), but nobody had shown they could produce competitive image quality. The DDPM paper changed that.
The forward process: dissolving images into noise
The forward process is a predetermined recipe — no learning happens here. Starting from a clean image , we add a tiny amount of Gaussian noise at each step, producing a sequence . After enough steps (), the image is indistinguishable from pure static.
Think of it as slowly turning up the static on a television: at step 1 you can still see the picture clearly; by step 500 it is getting fuzzy; by step 1000 it is pure snow.
Each step is a Markov transition — the amount of noise added at step depends only on the image at step , not on any earlier step. This Markov property is what makes the math tractable.
The noise schedule: controlling the pace of destruction
The sequence is the — it controls how quickly the image is destroyed. Ho et al. used a linear schedule from to with steps.
Why does the schedule matter? If noise is added too fast, the must make large corrections and is prone to error. If too slow, training is inefficient and sampling takes forever. The linear schedule strikes a balance: gentle at the start (preserve structure), accelerating at the end (finish destroying the signal).
Later work (Improved DDPM, Nichol & Dhariwal 2021) explored cosine schedules that keep the dropping more uniformly, but the original linear schedule was sufficient for state-of-the-art results.
The reverse process: sculpting images from static
The reverse process is where the magic happens — and where all the learning occurs. Starting from pure noise , the model iteratively denoises: .
At each step, the model must answer: "Given this noisy image and the current timestep , what does a slightly less noisy version look like?" This is parameterized as a Gaussian:
Instead of predicting directly, Ho et al. made a key design choice: reparameterize the mean in terms of the predicted noise . The network takes a noisy image and the timestep, and outputs its best guess of the noise that was added. The mean is then computed from this prediction:
The training objective: simplicity wins
The full variational lower bound () for diffusion models involves a sum of KL divergences across all timesteps — complex and expensive. Ho et al.'s breakthrough was showing that a drastically simplified loss works just as well (and often better):
- Sample a clean image from the dataset
- Sample a random timestep
- Sample random noise
- Construct the noisy image:
- Ask the network to predict the noise:
- Compute the loss:
That is it. between the true noise and the predicted noise. No adversarial games, no reconstruction-vs-KL balancing, no instability. Just: "can you guess what noise I added?"
The U-Net backbone
The neural network that predicts the noise is a U-Net — an architecture with skip connections, originally designed for biomedical image segmentation. The U-Net is ideal for because it operates at the same resolution as the input (unlike a classifier which compresses to a single label) and its skip connections preserve fine spatial details across the encoding and decoding stages.
DDPM augments the basic U-Net with two critical additions:
Sinusoidal timestep embeddings — the network must know how noisy the input is. The timestep is encoded using sinusoidal positional embeddings (the same idea from the ) and injected into every . This lets the same network handle all noise levels, from barely-touched () to pure static ().
at lower resolutions — at the 16×16 resolution level, self-attention layers let distant spatial regions communicate, improving global coherence (e.g., both eyes facing the same direction in a face).
The deeper theory: score matching and Langevin dynamics
DDPM's simplified loss hides a deep connection to score-based generative models. The "score" of a distribution is the gradient of its log-probability: . It points toward regions of higher density — "which direction should I move to make this sample more likely?"
The noise-prediction network is implicitly learning a scaled version of the . Specifically: .
This means DDPM's reverse process is closely related to — a physics-inspired sampling method where you follow the score (the gradient) with added noise to draw samples from a distribution. Song & Ermon (2019) had shown this approach works; Ho et al. showed the diffusion framework achieves it naturally and produces better samples.
This connection is what makes diffusion models so theoretically grounded: they are simultaneously models (optimizing a bound on likelihood), denoising autoencoders (predicting clean from noisy), and score-matching models (learning the gradient of the data distribution).
Sampling: generating images step by step
Once trained, generating an image follows the reverse process:
- Sample — start from pure noise
- For : — Predict — Compute using the formula above — Sample
- Return — the generated image
This is the algorithm. It requires running the full T = 1000 steps, making it slow compared to a single-pass . However, each step is a simple through the U-Net, and the quality of the result is remarkable.
Results: diffusion matches the best GANs
On unconditional image generation on CIFAR-10, DDPM achieved an score of 3.17 — matching ProgressiveGAN which held the state of the art at the time. On LSUN (bedrooms, churches, horses), it produced diverse, high-fidelity 256×256 samples.
Crucially, DDPM achieved this without any adversarial training — no discriminator, no minimax game, no mode collapse. Training was as stable as training a classifier: just on the simple loss.
The paper also demonstrated that diffusion models naturally produce a progressive lossy decompression scheme: each reverse step reveals more detail, from coarse structure to fine texture, analogous to how progressive JPEG decoding works.
The same idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn.functional as F
def ddpm_train_step(model, x_0, betas, alpha_bars):
"""One training step: sample noise, predict it, compute loss."""
T = len(betas)
# 1. Sample a random timestep
t = torch.randint(0, T, (x_0.shape[0],))
# 2. Sample Gaussian noise
epsilon = torch.randn_like(x_0)
# 3. Construct noisy image via reparameterization
alpha_bar_t = alpha_bars[t].view(-1, 1, 1, 1)
x_t = torch.sqrt(alpha_bar_t) * x_0 + torch.sqrt(1 - alpha_bar_t) * epsilon
# 4. Predict the noise
epsilon_pred = model(x_t, t)
# 5. Simple MSE loss — that's it!
loss = F.mse_loss(epsilon_pred, epsilon)
return loss
@torch.no_grad()
def ddpm_sample(model, shape, betas, alpha_bars):
"""Generate an image by iterative denoising."""
T = len(betas)
alphas = 1.0 - betas
# Start from pure noise
x = torch.randn(shape)
for t in reversed(range(T)):
# Predict and remove noise
eps = model(x, torch.tensor([t]))
mean = (1 / alphas[t].sqrt()) * (
x - (betas[t] / (1 - alpha_bars[t]).sqrt()) * eps
)
if t > 0:
x = mean + betas[t].sqrt() * torch.randn_like(x)
else:
x = mean # no noise at final step
return xWhy it mattered
2015
Sohl-Dickstein et al. — Diffusion Conceived
First paper proposing diffusion probabilistic models. Proved the theory works but image quality was far from competitive.
2019
Song & Ermon — Score Matching with Langevin
Showed that learning the score function at multiple noise levels enables high-quality image generation via Langevin dynamics.
2020
Ho et al. — DDPM
This paper. Simplified the loss, used a U-Net backbone, and matched GAN quality. Ignited the diffusion revolution.
2020
DDIM — Fast Sampling
Song et al. showed how to skip steps in the reverse process, reducing sampling from 1000 to 50 steps with minimal quality loss.
2021
Dhariwal & Nichol — Diffusion Beats GANs
Classifier guidance and architectural improvements pushed diffusion past GANs on ImageNet FID for the first time.
2022
Latent Diffusion / Stable Diffusion
Rombach et al. moved diffusion to a compressed latent space, enabling high-resolution text-to-image generation at consumer scale.
2022
DALL·E 2 & Imagen
Text-to-image systems from OpenAI and Google, both built on diffusion. Showed the paradigm generalizes to language-guided generation.
2023
Diffusion Policy
Chi et al. applied diffusion to robot action planning — generating smooth, multi-modal trajectories instead of images.
From a single idea — "learn to predict and remove noise" — an entire ecosystem emerged. Every major text-to-image, video generation, and 3D synthesis system today traces its lineage back to this paper. DDPM proved that patience (1000 small steps) beats brute force (one-shot adversarial generation), and that simplicity in training objectives can yield extraordinary results.
CitationHo, Jain, Abbeel. Denoising Diffusion Probabilistic Models. NeurIPS, 2020.
Terms in this paper
- DDPMنموذج الانتشار الاحتمالي لإزالة التشويش
- Diffusion Modelنموذج الانتشار
- Denoisingإزالة الضوضاء
- Forward Processالعملية الأمامية
- Reverse Processالعملية العكسية
- Markov Chainسلسلة ماركوف الاحتمالية
- Noise Scheduleجدول الضوضاء
- Reparameterization Trickحيلة إعادة البَرمَتَة
- U-Netشبكة U-Net
- Ancestral Samplingالسحب الأصلي
- Score Matchingمطابقة النتيجة