Generative Models2020intermediate11 min read

Denoising Diffusion Probabilistic Models

نماذج الانتشار الاحتمالية لإزالة الضوضاء

Ho, J. · Jain, A. · Abbeel, P. — NeurIPS

The problem

By 2020, image generation was dominated by GANs — powerful but notoriously unstable. collapses, (generating only a few types of images), and a delicate adversarial balance made GANs hard to scale and reproduce. VAEs offered stability but produced blurry outputs. The field needed a that combined high visual fidelity with stable, principled training.

The contribution

: a generative model built on a of Gaussian steps. The gradually destroys an image by adding noise over T steps until only static remains. A neural network () learns to reverse each step — predicting and subtracting the noise. The key insight: instead of predicting the clean image directly, predict the noise that was added. A simplified training loss — just mean-squared error between true and predicted noise — made training stable and scalable. The result: sample quality rivaling the best GANs, with none of the instability.

The impact

DDPM ignited the diffusion revolution. DALL·E 2, Stable Diffusion, Midjourney, Imagen — all descend from this paper. It proved that iterative refinement could match and then surpass adversarial training, shifting the entire field of generative AI from GANs to diffusion. It also opened the door to controllable generation, video synthesis, and even robot planning via diffusion.

A sculptor working in marble removes material to reveal the statue hiding inside. DDPM works the same way: the "marble block" is pure noise, and the neural network is the sculptor — it chips away the randomness, one thin layer at a time, until a recognizable image emerges.

The genius of the method is that learning to remove noise is far easier than learning to create images from scratch, just as it is easier to teach someone to sand and polish than to carve freehand.

The generative landscape before DDPM

Before DDPM, two families of generative models dominated image synthesis:

GANs (Generative Adversarial Networks) pitted a against a in a game. They produced sharp images but suffered from training instability: mode collapse (the generator learns to produce only a few outputs), vanishing gradients for the generator, and extreme sensitivity to hyperparameters. Scaling GANs required years of engineering tricks.

VAEs (Variational Autoencoders) encoded images into a and decoded them back. Training was stable and principled (maximizing a variational lower bound), but the decoded images were systematically blurry — the model averaged over possible outputs rather than committing to one sharp reconstruction.

Diffusion models had been proposed theoretically by Sohl-Dickstein et al. (2015), but nobody had shown they could produce competitive image quality. The DDPM paper changed that.

Open in Lab
Compare the three generative paradigms: GAN, VAE, and Diffusion. Click each card to see its strengths and weaknesses.
The demo wakes as you arrive…

The forward process: dissolving images into noise

The forward process is a predetermined recipe — no learning happens here. Starting from a clean image x0x_0, we add a tiny amount of Gaussian noise at each step, producing a sequence x1,x2,…,xTx_1, x_2, \ldots, x_T. After enough steps (T≈1000T \approx 1000), the image is indistinguishable from pure static.

Think of it as slowly turning up the static on a television: at step 1 you can still see the picture clearly; by step 500 it is getting fuzzy; by step 1000 it is pure snow.

Each step is a Markov transition — the amount of noise added at step tt depends only on the image at step t−1t-1, not on any earlier step. This Markov property is what makes the math tractable.

q(xt∣xt−1)=N(xt; 1−βt xt−1,  βt I)q(x_t | x_{t-1}) = \mathcal{N}(x_t;\, \sqrt{1 - \beta_t}\, x_{t-1},\; \beta_t\, \mathbf{I})
Single forward step — adding a pinch of noise — At step t, the image is slightly shrunk by √(1−βₜ) and Gaussian noise with variance βₜ is added. βₜ is small (e.g. 0.0001 to 0.02), so each step barely changes the image — but 1000 such steps add up.
q(xt∣x0)=N(xt; αˉt x0,  (1−αˉt) I)q(x_t | x_0) = \mathcal{N}(x_t;\, \sqrt{\bar\alpha_t}\, x_0,\; (1 - \bar\alpha_t)\, \mathbf{I})
Jump to any timestep directly — ᾱₜ is the cumulative product of (1−β₁)·(1−β₂)·…·(1−βₜ). As t grows, ᾱₜ → 0, so x_t ≈ pure noise. As t → 0, ᾱₜ → 1 and x_t ≈ x_0.
Open in Lab
Drag the timestep slider to watch the image dissolve into noise. Notice how ᾱₜ drops toward zero.
The demo wakes as you arrive…

The noise schedule: controlling the pace of destruction

The sequence β1,β2,…,βT\beta_1, \beta_2, \ldots, \beta_T is the — it controls how quickly the image is destroyed. Ho et al. used a linear schedule from β1=10−4\beta_1 = 10^{-4} to βT=0.02\beta_T = 0.02 with T=1000T = 1000 steps.

Why does the schedule matter? If noise is added too fast, the must make large corrections and is prone to error. If too slow, training is inefficient and sampling takes forever. The linear schedule strikes a balance: gentle at the start (preserve structure), accelerating at the end (finish destroying the signal).

Later work (Improved DDPM, Nichol & Dhariwal 2021) explored cosine schedules that keep the dropping more uniformly, but the original linear schedule was sufficient for state-of-the-art results.

Open in Lab
Toggle between linear and cosine schedules. Watch how ᾱₜ drops differently — cosine preserves more signal in the middle timesteps.
The demo wakes as you arrive…

The reverse process: sculpting images from static

The reverse process is where the magic happens — and where all the learning occurs. Starting from pure noise xT∼N(0,I)x_T \sim \mathcal{N}(0, \mathbf{I}), the model iteratively denoises: xT→xT−1→…→x1→x0x_T \to x_{T-1} \to \ldots \to x_1 \to x_0.

At each step, the model must answer: "Given this noisy image xtx_t and the current timestep tt, what does a slightly less noisy version xt−1x_{t-1} look like?" This is parameterized as a Gaussian:

pθ(xt−1∣xt)=N(xt−1; μθ(xt,t),  σt2 I)p_\theta(x_{t-1} | x_t) = \mathcal{N}(x_{t-1};\, \mu_\theta(x_t, t),\; \sigma_t^2\, \mathbf{I})
Learned reverse step — The neural network predicts the mean μ_θ; the variance σ²ₜ is fixed to the forward-process posterior variance. The model only needs to learn which direction to "push" the noisy image.

Instead of predicting μθ\mu_\theta directly, Ho et al. made a key design choice: reparameterize the mean in terms of the predicted noise ϵθ(xt,t)\epsilon_\theta(x_t, t). The network takes a noisy image and the timestep, and outputs its best guess of the noise that was added. The mean is then computed from this prediction:

μθ(xt,t)=1αt(xt−βt1−αˉt ϵθ(xt,t))\mu_\theta(x_t, t) = \frac{1}{\sqrt{\alpha_t}} \left( x_t - \frac{\beta_t}{\sqrt{1 - \bar\alpha_t}}\, \epsilon_\theta(x_t, t) \right)
Predicting noise, then deriving the mean — Subtract the predicted noise (scaled appropriately) from the noisy input, then rescale. This is the heart of DDPM: "tell me what noise was added, and I will undo it."
Open in Lab
Click "Step" to watch the reverse process unfold — the model removes noise one step at a time. Compare the predicted noise to the actual noise.
The demo wakes as you arrive…

The training objective: simplicity wins

The full variational lower bound () for diffusion models involves a sum of KL divergences across all timesteps — complex and expensive. Ho et al.'s breakthrough was showing that a drastically simplified loss works just as well (and often better):

  1. Sample a clean image x0x_0 from the dataset
  2. Sample a random timestep t∼Uniform(1,T)t \sim \text{Uniform}(1, T)
  3. Sample random noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, \mathbf{I})
  4. Construct the noisy image: xt=αˉt x0+1−αˉt ϵx_t = \sqrt{\bar\alpha_t}\, x_0 + \sqrt{1 - \bar\alpha_t}\, \epsilon
  5. Ask the network to predict the noise: ϵθ(xt,t)\epsilon_\theta(x_t, t)
  6. Compute the loss: ∥ϵ−ϵθ(xt,t)∥2\|\epsilon - \epsilon_\theta(x_t, t)\|^2

That is it. between the true noise and the predicted noise. No adversarial games, no reconstruction-vs-KL balancing, no instability. Just: "can you guess what noise I added?"

Lsimple=Et, x0, ϵ ⁣[ ∥ϵ−ϵθ ⁣(αˉt x0+1−αˉt ϵ,  t)∥2 ]L_{\text{simple}} = \mathbb{E}_{t,\, x_0,\, \epsilon}\!\Big[\, \big\|\epsilon - \epsilon_\theta\!\big(\sqrt{\bar\alpha_t}\, x_0 + \sqrt{1 - \bar\alpha_t}\, \epsilon,\; t\big)\big\|^2\,\Big]
The simplified DDPM loss — the paper's key equation — Sample a clean image, corrupt it to timestep t using the reparameterization shortcut, predict the noise, measure the error. This single equation drives all of DDPM training.
Open in Lab
Walk through one training iteration: sample image, pick timestep, add noise, predict, compute loss.
The demo wakes as you arrive…

The U-Net backbone

The neural network ϵθ\epsilon_\theta that predicts the noise is a U-Net — an architecture with skip connections, originally designed for biomedical image segmentation. The U-Net is ideal for because it operates at the same resolution as the input (unlike a classifier which compresses to a single label) and its skip connections preserve fine spatial details across the encoding and decoding stages.

DDPM augments the basic U-Net with two critical additions:

Sinusoidal timestep embeddings — the network must know how noisy the input is. The timestep tt is encoded using sinusoidal positional embeddings (the same idea from the ) and injected into every . This lets the same network handle all noise levels, from barely-touched (t=1t = 1) to pure static (t=Tt = T).

at lower resolutions — at the 16×16 resolution level, self-attention layers let distant spatial regions communicate, improving global coherence (e.g., both eyes facing the same direction in a face).

Open in Lab
Click any layer in the U-Net to see what it does and how timestep information flows through.
The demo wakes as you arrive…

The deeper theory: score matching and Langevin dynamics

DDPM's simplified loss hides a deep connection to score-based generative models. The "score" of a distribution is the gradient of its log-probability: ∇xlog⁡p(x)\nabla_x \log p(x). It points toward regions of higher density — "which direction should I move to make this sample more likely?"

The noise-prediction network ϵθ\epsilon_\theta is implicitly learning a scaled version of the . Specifically: ϵθ(xt,t)≈−1−αˉt ∇xtlog⁡q(xt)\epsilon_\theta(x_t, t) \approx -\sqrt{1 - \bar\alpha_t}\, \nabla_{x_t} \log q(x_t).

This means DDPM's reverse process is closely related to — a physics-inspired sampling method where you follow the score (the gradient) with added noise to draw samples from a distribution. Song & Ermon (2019) had shown this approach works; Ho et al. showed the diffusion framework achieves it naturally and produces better samples.

This connection is what makes diffusion models so theoretically grounded: they are simultaneously models (optimizing a bound on likelihood), denoising autoencoders (predicting clean from noisy), and score-matching models (learning the gradient of the data distribution).

Sampling: generating images step by step

Once trained, generating an image follows the reverse process:

  1. Sample xT∼N(0,I)x_T \sim \mathcal{N}(0, \mathbf{I}) — start from pure noise
  2. For t=T,T−1,…,1t = T, T-1, \ldots, 1: — Predict ϵθ(xt,t)\epsilon_\theta(x_t, t) — Compute μθ(xt,t)\mu_\theta(x_t, t) using the formula above — Sample xt−1∼N(μθ,σt2I)x_{t-1} \sim \mathcal{N}(\mu_\theta, \sigma_t^2 \mathbf{I})
  3. Return x0x_0 — the generated image

This is the algorithm. It requires running the full T = 1000 steps, making it slow compared to a single-pass . However, each step is a simple through the U-Net, and the quality of the result is remarkable.

Results: diffusion matches the best GANs

On unconditional image generation on CIFAR-10, DDPM achieved an score of 3.17 — matching ProgressiveGAN which held the state of the art at the time. On LSUN (bedrooms, churches, horses), it produced diverse, high-fidelity 256×256 samples.

Crucially, DDPM achieved this without any adversarial training — no discriminator, no minimax game, no mode collapse. Training was as stable as training a classifier: just on the simple loss.

The paper also demonstrated that diffusion models naturally produce a progressive lossy decompression scheme: each reverse step reveals more detail, from coarse structure to fine texture, analogous to how progressive JPEG decoding works.

The same idea in code

DDPM training loop — the complete algorithmpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn.functional as F

def ddpm_train_step(model, x_0, betas, alpha_bars):
    """One training step: sample noise, predict it, compute loss."""
    T = len(betas)

    # 1. Sample a random timestep
    t = torch.randint(0, T, (x_0.shape[0],))

    # 2. Sample Gaussian noise
    epsilon = torch.randn_like(x_0)

    # 3. Construct noisy image via reparameterization
    alpha_bar_t = alpha_bars[t].view(-1, 1, 1, 1)
    x_t = torch.sqrt(alpha_bar_t) * x_0 + torch.sqrt(1 - alpha_bar_t) * epsilon

    # 4. Predict the noise
    epsilon_pred = model(x_t, t)

    # 5. Simple MSE loss — that's it!
    loss = F.mse_loss(epsilon_pred, epsilon)
    return loss

@torch.no_grad()
def ddpm_sample(model, shape, betas, alpha_bars):
    """Generate an image by iterative denoising."""
    T = len(betas)
    alphas = 1.0 - betas

    # Start from pure noise
    x = torch.randn(shape)

    for t in reversed(range(T)):
        # Predict and remove noise
        eps = model(x, torch.tensor([t]))
        mean = (1 / alphas[t].sqrt()) * (
            x - (betas[t] / (1 - alpha_bars[t]).sqrt()) * eps
        )
        if t > 0:
            x = mean + betas[t].sqrt() * torch.randn_like(x)
        else:
            x = mean  # no noise at final step

    return x

Why it mattered

  1. 2015

    Sohl-Dickstein et al. — Diffusion Conceived

    First paper proposing diffusion probabilistic models. Proved the theory works but image quality was far from competitive.

  2. 2019

    Song & Ermon — Score Matching with Langevin

    Showed that learning the score function at multiple noise levels enables high-quality image generation via Langevin dynamics.

  3. 2020

    Ho et al. — DDPM

    This paper. Simplified the loss, used a U-Net backbone, and matched GAN quality. Ignited the diffusion revolution.

  4. 2020

    DDIM — Fast Sampling

    Song et al. showed how to skip steps in the reverse process, reducing sampling from 1000 to 50 steps with minimal quality loss.

  5. 2021

    Dhariwal & Nichol — Diffusion Beats GANs

    Classifier guidance and architectural improvements pushed diffusion past GANs on ImageNet FID for the first time.

  6. 2022

    Latent Diffusion / Stable Diffusion

    Rombach et al. moved diffusion to a compressed latent space, enabling high-resolution text-to-image generation at consumer scale.

  7. 2022

    DALL·E 2 & Imagen

    Text-to-image systems from OpenAI and Google, both built on diffusion. Showed the paradigm generalizes to language-guided generation.

  8. 2023

    Diffusion Policy

    Chi et al. applied diffusion to robot action planning — generating smooth, multi-modal trajectories instead of images.

From a single idea — "learn to predict and remove noise" — an entire ecosystem emerged. Every major text-to-image, video generation, and 3D synthesis system today traces its lineage back to this paper. DDPM proved that patience (1000 small steps) beats brute force (one-shot adversarial generation), and that simplicity in training objectives can yield extraordinary results.

CitationHo, Jain, Abbeel. Denoising Diffusion Probabilistic Models. NeurIPS, 2020.

Terms in this paper