Generative Models2021advanced12 min read

Diffusion Models Beat GANs on Image Synthesis

نماذج الانتشار تتفوّق على شبكات GAN في توليد الصور

Dhariwal, P. · Nichol, A. — NeurIPS

The problem

By 2021, GANs dominated image generation in terms of sample quality metrics like and . However, GANs suffered from mode collapse, instability, and limited diversity. Diffusion models showed promise — they covered the data distribution better and trained more stably — but their sample quality still lagged behind GANs on challenging datasets like ImageNet. Two key gaps remained: architectures had not been optimized as extensively as architectures, and there was no mechanism for diffusion models to trade off diversity for fidelity the way GANs do with truncation.

The contribution

Two complementary breakthroughs. First, a systematic that improved the architecture for diffusion models: deeper at multiple resolutions, more attention heads, BigGAN-style residual blocks, and . This alone achieved state-of-the-art on unconditional generation (LSUN). Second, — using gradients from a pretrained classifier to steer the process toward a target class. A single hyperparameter (the scale) smoothly trades diversity for fidelity. Together, these achieved FID 2.97 on ImageNet 128×128 and 4.59 on 256×256, surpassing BigGAN-deep while maintaining better distribution coverage.

The impact

This paper marked the moment diffusion models overtook GANs as the leading paradigm for image generation. Classifier guidance directly inspired (Ho & Salimans, 2021), which eliminated the need for a separate classifier and became the standard conditioning mechanism in DALL·E 2, Imagen, Stable Diffusion, and GLIDE. The improved U-Net architecture (ADM) became the backbone of nearly all subsequent diffusion image generators. The paper effectively ended the GAN era of dominance and launched the diffusion era that continues today.

Imagine two art studios competing to produce photorealistic paintings. Studio GAN uses a brilliant but temperamental artist paired with a harsh critic — together they produce stunning images, but the artist keeps painting the same few subjects and throws tantrums if the critic's feedback changes slightly.

Studio Diffusion takes a different approach: start with a canvas covered in static, and have a patient restorer clean it one careful brushstroke at a time. The restorer is stable, never throws tantrums, and naturally produces a wide variety of subjects. The problem? The restorations were slightly blurry compared to Studio GAN's paintings.

This paper gave Studio Diffusion two upgrades: a better set of restoration tools (architecture improvements) and a museum guide who whispers "make it look more like a corgi" during the cleaning process (classifier guidance). Suddenly, Studio Diffusion produced sharper paintings than Studio GAN — and with more variety.

Refresher: how diffusion models generate images

A diffusion model generates images by reversing a gradual noising process. During training, real images are progressively corrupted by adding Gaussian noise across TT timesteps, until the image becomes pure static. The model learns to undo this corruption one step at a time — predicting and removing the noise component at each timestep.

Think of it like a highway with TT exits. At exit 0, you have a clean image. At exit TT, you have pure noise. The drives the image from exit 0 to exit TT, adding a little noise at each exit. The model learns the reverse trip: starting from pure noise at exit TT and gradually denoising back to exit 0.

The noise predictor ϵθ(xt,t)\epsilon_\theta(x_t, t) takes a noisy image xtx_t and the current timestep tt as input, and predicts the noise ϵ\epsilon that was added. Training uses a simple loss between the predicted noise and the actual noise.

Lsimple=Et,x0,ϵ[∥ϵ−ϵθ(xt,t)∥2]L_{\text{simple}} = \mathbb{E}_{t, x_0, \epsilon}\left[\|\epsilon - \epsilon_\theta(x_t, t)\|^2\right]
DDPM training objective — predict the noise — For each training step, a clean image x0x_0 is sampled from the dataset, a random timestep tt is chosen, and Gaussian noise ϵ\epsilon is added to produce xtx_t. The model predicts ϵ\epsilon and the loss is the squared difference. This simple objective is surprisingly effective and is equivalent to denoising score matching.
Open in Lab
Watch noise gradually destroy and then restore an image. The forward process adds noise; the reverse process (learned by the model) removes it step by step.
The demo wakes as you arrive…

Breakthrough 1: a better U-Net through systematic ablations

The first hypothesis: the gap between diffusion models and GANs comes partly from under-optimized architectures. GAN architectures like StyleGAN2 and BigGAN had been refined through years of ablation studies. Diffusion models were still using the original U-Net from with minimal changes.

Dhariwal and Nichol conducted a systematic ablation study on ImageNet 128×128, testing one change at a time. They discovered that several modifications compound to give a substantial boost in FID, and most of these are inspired by best practices from GAN and Transformer literature.

Open in Lab
Toggle each architecture change to see its impact on FID. Changes compound — enable them all to see the full improvement.
The demo wakes as you arrive…

The key modifications, each tested independently and then combined, are:

More depth vs. width. Increasing depth (more residual blocks per resolution) while holding model size constant improves FID, but slows training. Wider models reach the same quality faster, so the authors chose width over depth.

Multi-resolution attention. Instead of applying attention only at 16×16 resolution (as in DDPM), the improved model applies attention at 32×32, 16×16, and 8×8 resolutions. This helps the model capture both fine details and global structure.

More attention heads. Increasing from 1 head to multiple heads with 64 channels per head significantly improves FID. This matches modern Transformer practice, where many small heads outperform few large ones.

BigGAN residual blocks. Using BigGAN-style and downsampling residual blocks instead of simple convolutions improves information flow through the network.

Adaptive (AdaGN). Instead of adding timestep and class embeddings via simple addition, AdaGN modulates the normalization parameters: AdaGN(h,y)=ys⋅GroupNorm(h)+yb\text{AdaGN}(h, y) = y_s \cdot \text{GroupNorm}(h) + y_b, where ysy_s and yby_b come from a linear projection of the timestep and class embedding. This is analogous to adaptive instance normalization from StyleGAN.

AdaGN(h,y)=ys⋅GroupNorm(h)+yb\text{AdaGN}(h, y) = y_s \cdot \text{GroupNorm}(h) + y_b
Adaptive Group Normalization — conditioning the denoiser — After the first convolution in each residual block, the intermediate features hh are normalized by GroupNorm. Then ysy_s (scale) and yby_b (bias), derived from a linear projection of the timestep and class embedding, modulate the normalized output. This allows each residual block to adapt its behavior to both the current noise level and the target class.

Breakthrough 2: classifier guidance

The second insight addresses a fundamental advantage that GANs had: they could trade off diversity for fidelity using the . A smaller truncation range produces sharper but less diverse images. Diffusion models had no equivalent knob.

The solution: classifier guidance. The idea is beautifully simple. Train a separate classifier pϕ(y∣xt,t)p_\phi(y|x_t, t) on noisy images at every noise level. During , use the classifier's gradients to nudge each denoising step toward a target class yy. The classifier acts like a museum guide whispering directions to the restorer: "this should look more like a corgi."

Mathematically, the conditional becomes:

pθ,ϕ(xt∣xt+1,y)=Z pθ(xt∣xt+1) pϕ(y∣xt)p_{\theta,\phi}(x_t | x_{t+1}, y) = Z \, p_\theta(x_t | x_{t+1}) \, p_\phi(y | x_t)
Conditional reverse step — combining denoiser with classifier — The unconditional denoising distribution pθ(xt∣xt+1)p_\theta(x_t|x_{t+1}) is multiplied by the classifier probability pϕ(y∣xt)p_\phi(y|x_t). This biases the sampling toward images that the classifier recognizes as class yy. ZZ is a normalizing constant. In practice, this product is approximated as a Gaussian with a shifted mean.

Under reasonable assumptions (the classifier's log-probability has low curvature compared to the of the Gaussian transition), this product simplifies elegantly. The conditional transition remains a Gaussian — like the unconditional one — but with its mean shifted by Σg\Sigma g, where g=∇xtlog⁡pϕ(y∣xt)g = \nabla_{x_t} \log p_\phi(y|x_t) is the classifier gradient and Σ\Sigma is the predicted .

The sampling algorithm is simple: at each timestep, compute the unconditional mean μ\mu and variance Σ\Sigma, then sample from N(μ+sΣg,Σ)\mathcal{N}(\mu + s\Sigma g, \Sigma), where ss is a gradient scale factor.

xt−1∼N ⁣(μθ(xt)+s Σθ(xt) ∇xt ⁣log⁡pϕ(y∣xt),  Σθ(xt))x_{t-1} \sim \mathcal{N}\!\bigl(\mu_\theta(x_t) + s\,\Sigma_\theta(x_t)\,\nabla_{x_t}\!\log p_\phi(y|x_t),\;\Sigma_\theta(x_t)\bigr)
Classifier-guided sampling step — At each step, the mean is nudged by s⋅Σ⋅gs \cdot \Sigma \cdot g, where ss is the guidance scale. When s=0s = 0, this is standard unconditional sampling. When s>0s > 0, the samples are pushed toward images the classifier is confident about. Larger ss produces sharper, more class-consistent images at the cost of diversity.
Open in Lab
Adjust the guidance scale to see how classifier gradients steer the denoising process. Higher scale = sharper images but less diversity.
The demo wakes as you arrive…

The gradient scale: a single knob for quality vs. diversity

A key practical discovery: using the classifier gradients at their natural scale (s=1s = 1) is not enough. While the classifier assigns around 50% probability to the target class, the resulting images do not visually match the class. The gradients need to be amplified.

Scaling the gradients by s>1s > 1 has a clean theoretical interpretation. It is equivalent to sampling from a sharpened classifier distribution p(y∣x)sp(y|x)^s. When s>1s > 1, this distribution becomes more peaked — it concentrates probability on the most confident class predictions. Think of it as turning up the "conviction" of the museum guide.

The authors found that the scale can be increased by an order of magnitude (up to s=10s = 10) without producing adversarial examples. This smoothly trades off (diversity) for precision (fidelity), with FID reaching its optimum at an intermediate scale.

s⋅∇xlog⁡p(y∣x)=∇xlog⁡1Zp(y∣x)ss \cdot \nabla_x \log p(y|x) = \nabla_x \log \frac{1}{Z} p(y|x)^s
Gradient scaling as distribution sharpening — Multiplying the classifier gradient by ss is mathematically equivalent to sampling from a distribution proportional to p(y∣x)sp(y|x)^s. When s>1s > 1, this sharpened distribution concentrates on images the classifier is most confident about. This is the mechanism that enables the diversity-fidelity trade-off.
Open in Lab
Drag the guidance scale and watch how precision and recall change. FID has a sweet spot where both fidelity and diversity are balanced.
The demo wakes as you arrive…

Results: surpassing GANs across the board

The improved architecture alone (ADM) achieved state-of-the-art FID on LSUN bedrooms (1.90), horses (2.57), and cats (5.57) — all unconditional, beating StyleGAN2. On ImageNet 64×64 the improved model reached FID 2.07.

With classifier guidance (ADM-G), diffusion models surpassed BigGAN-deep on ImageNet at every resolution: FID 2.97 at 128×128, 4.59 at 256×256, and 7.72 at 512×512. Crucially, the diffusion models maintained higher recall (distribution coverage) than GANs at comparable or better precision.

Even with only 25 sampling steps using , the guided model matched BigGAN-deep's FID while still covering more of the distribution.

Finally, combining classifier guidance with upsampling diffusion models gave the best overall results: FID 3.94 on 256×256 and 3.85 on 512×512. The authors showed that guidance and upsampling improve quality along different axes — guidance boosts fidelity while upsampling improves spatial resolution — so they complement each other.

Open in Lab
Compare FID scores across models and resolutions. Lower is better. ADM-G beats BigGAN-deep at every resolution.
The demo wakes as you arrive…

Classifier guidance for DDIM: the score-based view

The Gaussian perturbation derivation above only works for stochastic sampling. For deterministic samplers like DDIM, the authors use a score-based trick. The noise predictor ϵθ(xt)\epsilon_\theta(x_t) can be reinterpreted as a scaled :

∇xtlog⁡pθ(xt)=−11−αˉtϵθ(xt)\nabla_{x_t} \log p_\theta(x_t) = -\frac{1}{\sqrt{1-\bar\alpha_t}} \epsilon_\theta(x_t)

Adding the classifier gradient gives the joint score:

∇xtlog⁡(pθ(xt)⋅pϕ(y∣xt))=∇xtlog⁡pθ(xt)+∇xtlog⁡pϕ(y∣xt)\nabla_{x_t} \log(p_\theta(x_t) \cdot p_\phi(y|x_t)) = \nabla_{x_t}\log p_\theta(x_t) + \nabla_{x_t}\log p_\phi(y|x_t)

This translates into a modified noise prediction ϵ^(xt)=ϵθ(xt)−1−αˉt ∇xtlog⁡pϕ(y∣xt)\hat\epsilon(x_t) = \epsilon_\theta(x_t) - \sqrt{1-\bar\alpha_t}\,\nabla_{x_t}\log p_\phi(y|x_t) that can be plugged directly into the DDIM sampler. The classifier gradient literally "subtracts class-relevant noise" from the prediction, so the denoiser favors the target class.

Combining guidance with upsampling: complementary strengths

The authors also explored two-stage upsampling stacks, where a low-resolution diffusion model generates a small image and then a separate upsampling model enlarges it. They found that guidance and upsampling improve quality along different axes:

Upsampling improves spatial resolution and precision while maintaining high recall. Guidance improves class consistency and fidelity but reduces recall. When combined — applying guidance only to the low-resolution model before upsampling — the two approaches complement each other, yielding the best overall FIDs: 3.94 on 256×256 and 3.85 on 512×512.

This is like having a specialist who draws small but perfectly class-accurate sketches (guidance) and then a different specialist who enlarges sketches into detailed paintings (upsampling). Each handles a different aspect of quality.

Open in Lab
See how the two-stage pipeline works: low-resolution guided generation followed by high-resolution upsampling. Toggle guidance on and off to compare.
The demo wakes as you arrive…

The ripple effect: from classifier guidance to the diffusion revolution

This paper was a watershed moment for generative modeling. By proving that diffusion models could match and exceed GANs on their own turf — raw image quality — it shifted the entire field's center of gravity. Within months, the diffusion paradigm became dominant.

The most direct descendant is classifier-free guidance (Ho & Salimans, 2021). It observed that training a classifier on noisy images is cumbersome and limits the approach to labeled data. Their elegant alternative: train a single model that can operate both conditionally and unconditionally, then interpolate between the two during sampling. This preserved the diversity-fidelity trade-off without needing a separate classifier, and became the default guidance method in virtually all subsequent text-to-image models.

GLIDE (Nichol et al., 2021) applied both classifier guidance and classifier-free guidance to text-conditional image generation, showing that guided diffusion models could produce images from text prompts — laying the groundwork for DALL·E 2, Imagen, and Stable Diffusion.

  1. 2020

    DDPM: diffusion models produce high-quality images

    Ho et al. showed that diffusion models with a simple noise prediction objective can generate high-quality images, matching GANs on some benchmarks but not yet on ImageNet.

  2. 2021

    This paper: diffusion beats GANs

    Architecture improvements + classifier guidance push diffusion models past BigGAN-deep on ImageNet at all resolutions, with better distribution coverage.

  3. 2021

    Classifier-free guidance simplifies conditioning

    Ho and Salimans eliminate the need for a separate classifier by training a single model that supports both conditional and unconditional generation.

  4. 2021

    GLIDE: text-guided diffusion

    Nichol et al. combine guided diffusion with text conditioning, producing images from text prompts and showing diffusion can replace GANs for text-to-image too.

  5. 2022

    DALL·E 2, Imagen, Stable Diffusion

    The ADM architecture and guidance principles from this paper underpin the explosion of text-to-image diffusion models that reshaped creative AI.

Classifier-guided DDPM sampling (Algorithm 1 from the paper)python

Simplified to show the idea — not the real implementation.

# Classifier-guided diffusion sampling
# Given: diffusion model (mu_theta, Sigma_theta), classifier p_phi, scale s

def guided_sample(model, classifier, y, s, T):
    x = torch.randn_like(...)  # Start from pure noise x_T

    for t in reversed(range(1, T + 1)):
        # Step 1: Get unconditional denoising distribution
        mu, Sigma = model.predict(x, t)

        # Step 2: Compute classifier gradient
        x.requires_grad_(True)
        log_prob = classifier.log_prob(y, x, t)
        grad = torch.autograd.grad(log_prob, x)[0]

        # Step 3: Shift the mean by s * Sigma * gradient
        new_mu = mu + s * Sigma * grad

        # Step 4: Sample from the shifted Gaussian
        x = torch.normal(new_mu, Sigma.sqrt())

    return x  # Final denoised image

CitationDhariwal, Nichol. Diffusion Models Beat GANs on Image Synthesis. NeurIPS, 2021.

Terms in this paper