Generative Models2021advanced12 min read
Diffusion Models Beat GANs on Image Synthesis
نماذج الانتشار تتفوّق على شبكات GAN في توليد الصور
Dhariwal, P. · Nichol, A. — NeurIPS
The problem
By 2021, GANs dominated image generation in terms of sample quality metrics like and . However, GANs suffered from mode collapse, instability, and limited diversity. Diffusion models showed promise — they covered the data distribution better and trained more stably — but their sample quality still lagged behind GANs on challenging datasets like ImageNet. Two key gaps remained: architectures had not been optimized as extensively as architectures, and there was no mechanism for diffusion models to trade off diversity for fidelity the way GANs do with truncation.
The contribution
Two complementary breakthroughs. First, a systematic that improved the architecture for diffusion models: deeper at multiple resolutions, more attention heads, BigGAN-style residual blocks, and . This alone achieved state-of-the-art on unconditional generation (LSUN). Second, — using gradients from a pretrained classifier to steer the process toward a target class. A single hyperparameter (the scale) smoothly trades diversity for fidelity. Together, these achieved FID 2.97 on ImageNet 128×128 and 4.59 on 256×256, surpassing BigGAN-deep while maintaining better distribution coverage.
The impact
This paper marked the moment diffusion models overtook GANs as the leading paradigm for image generation. Classifier guidance directly inspired (Ho & Salimans, 2021), which eliminated the need for a separate classifier and became the standard conditioning mechanism in DALL·E 2, Imagen, Stable Diffusion, and GLIDE. The improved U-Net architecture (ADM) became the backbone of nearly all subsequent diffusion image generators. The paper effectively ended the GAN era of dominance and launched the diffusion era that continues today.
Imagine two art studios competing to produce photorealistic paintings. Studio GAN uses a brilliant but temperamental artist paired with a harsh critic — together they produce stunning images, but the artist keeps painting the same few subjects and throws tantrums if the critic's feedback changes slightly.
Studio Diffusion takes a different approach: start with a canvas covered in static, and have a patient restorer clean it one careful brushstroke at a time. The restorer is stable, never throws tantrums, and naturally produces a wide variety of subjects. The problem? The restorations were slightly blurry compared to Studio GAN's paintings.
This paper gave Studio Diffusion two upgrades: a better set of restoration tools (architecture improvements) and a museum guide who whispers "make it look more like a corgi" during the cleaning process (classifier guidance). Suddenly, Studio Diffusion produced sharper paintings than Studio GAN — and with more variety.
Refresher: how diffusion models generate images
A diffusion model generates images by reversing a gradual noising process. During training, real images are progressively corrupted by adding Gaussian noise across timesteps, until the image becomes pure static. The model learns to undo this corruption one step at a time — predicting and removing the noise component at each timestep.
Think of it like a highway with exits. At exit 0, you have a clean image. At exit , you have pure noise. The drives the image from exit 0 to exit , adding a little noise at each exit. The model learns the reverse trip: starting from pure noise at exit and gradually denoising back to exit 0.
The noise predictor takes a noisy image and the current timestep as input, and predicts the noise that was added. Training uses a simple loss between the predicted noise and the actual noise.
Breakthrough 1: a better U-Net through systematic ablations
The first hypothesis: the gap between diffusion models and GANs comes partly from under-optimized architectures. GAN architectures like StyleGAN2 and BigGAN had been refined through years of ablation studies. Diffusion models were still using the original U-Net from with minimal changes.
Dhariwal and Nichol conducted a systematic ablation study on ImageNet 128×128, testing one change at a time. They discovered that several modifications compound to give a substantial boost in FID, and most of these are inspired by best practices from GAN and Transformer literature.
The key modifications, each tested independently and then combined, are:
More depth vs. width. Increasing depth (more residual blocks per resolution) while holding model size constant improves FID, but slows training. Wider models reach the same quality faster, so the authors chose width over depth.
Multi-resolution attention. Instead of applying attention only at 16×16 resolution (as in DDPM), the improved model applies attention at 32×32, 16×16, and 8×8 resolutions. This helps the model capture both fine details and global structure.
More attention heads. Increasing from 1 head to multiple heads with 64 channels per head significantly improves FID. This matches modern Transformer practice, where many small heads outperform few large ones.
BigGAN residual blocks. Using BigGAN-style and downsampling residual blocks instead of simple convolutions improves information flow through the network.
Adaptive (AdaGN). Instead of adding timestep and class embeddings via simple addition, AdaGN modulates the normalization parameters: , where and come from a linear projection of the timestep and class embedding. This is analogous to adaptive instance normalization from StyleGAN.
Breakthrough 2: classifier guidance
The second insight addresses a fundamental advantage that GANs had: they could trade off diversity for fidelity using the . A smaller truncation range produces sharper but less diverse images. Diffusion models had no equivalent knob.
The solution: classifier guidance. The idea is beautifully simple. Train a separate classifier on noisy images at every noise level. During , use the classifier's gradients to nudge each denoising step toward a target class . The classifier acts like a museum guide whispering directions to the restorer: "this should look more like a corgi."
Mathematically, the conditional becomes:
Under reasonable assumptions (the classifier's log-probability has low curvature compared to the of the Gaussian transition), this product simplifies elegantly. The conditional transition remains a Gaussian — like the unconditional one — but with its mean shifted by , where is the classifier gradient and is the predicted .
The sampling algorithm is simple: at each timestep, compute the unconditional mean and variance , then sample from , where is a gradient scale factor.
The gradient scale: a single knob for quality vs. diversity
A key practical discovery: using the classifier gradients at their natural scale () is not enough. While the classifier assigns around 50% probability to the target class, the resulting images do not visually match the class. The gradients need to be amplified.
Scaling the gradients by has a clean theoretical interpretation. It is equivalent to sampling from a sharpened classifier distribution . When , this distribution becomes more peaked — it concentrates probability on the most confident class predictions. Think of it as turning up the "conviction" of the museum guide.
The authors found that the scale can be increased by an order of magnitude (up to ) without producing adversarial examples. This smoothly trades off (diversity) for precision (fidelity), with FID reaching its optimum at an intermediate scale.
Results: surpassing GANs across the board
The improved architecture alone (ADM) achieved state-of-the-art FID on LSUN bedrooms (1.90), horses (2.57), and cats (5.57) — all unconditional, beating StyleGAN2. On ImageNet 64×64 the improved model reached FID 2.07.
With classifier guidance (ADM-G), diffusion models surpassed BigGAN-deep on ImageNet at every resolution: FID 2.97 at 128×128, 4.59 at 256×256, and 7.72 at 512×512. Crucially, the diffusion models maintained higher recall (distribution coverage) than GANs at comparable or better precision.
Even with only 25 sampling steps using , the guided model matched BigGAN-deep's FID while still covering more of the distribution.
Finally, combining classifier guidance with upsampling diffusion models gave the best overall results: FID 3.94 on 256×256 and 3.85 on 512×512. The authors showed that guidance and upsampling improve quality along different axes — guidance boosts fidelity while upsampling improves spatial resolution — so they complement each other.
Classifier guidance for DDIM: the score-based view
The Gaussian perturbation derivation above only works for stochastic sampling. For deterministic samplers like DDIM, the authors use a score-based trick. The noise predictor can be reinterpreted as a scaled :
Adding the classifier gradient gives the joint score:
This translates into a modified noise prediction that can be plugged directly into the DDIM sampler. The classifier gradient literally "subtracts class-relevant noise" from the prediction, so the denoiser favors the target class.
Combining guidance with upsampling: complementary strengths
The authors also explored two-stage upsampling stacks, where a low-resolution diffusion model generates a small image and then a separate upsampling model enlarges it. They found that guidance and upsampling improve quality along different axes:
Upsampling improves spatial resolution and precision while maintaining high recall. Guidance improves class consistency and fidelity but reduces recall. When combined — applying guidance only to the low-resolution model before upsampling — the two approaches complement each other, yielding the best overall FIDs: 3.94 on 256×256 and 3.85 on 512×512.
This is like having a specialist who draws small but perfectly class-accurate sketches (guidance) and then a different specialist who enlarges sketches into detailed paintings (upsampling). Each handles a different aspect of quality.
The ripple effect: from classifier guidance to the diffusion revolution
This paper was a watershed moment for generative modeling. By proving that diffusion models could match and exceed GANs on their own turf — raw image quality — it shifted the entire field's center of gravity. Within months, the diffusion paradigm became dominant.
The most direct descendant is classifier-free guidance (Ho & Salimans, 2021). It observed that training a classifier on noisy images is cumbersome and limits the approach to labeled data. Their elegant alternative: train a single model that can operate both conditionally and unconditionally, then interpolate between the two during sampling. This preserved the diversity-fidelity trade-off without needing a separate classifier, and became the default guidance method in virtually all subsequent text-to-image models.
GLIDE (Nichol et al., 2021) applied both classifier guidance and classifier-free guidance to text-conditional image generation, showing that guided diffusion models could produce images from text prompts — laying the groundwork for DALL·E 2, Imagen, and Stable Diffusion.
2020
DDPM: diffusion models produce high-quality images
Ho et al. showed that diffusion models with a simple noise prediction objective can generate high-quality images, matching GANs on some benchmarks but not yet on ImageNet.
2021
This paper: diffusion beats GANs
Architecture improvements + classifier guidance push diffusion models past BigGAN-deep on ImageNet at all resolutions, with better distribution coverage.
2021
Classifier-free guidance simplifies conditioning
Ho and Salimans eliminate the need for a separate classifier by training a single model that supports both conditional and unconditional generation.
2021
GLIDE: text-guided diffusion
Nichol et al. combine guided diffusion with text conditioning, producing images from text prompts and showing diffusion can replace GANs for text-to-image too.
2022
DALL·E 2, Imagen, Stable Diffusion
The ADM architecture and guidance principles from this paper underpin the explosion of text-to-image diffusion models that reshaped creative AI.
Simplified to show the idea — not the real implementation.
# Classifier-guided diffusion sampling
# Given: diffusion model (mu_theta, Sigma_theta), classifier p_phi, scale s
def guided_sample(model, classifier, y, s, T):
x = torch.randn_like(...) # Start from pure noise x_T
for t in reversed(range(1, T + 1)):
# Step 1: Get unconditional denoising distribution
mu, Sigma = model.predict(x, t)
# Step 2: Compute classifier gradient
x.requires_grad_(True)
log_prob = classifier.log_prob(y, x, t)
grad = torch.autograd.grad(log_prob, x)[0]
# Step 3: Shift the mean by s * Sigma * gradient
new_mu = mu + s * Sigma * grad
# Step 4: Sample from the shifted Gaussian
x = torch.normal(new_mu, Sigma.sqrt())
return x # Final denoised image
CitationDhariwal, Nichol. Diffusion Models Beat GANs on Image Synthesis. NeurIPS, 2021.
Terms in this paper
- Diffusion Modelنموذج الانتشار
- Classifier Guidanceالتوجيه بالمصنِّف
- FIDمسافة فريشيه للبداية
- U-Netشبكة U-Net
- Ablation Studyدراسة الاستئصال
- Generative Adversarial Network (GAN)الشبكات التوليدية التنافسية
- Denoisingإزالة الضوضاء
- Noise Scheduleجدول الضوضاء
- Samplingاختيار العينات الاحتمالية
- Classifier-Free Guidanceالتوجيه بلا مُصنِّف
- Precisionمدى المحركية (الدقة المحددة لفئة)
- Recallالقدرة على الاستدعاء الشامل
- Inception Scoreمعيار إنسبشن
- Adaptive Group Normalizationالتسوية التكيُّفية للمجموعات
- Residual Blockالكتلة المتبقّية