Generative Models2022intermediate12 min read
Classifier-Free Diffusion Guidance
التوجيه بدون مُصنِّف في نماذج الانتشار
Ho, J. · Salimans, T. — NeurIPS Workshop (2021) / arXiv (2022)
The problem
By 2021, diffusion models were producing state-of-the-art images, but controlling what they generated — making them produce a specific class with high fidelity — required a separate classifier network. This "" approach of Dhariwal & Nichol was effective but complicated: the classifier had to be trained on noisy images, added a second model to maintain, and its gradients resembled adversarial attacks on the classifier itself, raising questions about whether the quality gains were genuine or just gaming classifier-based metrics.
The contribution
A strikingly simple alternative: instead of a separate classifier, jointly train one to handle both conditional generation (with a class label) and generation (without). During , extrapolate away from the unconditional prediction and toward the conditional one using a guidance weight w. This single-line code change during training — randomly dropping the class label — and a simple mixing formula during sampling achieves the same quality-diversity tradeoff as classifier guidance, but with a pure and zero classifier gradients.
The impact
became the de facto standard for conditional generation in virtually every major diffusion system that followed. Imagen, Stable Diffusion (Latent Diffusion), DALL·E 2, and nearly all text-to-image models use it as their primary mechanism for controlling sample quality. The idea's simplicity — one extra line of code — made it trivially adoptable, and its generality extended beyond images to audio, video, and 3D generation. It is arguably the single most impactful sampling technique in modern generative AI.
Imagine you're a photographer trying to take the perfect picture of a golden retriever.
The old way (classifier guidance): you bring along an art critic who keeps whispering "more dog-like, more dog-like" as you adjust the camera. The critic helps, but you have to pay and transport them, they only judge through their own lens, and sometimes their advice just makes photos look good to critics rather than actually better.
The new way (classifier-free guidance): you take two shots yourself — one where you try to capture "golden retriever" and one where you just point the camera randomly. Then you look at what's different between them and push harder in that direction. No critic needed. The difference between "what I asked for" and "nothing in particular" is the guidance signal.
The problem: high quality or high diversity — pick one?
Generative models face a fundamental tension. When asked to generate images of a class — say "golden retriever" — they can produce diverse samples that cover the full range of what golden retrievers look like (different poses, backgrounds, lighting). But diversity comes at a cost: some samples will be off-target, blurry, or ambiguous. Alternatively, they can concentrate on producing sharp, unmistakable golden retrievers — but then every output starts looking the same.
In GANs, this tradeoff is controlled by truncation: shrinking the range of random noise inputs. Small noise range means less variety but crisper images. In flow models, serves the same role: lower temperature, less diversity, sharper outputs.
But diffusion models had no equivalent knob. Naive attempts — scaling the model's score vectors or reducing sampling noise — produced blurry, washed-out images rather than the crisp, high-fidelity samples that truncation gives GANs.
Classifier guidance: the predecessor
Dhariwal & Nichol (2021) solved the truncation problem for diffusion models with a clever trick: during the process, mix the diffusion model's score estimate with the of a separately trained classifier. The classifier says "this noisy image looks 60% like a golden retriever — here's which direction to push the pixels to make it more retriever-like." By varying the classifier gradient strength, they could smoothly trade off sample diversity (measured by FID) against sample fidelity (measured by ).
This worked beautifully but introduced three problems. First, the classifier must be trained on noisy images at every noise level — you cannot simply plug in an off-the-shelf ImageNet classifier. Second, the training pipeline doubles in complexity because you now maintain two models. Third, and most subtle: the sampling process literally takes gradient steps to maximize a classifier's confidence, which is structurally identical to an . Are the images actually better, or do they just fool the classifier?
The core idea: guidance from the model itself
The key insight is elegant: you don't need an external classifier to know "what makes this image more dog-like." The model already knows, because it has learned both what dogs look like (conditional score) and what generic images look like (unconditional score). The difference between these two scores is the classifier signal — it's an implicit classifier derived from Bayes' rule.
Here is the idea in three steps. Step 1 (Training): train a single diffusion model that sometimes sees the class label and sometimes does not. In practice, randomly replace the class label with a null token ∅ with probability p_uncond (a small value like 0.1 or 0.2). The model learns to denoise both conditionally and unconditionally.
Step 2 (Conditional score): at sampling time, run the model with the desired class label c to get the conditional noise prediction ε_θ(z_λ, c).
Step 3 (Guidance): also run the model without a label to get the unconditional prediction ε_θ(z_λ). Then combine them: push away from the unconditional prediction and toward the conditional one, scaled by a guidance weight w.
Training: one model, two modes
The training change is minimal. The model is a standard conditional diffusion model ε_θ(z_λ, c) — for example a that receives a class . The only modification is that during training, the class label c is randomly replaced with a null token ∅ with probability p_uncond. This means the network learns to denoise both with and without class information, using the same weights.
No extra model, no extra loss, no extra hyperparameter tuning. Just one line of code: with some probability, set c = ∅. The authors found that p_uncond values of 0.1 to 0.2 work well — only a small fraction of training needs to be unconditional. Even with p_uncond = 0.5, the model still achieves competitive performance, though the quality-diversity frontier shifts.
Simplified to show the idea — not the real implementation.
# Training loop for classifier-free guidance # The ONLY change: randomly drop the class label
for x, c in dataloader:
# With probability p_uncond, replace class with null
if random() < p_uncond:
c = NULL_TOKEN # ← the one-line change
t = sample_timestep() # Sample noise level
eps = randn_like(x) # Sample noise
z_t = alpha_t * x + sigma_t * eps
# Standard denoising loss — same as any diffusion model
loss = mse(model(z_t, c, t), eps)
loss.backward()
optimizer.step()Sampling: two passes, one formula
At sampling time, each step requires two forward passes of the model: one conditional pass with the desired class label c, and one unconditional pass with the null token ∅. The guided noise prediction is then computed using the formula from above.
Think of it as a compass. The conditional prediction points toward "images of class c." The unconditional prediction points toward "generic images." The guidance formula says: start at the conditional direction, then continue past it by an amount proportional to how far it differs from the unconditional direction. You're overshooting the conditional target, amplifying what makes it special.
This doubles the computational cost per step compared to standard sampling (two forward passes instead of one), but it can be partially offset by using fewer total steps. A potential future optimization mentioned in the paper is injecting late in the network, so parts of the computation can be shared.
Simplified to show the idea — not the real implementation.
# Sampling with classifier-free guidance
z = randn(shape) # Start from pure noise
for t in reversed(timesteps):
# Two forward passes per step
eps_cond = model(z, c, t) # Conditional prediction
eps_uncond = model(z, NULL, t) # Unconditional prediction
# The guidance formula
eps_guided = (1 + w) * eps_cond - w * eps_uncond
# Standard DDPM/DDIM update using eps_guided
z = denoise_step(z, eps_guided, t)
return zWhy it works: the implicit classifier
The mathematical justification connects classifier-free guidance to classifier guidance through Bayes' rule. An implicit classifier can be defined from any conditional and unconditional generative model as the ratio of their likelihoods.
In score space (working with gradients of log-densities), this becomes a difference between the conditional and unconditional scores. Classifier-free guidance with weight w is equivalent to classifier guidance using this implicit classifier with the same weight w — at least when the model estimates are exact.
In practice the estimates are not exact, and the score estimates come from unconstrained neural networks that don't form conservative vector fields. So the guided score ε̃_θ cannot literally be the gradient of any classifier's . But the empirical results show it works just as well — the guidance signal doesn't need to be a true classifier gradient.
Experiments: proof that it works
The authors evaluate on class-conditional ImageNet generation at 64×64 and 128×128 resolution, the standard benchmark for studying quality-diversity tradeoffs since BigGAN. They sweep the guidance weight w from 0 to 4 and measure FID (lower is better for diversity) and Inception Score (higher is better for fidelity).
Key findings are as follows. At w=0 (no guidance), the model produces diverse but imperfect samples. A small amount of guidance (w=0.1–0.3) dramatically improves FID, often beating classifier-guided models. Strong guidance (w≥4) maximizes IS at the expense of FID, exactly paralleling the truncation curve of GANs. On 128×128 ImageNet, classifier-free guidance at w=0.3 outperforms ADM-G (the classifier-guided baseline) on FID; at w=4.0 it outperforms BigGAN-deep on both FID and IS simultaneously.
Regarding the unconditional training probability p_uncond, values of 0.1 and 0.2 perform roughly equally; p_uncond=0.5 is consistently worse, confirming that only a small fraction of training capacity needs to be dedicated to unconditional generation. For sampling steps T, the model benefits from more steps (T=256 is a good balance), but each step costs two forward passes, so the fair speed comparison with classifier guidance is at half the step count.
Visual intuition: guidance on Gaussians
The paper provides a beautiful toy example. Imagine three classes, each represented by an isotropic Gaussian in 2D. Without guidance, the density is a mixture of three overlapping blobs. As guidance strength increases, each class's density sharpens, pulls away from the others, and concentrates into a tighter region. The shapes become distinctly non-Gaussian — each class "claims territory" by pushing probability mass away from neighboring classes.
This is a miniature version of what happens in high-dimensional image space: guidance makes each class more distinctive and concentrated, at the cost of no longer covering the full range of possible samples.
Why classifier-free won: simplicity and purity
The paper identifies several advantages of classifier-free guidance over classifier guidance.
Simplicity: one line of code during training (random label dropout), one formula during sampling. No second model to train, validate, or deploy.
Purity: the sampling process uses only generative model scores. No classifier gradients are involved, so the results cannot be dismissed as adversarial artifacts. This proves that pure generative models are capable of high-fidelity conditional generation.
Flexibility: the approach generalizes to any conditioning signal, not just class labels. This is why it became the backbone of text-to-image models like Imagen and Stable Diffusion, where the "class" is replaced by a text embedding.
The main disadvantage is sampling speed: each step requires two full forward passes instead of one (the generative model is typically much larger than a classifier). However, this cost proved acceptable in practice, and techniques like late conditioning injection or distillation can mitigate it.
Timeline: from idea to industry standard
2020
DDPM — modern diffusion foundations
Ho, Jain & Abbeel formalize denoising diffusion probabilistic models, establishing the training and sampling framework on which guidance builds.
2021
Diffusion beats GANs + Classifier guidance
Dhariwal & Nichol show diffusion models can outperform GANs on ImageNet, introducing classifier guidance as the mechanism for quality-diversity tradeoff.
2021
Classifier-free guidance proposed
Ho & Salimans present the classifier-free alternative at the NeurIPS 2021 Workshop. The method eliminates the need for a separate classifier with a single-line training change.
2022
Imagen and DALL·E 2 adopt CFG
Google's Imagen and OpenAI's DALL·E 2 both use classifier-free guidance as their core conditioning mechanism for text-to-image generation, validating the approach at scale.
2022
Stable Diffusion brings CFG to the masses
Latent Diffusion / Stable Diffusion makes classifier-free guidance accessible to millions of users. The "CFG scale" slider becomes a familiar control in every AI art interface.
Few papers in machine learning have achieved so much with so little. Classifier-free guidance is not a new architecture, not a new , not even a new model. It is a training trick (random label dropout) and a sampling formula (linear extrapolation of scores). Yet it became the default conditioning mechanism in the entire diffusion model ecosystem. The lesson is powerful: sometimes the most impactful ideas are not the most complex — they are the ones that remove complexity while preserving the result.
CitationHo, Salimans. Classifier-Free Diffusion Guidance. NeurIPS Workshop 2021 / arXiv 2022, 2022.
Terms in this paper
- Classifier-Free Guidanceالتوجيه بلا مُصنِّف
- Diffusion Modelنموذج الانتشار
- Score Functionدالة الرصيد
- Conditioningالتوجيه
- FID Scoreمقياس FID
- Unconditionalغير مشروط
- Noise Scheduleجدول الضوضاء
- Denoisingإزالة الضوضاء
- Generative Modelالنموذج التوليدي
- Samplingاختيار العينات الاحتمالية
- Mode Collapseانهيار الأنماط
- Temperatureالحرارة
- Forward Diffusionالانتشار الأمامي (إضافة الضجيج)
- Reverse Diffusionالانتشار خارجي
- DDPMنموذج الانتشار الاحتمالي لإزالة التشويش