Generative Models2019advanced10 min read
Generative Modeling by Estimating Gradients of the Data Distribution
النمذجة التوليدية عبر تقدير تدرّجات توزيع البيانات
Song, Y. · Ermon, S. — NeurIPS
The problem
By 2019, generative models had split into two camps with opposing trade-offs. GANs produced sharp images but suffered from and unstable . Likelihood-based models (VAEs, normalizing flows) were principled but imposed strict architectural constraints — invertibility for flows, restrictive designs for VAEs — and often produced blurry samples. Neither camp offered a model that was simultaneously flexible in architecture, stable in training, and capable of high-quality generation without adversarial games.
The contribution
A fundamentally new generative framework: instead of learning a density or an adversarial game, learn the — the of the log-density — at multiple levels using . then proceeds via annealed , starting from pure noise and following the learned gradients downhill through decreasing noise scales until reaching the data . The architecture (a Noise Conditional Score Network, NCSN) is unconstrained: any that outputs a vector field will work. No adversarial training, no invertibility requirement, no . The method achieved state-of-the-art Inception and FID scores on CIFAR-10 among non-adversarial models.
The impact
This paper opened the "score-based" branch of generative modeling that converged with diffusion probabilistic models (DDPM) into today's diffusion-model family. The unified SDE framework (Song et al., 2021) showed NCSN and DDPM as two discretizations of the same continuous process. The ideas directly enabled flow matching, consistency models, and the diffusion engines behind DALL·E, Stable Diffusion, and Sora. Score-based thinking also spread beyond generation into inverse problems, molecular design, and audio synthesis.
Imagine a foggy mountain range at night. The valleys are where all the towns (your data) sit, but you can't see them. All you have is a compass that, at every point, tells you the steepest downhill direction.
Now release a thousand hikers from random spots on the mountain. Each one follows their compass downhill step by step. Eventually, they all cluster in the valleys — they've "generated" the towns without ever seeing a map.
But there's a catch: on flat plateaus high above the valleys, the compass barely moves — the signal is too faint. The paper's trick is to reshape the landscape at multiple blur levels: first create big, obvious valleys (heavy noise), let hikers find the general area, then sharpen the terrain (less noise) so they settle into the exact town centers.
The core idea: learning slopes instead of heights
Traditional generative models try to learn the probability density — essentially, how tall the landscape is at each point. But computing requires a constant (the partition function) that sums over all possible data points — an intractable integral in high dimensions.
This paper takes a different approach: instead of learning the height, learn the slope. The score function is defined as the gradient of the log-density:
This is a vector field: at every point in data space, it points toward higher-density regions. Crucially, the gradient of the log-density does not depend on the normalization constant — when you take the gradient, the constant disappears because the derivative of a constant is zero. This means we can learn the score without ever computing the intractable partition function.
Training: score matching without knowing the score
There's an apparent paradox: to train the network to approximate the true score , we would need to know — which is exactly what we're trying to avoid. Score matching resolves this elegantly.
The naive objective would be to minimize the Fisher divergence — the expected squared distance between the model's output and the true score. Through integration by parts (Hyvärinen, 2005), this objective can be rewritten into a form that depends only on the model's and values at data points — not on itself.
The paper uses denoising score matching (Vincent, 2011), which is even simpler: perturb each data point with known Gaussian noise to get . For a Gaussian , the true score of the noisy distribution has a closed form — it is simply the direction back to the clean data point, scaled by the noise variance:
The training loss is therefore refreshingly simple: add noise to a data point, ask the network to predict the score, and penalize the difference:
The manifold problem: why one noise level is not enough
Real data occupies a thin manifold in high-dimensional space — images of faces lie on a curved surface far lower-dimensional than the pixel space. In vast regions between data clusters, the score is essentially undefined because there is virtually zero probability density there.
A score network trained with a single small noise level will be accurate near the data but unreliable in the empty regions. Langevin dynamics initialized from random noise would have to traverse these empty deserts with no reliable gradient signal — the random walker would wander aimlessly, never finding the data.
Think of it spatially: with a tiny noise level, each data point creates a narrow "cone of attraction." If your random hiker starts far away, none of the cones reach them — they feel no pull toward any valley.
The solution is elegant: use multiple noise levels. Large noise inflates each data point into a wide hill that covers the empty regions, giving the hiker a coarse direction. Small noise sharpens the landscape for precise positioning. The paper uses a geometric sequence of noise levels from large to small .
The Noise Conditional Score Network (NCSN)
Instead of training a separate score network per noise level, the paper trains one shared network conditioned on the noise level . The Noise Conditional Score Network (NCSN) takes both a (noisy) data point and the noise scale as input.
The combined training objective sums the denoising score matching loss across all noise levels, weighted by to balance the different scales:
Sampling: annealed Langevin dynamics
Langevin dynamics is a sampling algorithm from physics: to draw samples from a distribution, start from random noise and iteratively take small steps in the direction of the score, plus a bit of random noise for exploration:
Think of it as a ball rolling downhill on the log-density surface, but with random jitter so it doesn't get stuck in local bumps. Given enough steps and a small enough step size , the ball's position converges to a sample from the true distribution.
But plain Langevin dynamics with one noise level fails for the manifold reasons discussed above. The paper's key sampling innovation is annealed Langevin dynamics: run Langevin dynamics in stages. Start with the score at the largest noise (wide smooth landscape), run steps to find the general region, then switch to (slightly sharper landscape), run more steps, and so on — gradually reducing the noise until you reach and settle precisely on the data manifold.
Simplified to show the idea — not the real implementation.
def annealed_langevin(score_net, sigmas, T=100, eps=0.00005):
"""Sample via annealed Langevin dynamics."""
# Start from pure noise
x = torch.randn(batch_size, *data_shape)
for sigma in sigmas: # From largest to smallest noise
alpha = eps * (sigma / sigmas[-1]) ** 2 # Step size scales with noise
for t in range(T):
z = torch.randn_like(x) # Fresh noise each step
score = score_net(x, sigma)
x = x + (alpha / 2) * score + torch.sqrt(alpha) * z
return x # Final samples on the data manifoldThe complete pipeline: train → sample → generate
Choosing noise levels: the geometric schedule
The noise levels form a geometric sequence: each level is a fixed ratio of the previous one. The paper offers two design principles for selecting the endpoints.
The largest noise should be large enough that the noisy distribution is nearly indistinguishable from the prior — pure Gaussian noise. This ensures that Langevin dynamics at the first scale can mix well starting from anywhere in space.
The smallest noise should be small enough that the noisy distribution is nearly indistinguishable from the true data — so that the final samples are clean. In practice, is set to a small fraction of the typical distance between nearest data points.
With levels and a geometric ratio between these extremes, the sequence creates a smooth bridge from "everything is one big blob" to "sharp, realistic images."
Results: beating likelihood-based models on image quality
On CIFAR-10, the NCSN achieved an of 8.87 and an FID of 25.32 — the best among all non-adversarial generative models at the time. For comparison, the best likelihood-based model (Glow) had an FID of 46.90, and the best (SNGAN) had an FID of 21.70.
On CelebA (128×128), the model generated faces with fine details like hair strands and facial features — quality competitive with GANs but without the mode collapse or training instability that plagued adversarial methods.
Perhaps more important than the numbers: the model generated diverse samples. Unlike GANs, which often converge to a subset of modes, the score-based approach covers the full data distribution because Langevin dynamics is theoretically guaranteed to converge to the correct distribution given enough steps.
Score-based vs GANs vs flows vs VAEs
Legacy: the road from scores to modern diffusion
2005
Score matching (Hyvärinen)
Introduced the mathematical framework for estimating score functions without knowing the normalization constant. The theoretical foundation that this paper builds upon.
2011
Denoising score matching (Vincent)
Connected score matching with denoising — showing that adding noise and training to reverse it is equivalent to score estimation. Made score matching practical for deep networks.
2019
This paper — NCSN (Song & Ermon)
Combined multi-scale noise with score matching and annealed Langevin dynamics to create the first practical score-based generative model.
2020
DDPM (Ho et al.)
Denoising Diffusion Probabilistic Models arrived at a similar framework from a different angle — viewing generation as iterative denoising through a learned reverse process.
2021
Score-SDE unification (Song et al.)
Showed that NCSN and DDPM are two discretizations of the same continuous stochastic differential equation. Unified score-based and diffusion models into a single framework.
2022
Flow matching (Lipman et al.)
Extended the score-based idea by learning velocity fields instead of score functions, enabling straighter transport paths and faster sampling.
2023
Consistency models (Song et al.)
Distilled the multi-step score-based sampling into single-step generation while maintaining quality — bridging the speed gap between diffusion and GANs.
What Song and Ermon demonstrated was that we don't need adversarial training, invertible architectures, or explicit density computation to build powerful generative models. We just need to learn which direction is "downhill" — toward the data — and follow it. This simple, physics-inspired insight turned out to be the key that unlocked the diffusion revolution.
CitationSong, Y. and Ermon, S.. Generative Modeling by Estimating Gradients of the Data Distribution. NeurIPS, 2019.
Terms in this paper
- Score Functionدالة الرصيد
- Score Matchingمطابقة النتيجة
- Langevin Dynamicsديناميكيات لانجفان
- Noise Scheduleجدول الضوضاء
- Denoisingإزالة الضوضاء
- Generative Modelالنموذج التوليدي
- Density Estimationتقدير الكثافة
- Gradientالتدرج التفاضلي
- Manifoldالمتشعب الهندسي
- Diffusion Modelنموذج الانتشار
- Samplingاختيار العينات الاحتمالية
- Energy Functionدالة الطاقة
- Normalizationالمعايرة القياسية للبيانات
- Annealingالتدريج التصاعدي
- Perturbationاضطراب