Generative Models2019advanced10 min read

Generative Modeling by Estimating Gradients of the Data Distribution

النمذجة التوليدية عبر تقدير تدرّجات توزيع البيانات

Song, Y. · Ermon, S. — NeurIPS

The problem

By 2019, generative models had split into two camps with opposing trade-offs. GANs produced sharp images but suffered from and unstable . Likelihood-based models (VAEs, normalizing flows) were principled but imposed strict architectural constraints — invertibility for flows, restrictive designs for VAEs — and often produced blurry samples. Neither camp offered a model that was simultaneously flexible in architecture, stable in training, and capable of high-quality generation without adversarial games.

The contribution

A fundamentally new generative framework: instead of learning a density or an adversarial game, learn the — the of the log-density — at multiple levels using . then proceeds via annealed , starting from pure noise and following the learned gradients downhill through decreasing noise scales until reaching the data . The architecture (a Noise Conditional Score Network, NCSN) is unconstrained: any that outputs a vector field will work. No adversarial training, no invertibility requirement, no . The method achieved state-of-the-art Inception and FID scores on CIFAR-10 among non-adversarial models.

The impact

This paper opened the "score-based" branch of generative modeling that converged with diffusion probabilistic models (DDPM) into today's diffusion-model family. The unified SDE framework (Song et al., 2021) showed NCSN and DDPM as two discretizations of the same continuous process. The ideas directly enabled flow matching, consistency models, and the diffusion engines behind DALL·E, Stable Diffusion, and Sora. Score-based thinking also spread beyond generation into inverse problems, molecular design, and audio synthesis.

Imagine a foggy mountain range at night. The valleys are where all the towns (your data) sit, but you can't see them. All you have is a compass that, at every point, tells you the steepest downhill direction.

Now release a thousand hikers from random spots on the mountain. Each one follows their compass downhill step by step. Eventually, they all cluster in the valleys — they've "generated" the towns without ever seeing a map.

But there's a catch: on flat plateaus high above the valleys, the compass barely moves — the signal is too faint. The paper's trick is to reshape the landscape at multiple blur levels: first create big, obvious valleys (heavy noise), let hikers find the general area, then sharpen the terrain (less noise) so they settle into the exact town centers.

The core idea: learning slopes instead of heights

Traditional generative models try to learn the probability density p(x)p(x) — essentially, how tall the landscape is at each point. But computing p(x)p(x) requires a constant (the partition function) that sums over all possible data points — an intractable integral in high dimensions.

This paper takes a different approach: instead of learning the height, learn the slope. The score function is defined as the gradient of the log-density:

s(x)=∇xlog⁡p(x)s(x) = \nabla_x \log p(x)

This is a vector field: at every point in data space, it points toward higher-density regions. Crucially, the gradient of the log-density does not depend on the normalization constant — when you take the gradient, the constant disappears because the derivative of a constant is zero. This means we can learn the score without ever computing the intractable partition function.

Open in Lab
The score function is a vector field pointing toward high-density regions. Toggle between the density view and the score (gradient) view to see how arrows guide samples toward data clusters.
The demo wakes as you arrive…
sθ(x)≈∇xlog⁡p(x)s_\theta(x) \approx \nabla_x \log p(x)
Score estimation — a neural network approximates the data gradient — A neural network sθs_\theta with parameters θ\theta is trained to output a vector at each point xx that approximates the true score. The network takes a data point as input and produces a vector of the same dimensionality, representing the direction of steepest ascent in log-density.

Training: score matching without knowing the score

There's an apparent paradox: to train the network to approximate the true score ∇xlog⁡p(x)\nabla_x \log p(x), we would need to know p(x)p(x) — which is exactly what we're trying to avoid. Score matching resolves this elegantly.

The naive objective would be to minimize the Fisher divergence — the expected squared distance between the model's output and the true score. Through integration by parts (Hyvärinen, 2005), this objective can be rewritten into a form that depends only on the model's and values at data points — not on p(x)p(x) itself.

The paper uses denoising score matching (Vincent, 2011), which is even simpler: perturb each data point xx with known Gaussian noise to get x~\tilde{x}. For a Gaussian , the true score of the noisy distribution has a closed form — it is simply the direction back to the clean data point, scaled by the noise variance:

∇x~log⁡qσ(x~∣x)=−x~−xσ2\nabla_{\tilde{x}} \log q_\sigma(\tilde{x} | x) = -\frac{\tilde{x} - x}{\sigma^2}
Denoising score — the ground truth comes free from Gaussian noise — Given a clean point xx and its noisy version x~=x+σϵ\tilde{x} = x + \sigma\epsilon where ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I), the score of the noisy conditional distribution simply points from x~\tilde{x} back toward xx, divided by σ2\sigma^2. This gives us a free training label at every noise level — no density estimation required.
Open in Lab
Drag the noisy point and see the denoising score arrow always points back to the clean data. This arrow is the training signal for the score network.
The demo wakes as you arrive…

The training loss is therefore refreshingly simple: add noise to a data point, ask the network to predict the score, and penalize the difference:

L(θ;σ)=12Epdata(x)Ex~∼N(x,σ2I)∥sθ(x~,σ)+x~−xσ2∥2\mathcal{L}(\theta; \sigma) = \frac{1}{2}\mathbb{E}_{p_{\text{data}}(x)} \mathbb{E}_{\tilde{x} \sim \mathcal{N}(x, \sigma^2 I)} \left\| s_\theta(\tilde{x}, \sigma) + \frac{\tilde{x} - x}{\sigma^2} \right\|^2
Denoising score matching loss — the network learns to point homeward — Sample a clean point xx, perturb it to x~\tilde{x}, and train the network sθs_\theta to match the known score. The ++ sign makes the network learn to output −(x~−x)/σ2-(\tilde{x}-x)/\sigma^2, i.e., the direction from the noisy point back to the clean data. This is the entire training procedure — no adversary, no decoder, no partition function.

The manifold problem: why one noise level is not enough

Real data occupies a thin manifold in high-dimensional space — images of faces lie on a curved surface far lower-dimensional than the pixel space. In vast regions between data clusters, the score is essentially undefined because there is virtually zero probability density there.

A score network trained with a single small noise level will be accurate near the data but unreliable in the empty regions. Langevin dynamics initialized from random noise would have to traverse these empty deserts with no reliable gradient signal — the random walker would wander aimlessly, never finding the data.

Think of it spatially: with a tiny noise level, each data point creates a narrow "cone of attraction." If your random hiker starts far away, none of the cones reach them — they feel no pull toward any valley.

The solution is elegant: use multiple noise levels. Large noise inflates each data point into a wide hill that covers the empty regions, giving the hiker a coarse direction. Small noise sharpens the landscape for precise positioning. The paper uses a geometric sequence of noise levels from large σ1\sigma_1 to small σL\sigma_L.

Open in Lab
Switch between noise levels to see how larger noise fills the gaps between data clusters. At high noise the landscape is smooth and easy to navigate; at low noise the peaks are sharp but valleys are unreachable from afar.
The demo wakes as you arrive…

The Noise Conditional Score Network (NCSN)

Instead of training a separate score network per noise level, the paper trains one shared network conditioned on the noise level σ\sigma. The Noise Conditional Score Network (NCSN) sθ(x,σ)s_\theta(x, \sigma) takes both a (noisy) data point and the noise scale as input.

The combined training objective sums the denoising score matching loss across all noise levels, weighted by λ(σi)=σi2\lambda(\sigma_i) = \sigma_i^2 to balance the different scales:

L(θ)=1L∑i=1Lλ(σi) E∥sθ(x~,σi)+x~−xσi2∥2\mathcal{L}(\theta) = \frac{1}{L}\sum_{i=1}^{L} \lambda(\sigma_i) \,\mathbb{E}\left\| s_\theta(\tilde{x}, \sigma_i) + \frac{\tilde{x} - x}{\sigma_i^2}\right\|^2
Multi-scale objective — one network, all noise levels — The weighting λ(σi)=σi2\lambda(\sigma_i) = \sigma_i^2 ensures the loss contribution from each noise level has comparable magnitude. Without it, large-σ\sigma terms would dominate because the score itself scales as 1/σ21/\sigma^2. The network shares all parameters across noise levels — σ\sigma is provided as conditioning input.

Sampling: annealed Langevin dynamics

Langevin dynamics is a sampling algorithm from physics: to draw samples from a distribution, start from random noise and iteratively take small steps in the direction of the score, plus a bit of random noise for exploration:

xt+1=xt+α2sθ(xt)+α zt,zt∼N(0,I)x_{t+1} = x_t + \frac{\alpha}{2} s_\theta(x_t) + \sqrt{\alpha}\, z_t, \quad z_t \sim \mathcal{N}(0, I)

Think of it as a ball rolling downhill on the log-density surface, but with random jitter so it doesn't get stuck in local bumps. Given enough steps and a small enough step size α\alpha, the ball's position converges to a sample from the true distribution.

But plain Langevin dynamics with one noise level fails for the manifold reasons discussed above. The paper's key sampling innovation is annealed Langevin dynamics: run Langevin dynamics in stages. Start with the score at the largest noise σ1\sigma_1 (wide smooth landscape), run TT steps to find the general region, then switch to σ2\sigma_2 (slightly sharper landscape), run TT more steps, and so on — gradually reducing the noise until you reach σL\sigma_L and settle precisely on the data manifold.

Open in Lab
Watch annealed Langevin dynamics in action. Particles start from random noise and are guided through decreasing noise levels until they settle into data clusters. Compare with single-noise Langevin which gets stuck.
The demo wakes as you arrive…
Annealed Langevin dynamics — the full sampling looppython

Simplified to show the idea — not the real implementation.

def annealed_langevin(score_net, sigmas, T=100, eps=0.00005):
    """Sample via annealed Langevin dynamics."""
    # Start from pure noise
    x = torch.randn(batch_size, *data_shape)

    for sigma in sigmas:              # From largest to smallest noise
        alpha = eps * (sigma / sigmas[-1]) ** 2   # Step size scales with noise
        for t in range(T):
            z = torch.randn_like(x)   # Fresh noise each step
            score = score_net(x, sigma)
            x = x + (alpha / 2) * score + torch.sqrt(alpha) * z

    return x  # Final samples on the data manifold

The complete pipeline: train → sample → generate

Open in Lab
The full NCSN pipeline: training (left) estimates the score at all noise levels; sampling (right) uses annealed Langevin dynamics to go from noise to data. Click each stage to see details.
The demo wakes as you arrive…

Choosing noise levels: the geometric schedule

The noise levels {σi}i=1L\{\sigma_i\}_{i=1}^{L} form a geometric sequence: each level is a fixed ratio of the previous one. The paper offers two design principles for selecting the endpoints.

The largest noise σ1\sigma_1 should be large enough that the noisy distribution qσ1(x)≈N(0,σ12I)q_{\sigma_1}(x) \approx \mathcal{N}(0, \sigma_1^2 I) is nearly indistinguishable from the prior — pure Gaussian noise. This ensures that Langevin dynamics at the first scale can mix well starting from anywhere in space.

The smallest noise σL\sigma_L should be small enough that the noisy distribution qσLq_{\sigma_L} is nearly indistinguishable from the true data — so that the final samples are clean. In practice, σL\sigma_L is set to a small fraction of the typical distance between nearest data points.

With L=10L = 10 levels and a geometric ratio between these extremes, the sequence creates a smooth bridge from "everything is one big blob" to "sharp, realistic images."

Open in Lab
Adjust the geometric noise schedule. Watch how σ₁ (largest noise) and σ_L (smallest noise) define the two endpoints of the annealing bridge.
The demo wakes as you arrive…

Results: beating likelihood-based models on image quality

On CIFAR-10, the NCSN achieved an of 8.87 and an FID of 25.32 — the best among all non-adversarial generative models at the time. For comparison, the best likelihood-based model (Glow) had an FID of 46.90, and the best (SNGAN) had an FID of 21.70.

On CelebA (128×128), the model generated faces with fine details like hair strands and facial features — quality competitive with GANs but without the mode collapse or training instability that plagued adversarial methods.

Perhaps more important than the numbers: the model generated diverse samples. Unlike GANs, which often converge to a subset of modes, the score-based approach covers the full data distribution because Langevin dynamics is theoretically guaranteed to converge to the correct distribution given enough steps.

Score-based vs GANs vs flows vs VAEs

Open in Lab
Compare the four generative paradigms: what each learns, its training requirements, and its trade-offs. Click each paradigm to explore.
The demo wakes as you arrive…

Legacy: the road from scores to modern diffusion

  1. 2005

    Score matching (Hyvärinen)

    Introduced the mathematical framework for estimating score functions without knowing the normalization constant. The theoretical foundation that this paper builds upon.

  2. 2011

    Denoising score matching (Vincent)

    Connected score matching with denoising — showing that adding noise and training to reverse it is equivalent to score estimation. Made score matching practical for deep networks.

  3. 2019

    This paper — NCSN (Song & Ermon)

    Combined multi-scale noise with score matching and annealed Langevin dynamics to create the first practical score-based generative model.

  4. 2020

    DDPM (Ho et al.)

    Denoising Diffusion Probabilistic Models arrived at a similar framework from a different angle — viewing generation as iterative denoising through a learned reverse process.

  5. 2021

    Score-SDE unification (Song et al.)

    Showed that NCSN and DDPM are two discretizations of the same continuous stochastic differential equation. Unified score-based and diffusion models into a single framework.

  6. 2022

    Flow matching (Lipman et al.)

    Extended the score-based idea by learning velocity fields instead of score functions, enabling straighter transport paths and faster sampling.

  7. 2023

    Consistency models (Song et al.)

    Distilled the multi-step score-based sampling into single-step generation while maintaining quality — bridging the speed gap between diffusion and GANs.

What Song and Ermon demonstrated was that we don't need adversarial training, invertible architectures, or explicit density computation to build powerful generative models. We just need to learn which direction is "downhill" — toward the data — and follow it. This simple, physics-inspired insight turned out to be the key that unlocked the diffusion revolution.

CitationSong, Y. and Ermon, S.. Generative Modeling by Estimating Gradients of the Data Distribution. NeurIPS, 2019.

Terms in this paper