Generative Models2020advanced14 min read

Analyzing and Improving the Image Quality of StyleGAN

تحليل جودة الصور في StyleGAN وتحسينها

Karras, T. · Laine, S. · Aittala, M. · Hellsten, J. · Lehtinen, J. · Aila, T. — CVPR

The problem

StyleGAN achieved state-of-the-art results in unconditional image generation, but closer inspection revealed systematic quality issues. Generated images contained characteristic water-droplet artifacts caused by the AdaIN normalization. Progressive growing — the technique that enabled high-resolution — caused phase artifacts where facial features like eyes and teeth appeared glued to fixed canvas positions instead of moving naturally. Furthermore, the mapping from latent codes to images was poorly conditioned: small changes in could cause unpredictable jumps in the output, making the difficult to invert and control.

The contribution

Three architectural and training improvements that together eliminate StyleGAN's quality issues. First, replaces AdaIN, removing the droplet artifacts by baking the style signal directly into the weights instead of normalizing activations. Second, a residual/skip-connection architecture replaces progressive growing, eliminating phase artifacts while maintaining training stability at high resolutions. Third, encourages the generator mapping to be well-conditioned — a fixed-size step in latent space produces a fixed-magnitude change in the image — which makes the generator easier to invert and improves overall image quality. reduces computational overhead by computing the term only every 16 mini-batches.

The impact

StyleGAN2 became the dominant generative model for high-fidelity image synthesis, powering research in face editing, image-to-image translation, and domain adaptation. Its weight demodulation technique influenced subsequent architectures, and its emphasis on generator conditioning and invertibility laid the groundwork for inversion methods. The quality bar it set — measured by on FFHQ — directly motivated the development of diffusion models that eventually surpassed GANs. The paper also introduced projection-based network attribution, enabling reliable detection of GAN-generated images.

Imagine a portrait studio with three problems. First, the painter's brush leaves tiny water droplets on the canvas — invisible from afar, but unmistakable up close. Second, no matter how the subject turns their head, the eyes and teeth stay stuck in the same spot on the canvas, as if glued there. Third, when the painter tries to slightly adjust a face — say, rotating it two degrees — the result is unpredictable: sometimes the whole face reshapes wildly.

StyleGAN2 is the story of diagnosing each of these flaws and prescribing a surgical fix. The droplets come from how the painter normalizes brush strokes — fix the normalization. The stuck features come from how the painter learned to paint progressively — replace that learning schedule. The unpredictable adjustments come from a poorly calibrated mapping between ideas and brushwork — add a regularizer that forces the mapping to be smooth.

Artifact 1: water droplets from normalization

StyleGAN uses (AdaIN) to inject style information into each layer of the generator. AdaIN works by normalizing each to zero mean and unit variance, then scaling and shifting it using parameters derived from the style vector. The concept is elegant: the content of the feature map is preserved, but its statistical signature — its "style" — is overwritten by the .

But the authors discovered a hidden side effect. When the generator wants to encode a strong signal in one feature map, washes it away because it forces unit variance. The generator learns a workaround: it creates a single dominant spike — a localized pixel with enormous value — that takes over the statistics of the entire feature map. After normalization, this spike is suppressed in magnitude, but it has already distorted the mean and variance, effectively letting the generator sneak its desired signal through the normalization gate. These spikes manifest as the characteristic "water droplet" artifacts visible in generated images.

Think of it like a classroom where the teacher forces everyone to speak at the same volume. One student wants to shout — so they create a distraction (the spike) that throws off the teacher's volume measurement, letting them get their message through. The distraction is the artifact.

Open in Lab
Toggle between AdaIN and weight demodulation to see how the normalization change eliminates the droplet artifacts from feature maps.
The demo wakes as you arrive…

The fix: weight demodulation instead of instance normalization

The key insight is: why normalize the activations after convolution when you can normalize the convolution weights before it? If the weights themselves are scaled so that the output has unit variance, there is no need for instance normalization — and no opportunity for the generator to create spike artifacts.

The solution has two steps. Modulation scales the convolution weights by the style vector, exactly like AdaIN would scale the activations — but it does so on the weights, not on the feature map. Demodulation then normalizes the modulated weights so that the output feature map will have approximately unit standard deviation, assuming the input has unit variance.

The mental model: imagine a pipeline where water (the signal) flows through adjustable valves (convolution weights). AdaIN adjusts the flow after the water passes through — it measures the water pressure and forces it to a target level. But by then, the pressure spikes have already damaged the pipes. Weight demodulation adjusts the valves before the water flows — the pressure is correct from the start, so no damage occurs.

Formally, let wi,j,kw_{i,j,k} be the original convolution weights where ii indexes input channels, jj indexes output channels, and kk indexes spatial positions (kernel elements). The style vector provides a per-layer scaling factor sis_i for each input channel.

Step 1 — Modulation: Scale each weight by the corresponding style factor:

wi,j,k′=si⋅wi,j,kw'_{i,j,k} = s_i \cdot w_{i,j,k}
Modulation — style scaling applied to weights — Each weight is multiplied by the style factor sis_i for its input channel. This is equivalent to scaling the input activations by sis_i before the convolution — the same effect as AdaIN's scaling, but applied to the weights.

Step 2 — Demodulation: Normalize the modulated weights so the output has unit standard deviation:

wi,j,k′′=wi,j,k′∑i,kw′i,j,k 2+ϵw''_{i,j,k} = \frac{w'_{i,j,k}}{\sqrt{\sum_{i,k} {w'}_{i,j,k}^{\,2} + \epsilon}}
Demodulation — ensuring unit output variance — For each output channel jj, we divide by the L2L_2 norm of the modulated weights across all input channels and spatial positions. This ensures the standard deviation of the output feature map is approximately 1, removing the need for instance normalization entirely. The small constant ϵ\epsilon prevents division by zero.
Open in Lab
See how style modulation scales the convolution weights, then demodulation normalizes them — achieving the same style control as AdaIN without touching the activations.
The demo wakes as you arrive…

Artifact 2: phase artifacts from progressive growing

Progressive growing trains the generator by starting at low resolution (4×4) and gradually adding higher-resolution layers during training. This was a breakthrough in ProGAN and StyleGAN for stabilizing high-resolution training. But the authors of StyleGAN2 noticed a subtle consequence: facial features like eyes and teeth have a strong preference for specific pixel locations on the canvas.

When interpolating between two latent codes — smoothly transitioning from one face to another — features that should move (because the face is turning) instead stay anchored to fixed positions. The eyes slide into place at the same coordinates regardless of the face's pose. The effect is like a puppet show where the eyes are painted on the backdrop rather than on the puppet: the puppet moves, but the eyes do not.

The cause is that progressive growing trains each resolution independently at first. The lower-resolution layers learn to place features at specific absolute positions, and these become "anchor points" that the higher-resolution layers cannot override. The features are stuck because the early layers set where they go, and the later layers just refine details without moving the positions.

Open in Lab
Compare progressive growing with the skip/residual architecture. Notice how the skip architecture lets all resolutions contribute simultaneously from the start.
The demo wakes as you arrive…

StyleGAN2 replaces progressive growing with a fixed-depth architecture that trains all resolutions from the start. The generator uses skip connections — each resolution contributes to the final image via bilinear and addition. The uses residual connections with downsampling. This design is inspired by the MSG-GAN approach.

The crucial difference: with progressive growing, the 4×4 layer has a monopoly on structure during early training. With skip connections, every layer contributes from the beginning, and the balance of contributions shifts organically during training rather than being imposed by a rigid schedule.

Smoothing the latent space: path length regularization

Even after fixing the artifacts, there remains a deeper question: how well-behaved is the mapping from latent codes to images? Ideally, if you take a small step in any direction in the latent space W\mathcal{W}, the corresponding change in the generated image should be proportional and consistent — not wildly different depending on where you are or which direction you step.

A generator with this property is called well-conditioned. Well-conditioning matters for two reasons. First, it makes latent space interpolation smooth and meaningful — walking between two faces produces a natural morph rather than jarring jumps. Second, it makes the generator invertible: given a real image, you can find the latent code that produced it, because the mapping does not fold or collapse.

The authors measure conditioning by looking at the Jacobian Jw\mathbf{J}_w of the generator — the matrix of partial derivatives that describes how each pixel changes in response to each latent dimension. Good conditioning means this Jacobian has consistent scale everywhere: a step of length δ\delta in W\mathcal{W} always produces an image change of roughly the same magnitude.

Ew,y∼N(0,I)(∥JwTy∥2−a)2\mathbb{E}_{\mathbf{w}, \mathbf{y} \sim \mathcal{N}(0, \mathbf{I})} \left( \|\mathbf{J}_{\mathbf{w}}^T \mathbf{y}\|_2 - a \right)^2
Path length regularizer — encouraging uniform Jacobian scale — A random direction y\mathbf{y} is projected through the transposed Jacobian JwT\mathbf{J}_{\mathbf{w}}^T. The regularizer penalizes the deviation of this projected length from a running average aa. This forces the Jacobian to have a consistent scale everywhere in W\mathcal{W}: every direction, at every point, produces approximately the same magnitude of image change.
Open in Lab
Explore how path length regularization makes the latent space smoother. Drag through the latent space with and without regularization to see the difference.
The demo wakes as you arrive…

Practical trick: lazy regularization

Computing the path length regularizer (and the R1 gradient penalty for the discriminator) at every training step is expensive — it requires a full backward pass through the generator just for the regularization term. The authors discovered that these regularizers change slowly during training, so computing them every step is wasteful.

Lazy regularization computes the regularization term once every kk mini-batches (they use k=16k = 16) and scales the regularization weight by kk to compensate. This reduces training cost significantly with virtually no impact on final quality. The intuition: regularizers guide the landscape, and the landscape shifts slowly. You do not need to re-measure the terrain every step — every 16 steps is frequent enough to keep the training on track.

How well does the generator use its resolution?

The authors introduced a diagnostic to measure whether the generator actually uses the full output resolution or concentrates its detail at lower frequencies. By masking different frequency bands and measuring the perceptual impact, they found that the original StyleGAN underutilizes its higher resolutions — much of the fine detail is actually determined at lower layers.

This motivated training larger models (Config F → Config E → Config F with more channels) and confirmed that StyleGAN2's architectural changes let the generator distribute detail more effectively across all resolution levels. The final model achieves an FID of 2.84 on FFHQ at 1024×1024 resolution, a significant improvement over StyleGAN's 4.40.

Open in Lab
Compare the incremental improvements across StyleGAN2 configurations — from baseline (A) to the final model (F) — and see how each change affects FID and PPL.
The demo wakes as you arrive…

Projecting images back into the latent space

Because path length regularization makes the generator well-conditioned, it becomes possible to project a given image back into the latent space — finding the latent code w\mathbf{w} that, when passed through the generator, reproduces the image as closely as possible. This is done by optimizing over w\mathbf{w} to minimize the (LPIPS) between the generated and target images.

This projection has a practical application: network attribution. If you suspect an image was generated by a specific GAN, project it into that GAN's latent space. If the projection is successful (low reconstruction error), the image likely came from that network. If projection fails (high error), the image either is real or came from a different generator.

The well-conditioned generator produced by StyleGAN2 makes this projection far more reliable than with StyleGAN, because the smooth latent space means the optimizer can find the correct code without getting stuck in local minima.

Putting it all together: the StyleGAN2 architecture

The complete StyleGAN2 pipeline retains the two-stage structure of StyleGAN. A ff transforms the random input z∼N(0,I)\mathbf{z} \sim \mathcal{N}(0, I) into an intermediate latent code w∈W\mathbf{w} \in \mathcal{W}. The synthesis network gg then converts w\mathbf{w} into an image, with each layer receiving its own style derived from w\mathbf{w} via learned affine transformations.

What changes in StyleGAN2 relative to StyleGAN:

  • AdaIN is replaced by weight modulation + demodulation — style is injected through the weights, not through activation statistics.
  • Progressive growing is replaced by a skip-connection generator and a residual discriminator — all resolutions train simultaneously.
  • Path length regularization is added to the generator loss — the Jacobian is encouraged to have uniform scale.
  • Lazy regularization computes regularizers every 16 steps — saving compute with no quality loss.
  • Noise inputs and biases are moved outside the style block — a minor reorganization that simplifies the architecture.
Open in Lab
Full architecture diagram of StyleGAN2 showing the mapping network, weight modulation/demodulation, skip connections, and noise injection.
The demo wakes as you arrive…
Weight modulation and demodulation — PyTorch pseudocodepython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn.functional as F

def modulated_conv2d(x, weight, style, demodulate=True, eps=1e-8):
    """
    x:       input features  [batch, in_ch, H, W]
    weight:  conv kernel     [out_ch, in_ch, kH, kW]
    style:   style vector    [batch, in_ch]  (from mapping network)
    """
    batch, in_ch, H, W = x.shape
    out_ch, _, kH, kW = weight.shape

    # Step 1: Modulation — scale weights by style
    # style[:, None, :, None, None] broadcasts across [out_ch, in_ch, kH, kW]
    w = weight[None] * style[:, None, :, None, None]  # [B, out_ch, in_ch, kH, kW]

    # Step 2: Demodulation — normalize so output has unit variance
    if demodulate:
        sigma = (w.square().sum(dim=[2, 3, 4]) + eps).rsqrt()  # [B, out_ch]
        w = w * sigma[:, :, None, None, None]

    # Reshape for grouped convolution (one group per batch element)
    x = x.reshape(1, batch * in_ch, H, W)
    w = w.reshape(batch * out_ch, in_ch, kH, kW)
    out = F.conv2d(x, w, groups=batch, padding=kH // 2)
    return out.reshape(batch, out_ch, H, W)

Timeline: from ProGAN to diffusion

  1. 2018

    ProGAN — progressive growing

    Introduced progressive growing to stabilize high-resolution GAN training. Start from 4×4, gradually add layers up to 1024×1024.

  2. 2019

    StyleGAN — style-based architecture

    Introduced the mapping network, AdaIN-based style injection, and stochastic noise inputs. Achieved unprecedented image quality on faces (FFHQ).

  3. 2020

    StyleGAN2 — this paper

    Weight demodulation, skip/residual architecture, path length regularization. Eliminated artifacts, improved FID from 4.40 to 2.84 on FFHQ 1024².

  4. 2020

    StyleGAN2-ADA — limited data training

    Added adaptive discriminator augmentation to train with very limited data without overfitting. Made StyleGAN2 practical for small datasets.

  5. 2021

    StyleGAN3 — alias-free generation

    Addressed remaining spatial aliasing issues. Ensured continuous equivariance under translation and rotation for coherent video and animation generation.

  6. 2021

    Diffusion beats GANs

    Dhariwal & Nichol showed that diffusion models can surpass the FID scores set by StyleGAN2, marking the shift from GAN dominance to diffusion dominance in image generation.

StyleGAN2's legacy is paradoxical: it set the quality bar so high that it motivated the very class of models — diffusion — that eventually displaced GANs from the top of the leaderboard. Yet StyleGAN2's architectural lessons endure. Weight demodulation influenced normalization design in subsequent architectures. Path length regularization established the importance of well-conditioned generators. And the idea that image quality can be improved by diagnosing and surgically removing specific artifacts — rather than simply scaling up — remains a powerful principle in generative modeling.

CitationKarras, Laine, Aittala, Hellsten, Lehtinen, Aila. Analyzing and Improving the Image Quality of StyleGAN. CVPR, 2020.

Terms in this paper