Generative Models2020advanced14 min read
Analyzing and Improving the Image Quality of StyleGAN
تحليل جودة الصور في StyleGAN وتحسينها
Karras, T. · Laine, S. · Aittala, M. · Hellsten, J. · Lehtinen, J. · Aila, T. — CVPR
The problem
StyleGAN achieved state-of-the-art results in unconditional image generation, but closer inspection revealed systematic quality issues. Generated images contained characteristic water-droplet artifacts caused by the AdaIN normalization. Progressive growing — the technique that enabled high-resolution — caused phase artifacts where facial features like eyes and teeth appeared glued to fixed canvas positions instead of moving naturally. Furthermore, the mapping from latent codes to images was poorly conditioned: small changes in could cause unpredictable jumps in the output, making the difficult to invert and control.
The contribution
Three architectural and training improvements that together eliminate StyleGAN's quality issues. First, replaces AdaIN, removing the droplet artifacts by baking the style signal directly into the weights instead of normalizing activations. Second, a residual/skip-connection architecture replaces progressive growing, eliminating phase artifacts while maintaining training stability at high resolutions. Third, encourages the generator mapping to be well-conditioned — a fixed-size step in latent space produces a fixed-magnitude change in the image — which makes the generator easier to invert and improves overall image quality. reduces computational overhead by computing the term only every 16 mini-batches.
The impact
StyleGAN2 became the dominant generative model for high-fidelity image synthesis, powering research in face editing, image-to-image translation, and domain adaptation. Its weight demodulation technique influenced subsequent architectures, and its emphasis on generator conditioning and invertibility laid the groundwork for inversion methods. The quality bar it set — measured by on FFHQ — directly motivated the development of diffusion models that eventually surpassed GANs. The paper also introduced projection-based network attribution, enabling reliable detection of GAN-generated images.
Imagine a portrait studio with three problems. First, the painter's brush leaves tiny water droplets on the canvas — invisible from afar, but unmistakable up close. Second, no matter how the subject turns their head, the eyes and teeth stay stuck in the same spot on the canvas, as if glued there. Third, when the painter tries to slightly adjust a face — say, rotating it two degrees — the result is unpredictable: sometimes the whole face reshapes wildly.
StyleGAN2 is the story of diagnosing each of these flaws and prescribing a surgical fix. The droplets come from how the painter normalizes brush strokes — fix the normalization. The stuck features come from how the painter learned to paint progressively — replace that learning schedule. The unpredictable adjustments come from a poorly calibrated mapping between ideas and brushwork — add a regularizer that forces the mapping to be smooth.
Artifact 1: water droplets from normalization
StyleGAN uses (AdaIN) to inject style information into each layer of the generator. AdaIN works by normalizing each to zero mean and unit variance, then scaling and shifting it using parameters derived from the style vector. The concept is elegant: the content of the feature map is preserved, but its statistical signature — its "style" — is overwritten by the .
But the authors discovered a hidden side effect. When the generator wants to encode a strong signal in one feature map, washes it away because it forces unit variance. The generator learns a workaround: it creates a single dominant spike — a localized pixel with enormous value — that takes over the statistics of the entire feature map. After normalization, this spike is suppressed in magnitude, but it has already distorted the mean and variance, effectively letting the generator sneak its desired signal through the normalization gate. These spikes manifest as the characteristic "water droplet" artifacts visible in generated images.
Think of it like a classroom where the teacher forces everyone to speak at the same volume. One student wants to shout — so they create a distraction (the spike) that throws off the teacher's volume measurement, letting them get their message through. The distraction is the artifact.
The fix: weight demodulation instead of instance normalization
The key insight is: why normalize the activations after convolution when you can normalize the convolution weights before it? If the weights themselves are scaled so that the output has unit variance, there is no need for instance normalization — and no opportunity for the generator to create spike artifacts.
The solution has two steps. Modulation scales the convolution weights by the style vector, exactly like AdaIN would scale the activations — but it does so on the weights, not on the feature map. Demodulation then normalizes the modulated weights so that the output feature map will have approximately unit standard deviation, assuming the input has unit variance.
The mental model: imagine a pipeline where water (the signal) flows through adjustable valves (convolution weights). AdaIN adjusts the flow after the water passes through — it measures the water pressure and forces it to a target level. But by then, the pressure spikes have already damaged the pipes. Weight demodulation adjusts the valves before the water flows — the pressure is correct from the start, so no damage occurs.
Formally, let be the original convolution weights where indexes input channels, indexes output channels, and indexes spatial positions (kernel elements). The style vector provides a per-layer scaling factor for each input channel.
Step 1 — Modulation: Scale each weight by the corresponding style factor:
Step 2 — Demodulation: Normalize the modulated weights so the output has unit standard deviation:
Artifact 2: phase artifacts from progressive growing
Progressive growing trains the generator by starting at low resolution (4×4) and gradually adding higher-resolution layers during training. This was a breakthrough in ProGAN and StyleGAN for stabilizing high-resolution training. But the authors of StyleGAN2 noticed a subtle consequence: facial features like eyes and teeth have a strong preference for specific pixel locations on the canvas.
When interpolating between two latent codes — smoothly transitioning from one face to another — features that should move (because the face is turning) instead stay anchored to fixed positions. The eyes slide into place at the same coordinates regardless of the face's pose. The effect is like a puppet show where the eyes are painted on the backdrop rather than on the puppet: the puppet moves, but the eyes do not.
The cause is that progressive growing trains each resolution independently at first. The lower-resolution layers learn to place features at specific absolute positions, and these become "anchor points" that the higher-resolution layers cannot override. The features are stuck because the early layers set where they go, and the later layers just refine details without moving the positions.
StyleGAN2 replaces progressive growing with a fixed-depth architecture that trains all resolutions from the start. The generator uses skip connections — each resolution contributes to the final image via bilinear and addition. The uses residual connections with downsampling. This design is inspired by the MSG-GAN approach.
The crucial difference: with progressive growing, the 4×4 layer has a monopoly on structure during early training. With skip connections, every layer contributes from the beginning, and the balance of contributions shifts organically during training rather than being imposed by a rigid schedule.
Smoothing the latent space: path length regularization
Even after fixing the artifacts, there remains a deeper question: how well-behaved is the mapping from latent codes to images? Ideally, if you take a small step in any direction in the latent space , the corresponding change in the generated image should be proportional and consistent — not wildly different depending on where you are or which direction you step.
A generator with this property is called well-conditioned. Well-conditioning matters for two reasons. First, it makes latent space interpolation smooth and meaningful — walking between two faces produces a natural morph rather than jarring jumps. Second, it makes the generator invertible: given a real image, you can find the latent code that produced it, because the mapping does not fold or collapse.
The authors measure conditioning by looking at the Jacobian of the generator — the matrix of partial derivatives that describes how each pixel changes in response to each latent dimension. Good conditioning means this Jacobian has consistent scale everywhere: a step of length in always produces an image change of roughly the same magnitude.
Practical trick: lazy regularization
Computing the path length regularizer (and the R1 gradient penalty for the discriminator) at every training step is expensive — it requires a full backward pass through the generator just for the regularization term. The authors discovered that these regularizers change slowly during training, so computing them every step is wasteful.
Lazy regularization computes the regularization term once every mini-batches (they use ) and scales the regularization weight by to compensate. This reduces training cost significantly with virtually no impact on final quality. The intuition: regularizers guide the landscape, and the landscape shifts slowly. You do not need to re-measure the terrain every step — every 16 steps is frequent enough to keep the training on track.
How well does the generator use its resolution?
The authors introduced a diagnostic to measure whether the generator actually uses the full output resolution or concentrates its detail at lower frequencies. By masking different frequency bands and measuring the perceptual impact, they found that the original StyleGAN underutilizes its higher resolutions — much of the fine detail is actually determined at lower layers.
This motivated training larger models (Config F → Config E → Config F with more channels) and confirmed that StyleGAN2's architectural changes let the generator distribute detail more effectively across all resolution levels. The final model achieves an FID of 2.84 on FFHQ at 1024×1024 resolution, a significant improvement over StyleGAN's 4.40.
Projecting images back into the latent space
Because path length regularization makes the generator well-conditioned, it becomes possible to project a given image back into the latent space — finding the latent code that, when passed through the generator, reproduces the image as closely as possible. This is done by optimizing over to minimize the (LPIPS) between the generated and target images.
This projection has a practical application: network attribution. If you suspect an image was generated by a specific GAN, project it into that GAN's latent space. If the projection is successful (low reconstruction error), the image likely came from that network. If projection fails (high error), the image either is real or came from a different generator.
The well-conditioned generator produced by StyleGAN2 makes this projection far more reliable than with StyleGAN, because the smooth latent space means the optimizer can find the correct code without getting stuck in local minima.
Putting it all together: the StyleGAN2 architecture
The complete StyleGAN2 pipeline retains the two-stage structure of StyleGAN. A transforms the random input into an intermediate latent code . The synthesis network then converts into an image, with each layer receiving its own style derived from via learned affine transformations.
What changes in StyleGAN2 relative to StyleGAN:
- AdaIN is replaced by weight modulation + demodulation — style is injected through the weights, not through activation statistics.
- Progressive growing is replaced by a skip-connection generator and a residual discriminator — all resolutions train simultaneously.
- Path length regularization is added to the generator loss — the Jacobian is encouraged to have uniform scale.
- Lazy regularization computes regularizers every 16 steps — saving compute with no quality loss.
- Noise inputs and biases are moved outside the style block — a minor reorganization that simplifies the architecture.
Simplified to show the idea — not the real implementation.
import torch
import torch.nn.functional as F
def modulated_conv2d(x, weight, style, demodulate=True, eps=1e-8):
"""
x: input features [batch, in_ch, H, W]
weight: conv kernel [out_ch, in_ch, kH, kW]
style: style vector [batch, in_ch] (from mapping network)
"""
batch, in_ch, H, W = x.shape
out_ch, _, kH, kW = weight.shape
# Step 1: Modulation — scale weights by style
# style[:, None, :, None, None] broadcasts across [out_ch, in_ch, kH, kW]
w = weight[None] * style[:, None, :, None, None] # [B, out_ch, in_ch, kH, kW]
# Step 2: Demodulation — normalize so output has unit variance
if demodulate:
sigma = (w.square().sum(dim=[2, 3, 4]) + eps).rsqrt() # [B, out_ch]
w = w * sigma[:, :, None, None, None]
# Reshape for grouped convolution (one group per batch element)
x = x.reshape(1, batch * in_ch, H, W)
w = w.reshape(batch * out_ch, in_ch, kH, kW)
out = F.conv2d(x, w, groups=batch, padding=kH // 2)
return out.reshape(batch, out_ch, H, W)
Timeline: from ProGAN to diffusion
2018
ProGAN — progressive growing
Introduced progressive growing to stabilize high-resolution GAN training. Start from 4×4, gradually add layers up to 1024×1024.
2019
StyleGAN — style-based architecture
Introduced the mapping network, AdaIN-based style injection, and stochastic noise inputs. Achieved unprecedented image quality on faces (FFHQ).
2020
StyleGAN2 — this paper
Weight demodulation, skip/residual architecture, path length regularization. Eliminated artifacts, improved FID from 4.40 to 2.84 on FFHQ 1024².
2020
StyleGAN2-ADA — limited data training
Added adaptive discriminator augmentation to train with very limited data without overfitting. Made StyleGAN2 practical for small datasets.
2021
StyleGAN3 — alias-free generation
Addressed remaining spatial aliasing issues. Ensured continuous equivariance under translation and rotation for coherent video and animation generation.
2021
Diffusion beats GANs
Dhariwal & Nichol showed that diffusion models can surpass the FID scores set by StyleGAN2, marking the shift from GAN dominance to diffusion dominance in image generation.
StyleGAN2's legacy is paradoxical: it set the quality bar so high that it motivated the very class of models — diffusion — that eventually displaced GANs from the top of the leaderboard. Yet StyleGAN2's architectural lessons endure. Weight demodulation influenced normalization design in subsequent architectures. Path length regularization established the importance of well-conditioned generators. And the idea that image quality can be improved by diagnosing and surgically removing specific artifacts — rather than simply scaling up — remains a powerful principle in generative modeling.
CitationKarras, Laine, Aittala, Hellsten, Lehtinen, Aila. Analyzing and Improving the Image Quality of StyleGAN. CVPR, 2020.
Terms in this paper
- Generative Adversarial Network (GAN)الشبكات التوليدية التنافسية
- Generatorالموّلد التخليقي
- Discriminatorالـمُميِّز الحاكم
- Latent Spaceالفضاء الكامن
- Instance Normalizationتسوية النسخة
- Feature Mapخريطة السمات
- Adaptive Instance Normalizationالتطبيع التكيُّفي للنُّسَخ
- Weight Demodulationإزالة التشكيل من الأوزان
- Path Length Regularizationتسوية طول المسار
- Progressive Trainingالتدريب التدريجي
- Skip Connectionالاتصال التجاوزي
- Residual Connectionالوصلة التجاوزية
- Perceptual Lossالخسارة الإدراكية
- FIDمسافة فريشيه للبداية
- Bilinear Interpolationالاستيفاء الثنائي الخطي