Generative Models2018intermediate9 min read

Progressive Growing of GANs for Improved Quality, Stability, and Variation

النمو التدريجي لشبكات GAN لتحسين الجودة والاستقرار والتنوع

Karras, T. · Aila, T. · Laine, S. · Lehtinen, J. — ICLR

The problem

By 2017, Generative Adversarial Networks could produce sharp images at small resolutions (64×64 to 256×256), but scaling to higher resolutions like 1024×1024 was extremely unstable. The and would spiral into unhealthy competition — gradients would explode or vanish, the generator would collapse to producing a handful of near-identical outputs (), and would simply diverge. The fundamental problem: asking a randomly initialized network to learn all frequency scales simultaneously is too hard.

The contribution

A curriculum-learning approach to training: start both generator and discriminator at 4×4 resolution, then progressively add convolutional layers that double the resolution, with each new layer smoothly faded in via an alpha-blending parameter. This is complemented by three stabilization techniques: (runtime weight scaling instead of careful initialization), pixelwise feature in the generator, and minibatch standard deviation in the discriminator to fight mode collapse. The result: photorealistic 1024×1024 face images and a 5.4× training speedup.

The impact

Progressive GAN was the first model to generate photorealistic faces, proving that GANs can scale far beyond previous limits. Its methodology and stabilization techniques became standard practice and directly enabled StyleGAN, which inherited progressive growing and added style-based control. The CelebA-HQ dataset created for this paper remains a benchmark for face generation research.

Teaching a GAN to generate 1024×1024 images from scratch is like asking a first-year art student to paint a photorealistic mural on day one — overwhelmed by the sheer number of details, they freeze or produce chaos.

Progressive GAN takes the master-class approach: start with a tiny canvas (4×4 pixels), learn to get the broad strokes right (overall face shape, background color), then gradually enlarge the canvas while adding finer details at each step — eyes, nose, lips, skin texture, individual hairs. By the time the student reaches the full mural, they've built competence layer by layer.

The problem: GANs cannot scale to high resolution

By 2017 the GAN landscape was stuck at a resolution ceiling. Models like DCGAN could generate plausible 64×64 faces, and careful engineering pushed quality to 256×256, but going to 1024×1024 consistently failed. Three entangled problems were to blame:

  • Instability. At high resolution, the discriminator easily overpowers the generator early in training, producing extreme gradients that destabilize both networks. The adversarial game degenerates into an arms race where neither network converges.

  • Mode collapse. Under pressure from a strong discriminator, the generator learns to produce a small set of "safe" outputs that fool the discriminator instead of covering the full data distribution. Variation is sacrificed for survival.

  • Computational cost. Training at full resolution from the start means every iteration is maximally expensive, and most of that compute is wasted in the early epochs when the generator is producing noise anyway.

Open in Lab
Watch the generator grow from 4×4 to 1024×1024. Each stage adds new layers that are smoothly faded in, so existing knowledge is never disrupted.
The demo wakes as you arrive…

The core idea: grow the network resolution by resolution

The insight is a form of applied to image generation. Instead of attacking the full 1024×1024 problem from the start, decompose it into a sequence of progressively harder sub-problems:

Phase 1 (4×4): The generator maps a to a 4×4 and the discriminator classifies 4×4 images. At this resolution there are only ~48 pixels to get right — the network quickly learns the dominant colors and rough structure.

Phase 2 (8×8): New convolutional layers are added to both generator and discriminator, doubling the spatial resolution. The generator must now produce coherent 8×8 images, but it starts from a solid foundation — it already knows the broad strokes.

Phases 3–9 (16×16 → 1024×1024): The process repeats. Each new resolution adds finer detail — edges at 16×16, facial features at 64×64, skin pores at 512×512, individual hairs at 1024×1024.

The key engineering trick is that new layers are not added abruptly. Instead, they are faded in using a blending parameter α\alpha that linearly ramps from 0 to 1 over many training iterations. This avoids shocking the already-trained lower layers with sudden gradient changes.

Open in Lab
Drag the α slider to see how new layers are blended in. At α=0 the output comes entirely from the upsampled lower-resolution layer; at α=1 the new high-resolution layer takes full control.
The demo wakes as you arrive…
output=(1−α)⋅upsampled(old_layer)+α⋅new_layer\text{output} = (1 - \alpha) \cdot \text{upsampled}(\text{old\_layer}) + \alpha \cdot \text{new\_layer}
Fade-in blending — smooth transition between resolution stages — α ramps linearly from 0 to 1. At the start the new layer contributes nothing; by the end it has completely replaced the upsampled shortcut. All existing layers remain trainable throughout.

Think of it as a dimmer switch: the old output is at full brightness and the new layer's output is in darkness. Over thousands of iterations, you slowly turn the dimmer — fading out the old, fading in the new — until the new layer is fully "on" and the old shortcut is fully "off." The network never experiences a jarring transition.

Stabilization toolkit: three techniques that keep training healthy

Progressive growing alone is not enough. GANs are notorious for the signal magnitude escalation problem: the generator and discriminator amplify each other's outputs in an unhealthy feedback loop. The paper introduces three complementary techniques — none of which have learnable parameters — that together control this escalation and encourage diversity.

w^i=wi⋅c,c=2fan_in\hat{w}_i = w_i \cdot c, \quad c = \sqrt{\frac{2}{\text{fan\_in}}}
Equalized learning rate — runtime weight scaling — fan_in is the number of input connections to the layer. The scaling constant c is the same one used in He initialization, but applied at every forward pass instead of only at initialization. This makes all layers learn at the same speed.
Open in Lab
Compare standard initialization (left) vs. equalized learning rate (right). Notice how the equalized version keeps all layers updating at similar speeds.
The demo wakes as you arrive…
Open in Lab
The discriminator compares batch diversity. Left: a diverse batch (high minibatch std) looks realistic. Right: a collapsed batch (low std) is easy to catch as fake.
The demo wakes as you arrive…

Architecture walkthrough: from latent vector to megapixel face

The generator starts from a 512-dimensional latent vector sampled from a standard normal distribution. This vector is reshaped into a 4×4×512 feature map — think of it as an extremely compressed "seed" that contains all the variation the network will unfold.

Each resolution stage adds a block of two convolutional layers (3×3 filters) with activations, preceded by nearest-neighbor (2×) to double the spatial dimensions. The number of feature maps decreases as resolution grows: 512 at 4×4 down to 16 at 1024×1024. This mirrors how a painter works — broad color choices first, fine brush strokes last.

The discriminator is a mirror image: it takes a full-resolution image, applies strided convolutions to halve the spatial dimensions at each block, and ends with a single scalar verdict — real or fake. The toRGB and fromRGB layers (1×1 convolutions) bridge between the internal feature space and the RGB color space at each resolution.

Open in Lab
Click on any resolution stage to explore its layers, feature counts, and role in the progressive pipeline.
The demo wakes as you arrive…

Progressive growing in code

Progressive GAN — fade-in layer blendingpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn
import torch.nn.functional as F

class FadeInBlock(nn.Module):
    """Smooth transition when a new resolution block is added."""
    def __init__(self, old_block, new_block):
        super().__init__()
        self.old_block = old_block    # upsampled shortcut
        self.new_block = new_block    # new conv layers
        self.alpha = 0.0             # ramps from 0 → 1

    def forward(self, x):
        # Old path: upsample then 1×1 toRGB
        old_out = F.interpolate(self.old_block(x), scale_factor=2)
        # New path: full convolutional processing
        new_out = self.new_block(x)
        # Blend: start from old, slowly transition to new
        return (1 - self.alpha) * old_out + self.alpha * new_out

class PixelNorm(nn.Module):
    """Normalize feature vector at each pixel to unit length."""
    def forward(self, x):
        return x / (x.pow(2).mean(dim=1, keepdim=True).sqrt() + 1e-8)

class EqualizedConv2d(nn.Module):
    """Conv2d with runtime weight scaling for equalized learning rate."""
    def __init__(self, in_ch, out_ch, kernel_size, **kw):
        super().__init__()
        self.conv = nn.Conv2d(in_ch, out_ch, kernel_size, **kw)
        nn.init.normal_(self.conv.weight)          # N(0,1) init
        self.scale = (2 / (in_ch * kernel_size**2)) ** 0.5  # He constant

    def forward(self, x):
        return self.conv(x * self.scale)  # scale at runtime, not init

Results: unprecedented quality and speed

Progressive GAN achieved several milestones simultaneously. On the newly created CelebA-HQ dataset, it generated the first-ever photorealistic 1024×1024 face images from a GAN — images convincing enough that human viewers struggled to distinguish them from real photographs.

Training speed improved dramatically: the progressive approach was 5.4× faster than training at full resolution from scratch for 1024×1024 images. This speedup comes from spending most training iterations at low resolutions where computation is cheap.

On CIFAR-10, the model achieved an Inception Score of 8.80, setting a new state-of-the-art for unsupervised image generation at the time. The paper also introduced the Sliced (SWD) as a multi-scale evaluation metric that measures both quality and variation, complementing the Fréchet Inception Distance.

Open in Lab
Progressive training vs. direct full-resolution training. The progressive approach reaches convergence much faster by spending early epochs at cheap low resolutions.
The demo wakes as you arrive…

Impact and legacy

Progressive GAN's influence extends far beyond its own results. The progressive training paradigm and the stabilization techniques it introduced became foundational building blocks for the next generation of generative models:

  1. 2014

    GAN (Goodfellow et al.)

    The original adversarial framework: a generator tries to fool a discriminator. Produced blurry, small images but established the foundational idea.

  2. 2016

    DCGAN

    Convolutional architecture guidelines for stable GAN training. Reached 64×64 with decent quality and became the standard baseline.

  3. 2017

    WGAN / WGAN-GP

    Wasserstein distance as loss + gradient penalty for stable training. Improved training stability but still limited in resolution.

  4. 2018

    Progressive GAN (this paper)

    Progressive training + stabilization techniques → first photorealistic 1024×1024 faces. Proved that curriculum learning unlocks high-resolution generation.

  5. 2019

    StyleGAN

    Built on Progressive GAN's foundation, adding Adaptive Instance Normalization for style control at each resolution. Produced even more convincing faces with disentangled control over features.

  6. 2020

    StyleGAN2

    Removed progressive growing in favor of skip connections, but kept equalized learning rate and many stabilization principles from the original ProGAN.

CitationKarras, Aila, Laine, Lehtinen. Progressive Growing of GANs for Improved Quality, Stability, and Variation. ICLR, 2018.

Terms in this paper