Generative Models2018advanced12 min read

Glow: Generative Flow with Invertible 1×1 Convolutions

Glow: التدفق التوليدي بالالتفافات القابلة للعكس 1×1

Kingma, D. P. · Dhariwal, P. — NeurIPS

The problem

Generative models in 2018 faced a trilemma. GANs produced sharp images but had no and couldn't compute likelihoods. VAEs computed a lower bound on the likelihood but produced blurry samples. Autoregressive models computed exact likelihoods but generated pixels one at a time — impossibly slow for large images. No single offered all three: , efficient parallel sampling, and a useful .

The contribution

Glow proposed three architectural innovations for : (1) activation normalization (actnorm) that replaces and works with any , (2) 1×1 convolutions that generalize fixed permutations into a learnable mixing operation, and (3) an within a multi-scale squeeze-and-split architecture. Together, these produced the first flow-based model capable of generating realistic high-resolution face images while computing exact log-likelihoods — proving that you don't need adversarial or approximate inference to get good images.

The impact

Glow established normalizing flows as a viable paradigm for high-resolution image generation, sitting alongside GANs and VAEs. Its invertible 1×1 became a standard building block in subsequent flow architectures. The demonstration that exact likelihood optimization alone could produce realistic images influenced the broader generative modeling landscape, including the development of flow matching and continuous normalizing flows that power modern diffusion-adjacent methods.

Imagine a glassblower shaping molten glass. She starts with a perfectly round bubble (a Gaussian) and applies a sequence of precise twists, stretches, and compressions — each one reversible — until the glass takes the shape of an intricate vase.

Because every move can be undone exactly, she can take any finished vase and reverse the steps to recover the original bubble. This reversibility is what makes normalizing flows special: unlike a sculptor who chips away marble (irreversible), the glassblower transforms without losing information.

Glow is a recipe for building such a glassblower. Its key innovation: instead of prescribing which way to twist the glass (fixed permutations), it learns the optimal twist at every step — the invertible 1×1 convolution.

The generative trilemma: why we needed flows

By 2018, three families of generative models had emerged, each brilliant in one way but fundamentally limited in another:

  • GANs produced the sharpest images, but they had no encoder — you couldn't map a real image to a latent code — and they offered no way to compute how likely a given image was. Training was also notoriously unstable.
  • VAEs had both an encoder and a , but they optimized a lower bound on the (the ), not the likelihood itself. This approximation gap meant blurrier samples.
  • Autoregressive models (like PixelCNN) computed exact likelihoods, but they generated one pixel at a time, left to right, top to bottom — making synthesis extremely slow and impossible to parallelize.

What if a model could offer all three properties simultaneously — exact likelihood, parallel synthesis, and a meaningful latent space? This is precisely where normalizing flows enter.

Open in Lab
Compare the three generative paradigms. Flows are the only family offering all three desirable properties at once.
The demo wakes as you arrive…

The engine: change of variables

The mathematical foundation of all normalizing flows is the formula. The intuition: if you stretch a rubber sheet, the density of ink dots on it changes — areas that get stretched thin have fewer dots per unit area, and areas that get compressed have more. The determinant measures exactly this stretching factor.

Start with a simple you can evaluate — typically a standard Gaussian pZ(z)=N(0,I)p_Z(\mathbf{z}) = \mathcal{N}(\mathbf{0}, \mathbf{I}). Apply an invertible transformation ff to produce a complex distribution over images x=f(z)\mathbf{x} = f(\mathbf{z}). The change of variables formula tells us the resulting density:

log⁡p(x)=log⁡pZ ⁣(f−1(x))+log⁡ ⁣∣det⁡ ⁣(∂f−1∂x)∣\log p(\mathbf{x}) = \log p_Z\!\bigl(f^{-1}(\mathbf{x})\bigr) + \log\!\left|\det\!\left(\frac{\partial f^{-1}}{\partial \mathbf{x}}\right)\right|
Change of variables — the master equation for normalizing flows — The log-probability of a data point x equals the log-probability of its latent code z under the base Gaussian, plus a correction term (the log-determinant of the Jacobian) that accounts for how the transformation stretches or compresses space.

When we compose KK invertible steps f=f1∘f2∘⋯∘fKf = f_1 \circ f_2 \circ \dots \circ f_K, the log-determinant simply sums: log⁡ ⁣∣det⁡ ⁣(∂f−1∂x)∣=∑i=1Klog⁡ ⁣∣det⁡ ⁣(∂fi−1∂hi−1)∣\log\!\left|\det\!\left(\frac{\partial f^{-1}}{\partial \mathbf{x}}\right)\right| = \sum_{i=1}^{K} \log\!\left|\det\!\left(\frac{\partial f_i^{-1}}{\partial \mathbf{h}_{i-1}}\right)\right|

The entire design challenge then becomes: how do you build invertible functions whose Jacobian determinant is cheap to compute? This is where Glow's three-part flow step comes in.

Open in Lab
Drag the slider to see how an invertible transformation warps a Gaussian into a complex distribution. The Jacobian correction accounts for stretching.
The demo wakes as you arrive…

Glow's flow step: actnorm → 1×1 conv → coupling

Every step of flow in Glow is a composition of exactly three invertible layers, applied in sequence. Think of it as a three-stage factory line: each stage handles a different aspect of the transformation, and because each stage is individually reversible, the whole pipeline is reversible too.

Open in Lab
Click each stage to learn what it does. The three stages compose into one invertible flow step.
The demo wakes as you arrive…
Open in Lab
See how the 1×1 convolution mixes channels. Drag the matrix entries to see how information flows between channels.
The demo wakes as you arrive…
ya=xa,yb=s(xa)⊙xb+t(xa)\mathbf{y}_a = \mathbf{x}_a, \qquad \mathbf{y}_b = \mathbf{s}(\mathbf{x}_a) \odot \mathbf{x}_b + \mathbf{t}(\mathbf{x}_a)
Affine coupling — half transforms, half stays — Split the channels into two halves. One half passes through unchanged; the other is scaled and shifted by learned functions of the unchanged half. The trick: the Jacobian stays triangular, so its determinant is trivially the product of the scale factors.

Multi-scale architecture: squeeze, flow, split

A single flow step cannot capture all the complexity of a natural image. Glow stacks KK flow steps into one level, then repeats for LL levels in a multi-scale architecture borrowed from RealNVP. At each level:

  1. Squeeze: reshape each 2×2×c2 \times 2 \times c spatial block into 1×1×4c1 \times 1 \times 4c. This trades spatial resolution for channel depth, letting subsequent 1×1 convolutions mix information from neighboring pixels.
  2. KK flow steps: apply the three-part flow step (actnorm → 1×1 conv → coupling) KK times.
  3. Split: divide channels in half. One half goes directly to the latent space (as a factored-out Gaussian). The other half continues to the next level.

This creates a hierarchical latent space: early splits capture coarse, global features (overall shape, lighting) while later splits encode fine details (textures, edges). The factoring-out also makes training and memory more efficient — half the data exits the pipeline early at each level.

Open in Lab
Explore the multi-scale pipeline. Watch how spatial resolution shrinks and channel depth grows at each level, with half the channels split off to the latent space.
The demo wakes as you arrive…

Why learn the permutation?

In RealNVP, the choice of which channels go to xa\mathbf{x}_a and which go to xb\mathbf{x}_b was fixed — either a simple reversal or a checkerboard pattern. This means the coupling layer always saw the same partition. Imagine always shuffling a deck of cards the same way: you'd never get a truly random arrangement.

Glow's invertible 1×1 convolution is a learned full-rank linear transformation of the channel dimension. It generalizes three possibilities in ascending expressiveness:

  1. Reverse ordering (NICE) — fixed, zero parameters
  2. Fixed random permutation (RealNVP) — fixed, zero learnable parameters
  3. Learned 1×1 convolution (Glow) — c2c^2 learnable parameters

The paper shows a clear progression in log-likelihood as you move from (1) to (2) to (3), confirming that letting the model learn how to partition channels significantly improves the flow's expressiveness.

log⁡ ⁣∣det⁡ ⁣(∂conv2D(x;W)∂x)∣=h⋅w⋅log⁡ ⁣∣det⁡(W)∣\log\!\left|\det\!\left(\frac{\partial \text{conv2D}(\mathbf{x}; W)}{\partial \mathbf{x}}\right)\right| = h \cdot w \cdot \log\!\left|\det(W)\right|
Log-determinant of the 1×1 convolution — Because the same W matrix is applied at every spatial position independently, the log-determinant of the full spatial convolution is simply h × w copies of log|det(W)|. With LU decomposition, computing det(W) costs O(c³) once, not O((h·w·c)³).

Training: just maximize log-likelihood

One of the most appealing aspects of Glow is how simple its training objective is. There is no adversarial game (as in GANs), no variational bound (as in VAEs), and no autoregressive decomposition. The is simply the negative log-likelihood averaged over the training set:

L=−1N∑i=1Nlog⁡pθ(xi)\mathcal{L} = -\frac{1}{N} \sum_{i=1}^{N} \log p_\theta(\mathbf{x}_i)

Each term decomposes via the change of variables formula into a Gaussian log-probability plus a sum of log-determinants — all quantities we can compute exactly.

Synthesis runs the flow in reverse: sample z∼N(0,T2I)\mathbf{z} \sim \mathcal{N}(\mathbf{0}, T^2\mathbf{I}) and apply f−1f^{-1}. The TT controls quality: lower TT (e.g. 0.7) samples closer to the mean, producing sharper but less diverse images.

Interpolation is equally simple: encode two real images as zA\mathbf{z}_A and zB\mathbf{z}_B, then decode (1−t)zA+tzB(1-t)\mathbf{z}_A + t\mathbf{z}_B for t∈[0,1]t \in [0, 1]. Because the latent space is a well-calibrated Gaussian, linear paths in latent space produce smooth, meaningful transitions in image space — smiles fade in, glasses appear, hair color changes gradually.

Open in Lab
Explore latent space interpolation. Move the slider to travel between two encoded points — notice how the transition is smooth and semantically meaningful.
The demo wakes as you arrive…

Attribute manipulation: algebra in latent space

Because Glow's encoder maps images to a structured Gaussian latent space, we can do arithmetic with semantic meaning. For a binary attribute like "smiling":

  1. Encode all smiling training images and compute their average z+\mathbf{z}_{+}.
  2. Encode all non-smiling images and compute z−\mathbf{z}_{-}.
  3. The attribute direction is Δz=z+−z−\Delta\mathbf{z} = \mathbf{z}_{+} - \mathbf{z}_{-}.

To make any face smile: encode it as z\mathbf{z}, add α⋅Δz\alpha \cdot \Delta\mathbf{z}, and decode. Varying α\alpha smoothly controls how much the person smiles. This works because the flow's exact invertibility ensures that latent vectors near each other decode to images near each other — a property GANs cannot guarantee.

Open in Lab
Slide the attribute vector to see how adding or subtracting a learned direction changes the reconstruction smoothly.
The demo wakes as you arrive…

Results: exact likelihood meets realistic synthesis

Glow achieved state-of-the-art log-likelihood on CIFAR-10 (3.35 bits/dim) and ImageNet 32×32 and 64×64 among flow-based models. But the qualitative results were arguably more striking: 256×256 face images that were, at the time, the most realistic ever produced by a non-adversarial model trained purely on maximum likelihood.

The architecture used for CelebA-HQ faces was substantial: L=6L = 6 levels, K=32K = 32 flow steps per level, approximately 200M parameters, and around 600 convolution layers. Training on 8 GPUs took about 2 weeks. This large capacity was necessary for modeling the long-range dependencies in high-resolution images.

Limitations and trade-offs

Normalizing flows come with inherent constraints. Every transformation must be invertible and must have a tractable Jacobian determinant — this significantly restricts the architectural design space. The invertibility requirement means that the latent space has the same dimensionality as the data space: a 256×256×3 image maps to a latent vector of 256×256×3 entries. This is much larger than a 's or 's compact latent code and raises questions about whether the is truly compressed.

Computationally, Glow is expensive. Generating a single 256×256 image requires running ~600 convolution layers in reverse. The model also needed about 200M parameters to achieve good results, substantially more than contemporary GANs of similar image quality. While GANs and later diffusion models surpassed Glow in image quality, the flow paradigm's guarantee of exact likelihood remains its unique strength.

Legacy: from Glow to modern flows

Glow demonstrated that the normalizing flow paradigm could scale to high-resolution image generation — a result that reinvigorated the entire field. Its invertible 1×1 convolution became a standard building block in virtually all subsequent flow architectures, from Flow++ to WaveGlow (for audio synthesis).

The paper also crystallized an important philosophical point: a model optimized with the simple, principled objective of maximum likelihood can produce compelling samples without adversarial training. This insight influenced the later development of continuous normalizing flows (Neural ODEs) and flow matching methods, which inherit the exact-likelihood and invertibility guarantees while using continuous-time dynamics instead of discrete steps.

  1. 2014

    NICE (Dinh et al.)

    Introduced additive coupling layers and the idea of composing simple invertible functions. First demonstration that flows can model image distributions.

  2. 2016

    RealNVP (Dinh et al.)

    Upgraded from additive to affine coupling, added multi-scale architecture with squeeze-and-split. Made flows practical for natural images.

  3. 2018

    Glow (Kingma & Dhariwal)

    Replaced fixed permutations with learned invertible 1×1 convolutions and introduced actnorm. First flow model to generate realistic 256×256 faces with exact likelihood.

  4. 2018

    Neural ODEs (Chen et al.)

    Generalized discrete flow steps to continuous-time dynamics. Continuous normalizing flows use an ODE solver instead of a finite chain of transformations.

  5. 2019

    Flow++ (Ho et al.)

    Improved on Glow with variational dequantization and logistic coupling, pushing flow-based log-likelihoods closer to autoregressive models.

  6. 2022

    Flow Matching (Lipman et al.)

    A simpler training framework for continuous normalizing flows. Regresses directly on the velocity field, avoiding expensive ODE solves during training. Powers modern diffusion-adjacent methods.

Glow sits at a pivotal moment in generative modeling: the point where normalizing flows proved they could compete on image quality, not just on mathematical elegance. The specific architectural tricks — actnorm, 1×1 convolutions, multi-scale splitting — became standard vocabulary in the field. And the demonstration that maximum likelihood alone suffices for realistic synthesis planted a seed that flowered into the flow matching revolution.

CitationKingma, D. P. & Dhariwal, P.. Glow: Generative Flow with Invertible 1×1 Convolutions. NeurIPS, 2018.

Terms in this paper