Generative Models2016intermediate13 min read
Pixel Recurrent Neural Networks
الشبكات العصبية التكرارية على مستوى البكسل
van den Oord, A. · Kalchbrenner, N. · Kavukcuoglu, K. — ICML
The problem
Modeling the distribution of natural images is a landmark problem in unsupervised learning. Before , generative models either sacrificed expressiveness for tractability (like mixtures of Gaussians) or gave up tractable entirely (like GANs). No existing model could simultaneously be expressive enough to capture the richness of natural images, provide exact and tractable likelihood computation, and scale to large datasets like ImageNet.
The contribution
PixelRNN models the joint distribution of pixels as a product of conditional distributions using the , predicting each sequentially in order. The paper introduces two novel 2D architectures — and — that capture spatial dependencies across the image. It also proposes , a faster purely convolutional alternative using masked convolutions. The models use a discrete 256-way for each color channel and deep residual connections across up to 12 recurrent layers, achieving state-of-the-art on MNIST, CIFAR-10, and ImageNet.
The impact
PixelRNN established autoregressive models as a serious contender in generative image modeling. It directly inspired WaveNet (the same autoregressive principle applied to audio), VQ- (which uses PixelCNN as its prior), and the entire family of improved PixelCNN variants. The idea that images can be generated one pixel at a time with opened a new paradigm distinct from GANs and VAEs.
Imagine a portrait painter who works in strict order: top-left corner first, then one dot at a time, left to right, row by row. Before placing each dot of paint, the painter looks at everything already on the canvas and thinks: "What color fits here best?"
An untrained painter might guess randomly. But a painter who has studied thousands of photographs learns that sky pixels tend to follow sky pixels, that edges continue in smooth curves, and that shadows fall in predictable directions.
PixelRNN is that trained painter — a that has memorized the statistical patterns of natural images so well that, dot by dot, it can paint new images that look startlingly real.
The challenge: how do you model an image?
A 32×32 color image has 3,072 values (32 × 32 × 3 channels). Each pixel channel takes one of 256 integer values (0–255). The number of possible images is — a number so vast it dwarfs the atoms in the observable universe.
A must learn which of these astronomical combinations look like real photographs. Before PixelRNN, the main approaches each had a critical weakness:
- Variational Autoencoders (VAEs) provide tractable likelihood bounds but tend to produce blurry samples because the averages over latent uncertainty.
- GANs produce sharp samples but give no likelihood at all — you cannot measure how "probable" an image is under the model.
- Mixture models are tractable but not expressive enough for the complexity of natural images.
The authors asked: can we model the exact joint distribution of all pixels, tractably, at scale?
The core idea: images as sequences
The key insight is deceptively simple: use the chain rule of probability to decompose the joint distribution of all pixels into a product of conditional distributions. Flatten the 2D image into a 1D sequence by scanning left to right, top to bottom (raster scan order), then model each pixel as a conditional distribution given all previous pixels.
This turns image generation into a sequence prediction problem — exactly the kind of problem LSTMs are designed for. And crucially, this decomposition is exact: no approximation, no lower bound. The product of all the conditional probabilities is the joint probability.
Think of it like a game of dominoes: the full pattern emerges from each piece fitting against the ones already placed. Pixel 1 is unconditional — the model picks a color freely. Pixel 2 is conditioned on pixel 1. Pixel 100 is conditioned on all 99 previous pixels. By the time the model reaches the last pixel, every prior pixel has had its say.
Color channels: R, then G, then B
Each pixel is not a single number but three: Red, Green, and Blue. The model decomposes each pixel's distribution further into a chain of three sub-pixel predictions:
The red channel is conditioned on all previous pixels. The green channel uses all previous pixels plus the current red value. The blue channel uses all previous pixels plus the current red and green values.
Three architectures, one principle
The paper introduces three architectures that all implement the same autoregressive principle but differ in how they capture the context of previous pixels:
Row LSTM processes the image row by row. For each pixel, it uses a 1D along the current row (masked to see only previous pixels) as the input-to-state component, and a recurrent connection from the previous row as the state-to-state component. This gives it a triangular — it sees all pixels above but has a blind spot in the upper-right.
Diagonal BiLSTM processes the image along diagonals by skewing the input map — shifting each row by one position relative to the previous row. Two directional LSTMs scan from opposite corners, and their outputs combine. This gives the ideal receptive field: every previous pixel is visible at every layer, with no blind spots.
PixelCNN replaces recurrence entirely with stacked masked convolutions. It is significantly faster to train because all pixels can be computed in parallel, but its bounded receptive field means it captures less context than the LSTM variants.
Row LSTM: fast rows, triangular vision
The Row LSTM is a unidirectional layer that processes the image row by row from top to bottom. For each row, the input-to-state computation is done for all positions at once using a k×1 — this is the parallelism trick. Then the state-to-state recurrence propagates information across the row.
Think of it as a typewriter that reads one line at a time: it can see the full page above the current line, but for the current line it types left to right, each character informed by the ones before it. The k×1 convolution lets it peek at a few pixels in the row above, giving it local spatial context.
Diagonal BiLSTM: complete context, diagonal sweep
The Diagonal BiLSTM achieves the ideal: at every layer, every pixel sees all previous pixels — no blind spots. The trick is a clever geometric transformation called skewing.
The input map is skewed by offsetting each row by one position relative to the previous row. This turns the rectangular grid into a parallelogram. Now, processing column by column is equivalent to processing the original image along its diagonals. Two directional LSTMs scan from opposite corners — one starts from the top-left, the other from the top-right — and their features are combined.
After the computation, the output is un-skewed back to the original rectangular shape. The beauty is that the 2×1 column-wise convolution in the skewed space processes one pixel from each diagonal at a time, maintaining the complete dependency field.
Masked convolutions: the autoregressive enforcer
In a standard convolution, the kernel sees pixels in all directions — including pixels the model hasn't generated yet. This violates the autoregressive property: a pixel must only depend on previous pixels.
The solution is masking: zero out the kernel weights that would connect to the current pixel or any future pixel. Two types of masks are used:
- Mask A (applied to the first layer): blocks the center pixel and everything after it. The model cannot see its own current value.
- Mask B (applied to subsequent layers): allows the center pixel. After the first layer, features at position (i, j) already represent valid past context, so they can be included.
The mask zeroes out specific weights in the convolution filter after each update, ensuring the autoregressive constraint is never violated.
PixelCNN: trading context for speed
PixelCNN replaces recurrence with stacked masked convolutions. Because convolutions have no sequential dependency, all pixel features within a layer can be computed in parallel — making dramatically faster than Row LSTM or Diagonal BiLSTM.
The trade-off is the receptive field: each convolutional layer only grows the receptive field by the kernel size, so even after many layers, a PixelCNN may not see distant pixels. The LSTM variants, by contrast, can theoretically capture dependencies across the entire image through their recurrent state.
In the paper's experiments, Diagonal BiLSTM achieved the best log-likelihood (3.00 bits/dim on CIFAR-10), followed by Row LSTM (3.06), and PixelCNN came last (3.14) — precisely the order of their receptive field sizes. This confirmed that larger receptive fields directly improve generative quality.
Residual connections in deep recurrent networks
The PixelRNN models use up to 12 LSTM layers stacked on top of each other. Training such deep recurrent networks is notoriously difficult — gradients can vanish or explode across layers. The paper adopts residual connections between LSTM layers: the input to each layer is added back to its output before passing to the next layer.
The residual block works as follows: the input map has 2h features. The LSTM layer reduces this to h features (the gate outputs). A 1×1 convolution projects the LSTM output back to 2h features, and the original input is added. This creates a shortcut path that lets gradients flow directly through the network, enabling stable training of deep models.
Multi-Scale PixelRNN: coarse to fine
The PixelRNN generates images in a coarse-to-fine fashion. An unconditional PixelRNN first generates a small s×s image. Then a conditional PixelRNN takes this small image as context (upsampling it via deconvolutional layers) and generates a larger n×n image.
Think of it like sketching: first you rough out the broad composition at low resolution — where the sky meets the ground, where the subject sits. Then you zoom in and fill in the details, using the rough sketch as a guide.
This approach allows the model to capture global structure at the coarse scale (where long-range dependencies span fewer pixels) and fill in local details at the fine scale.
The autoregressive principle in code
Simplified to show the idea — not the real implementation.
import numpy as np
def create_mask_a(kernel_size):
"""Mask A: blocks center pixel and everything after it."""
mask = np.ones((kernel_size, kernel_size))
center = kernel_size // 2
# Block center row after center, and all rows below
mask[center, center:] = 0
mask[center+1:, :] = 0
return mask
def create_mask_b(kernel_size):
"""Mask B: allows center pixel, blocks everything after."""
mask = create_mask_a(kernel_size)
center = kernel_size // 2
mask[center, center] = 1 # allow the center pixel
return mask
def masked_conv(image, kernel, mask):
"""Apply a masked convolution — the core of PixelCNN."""
return convolve2d(image, kernel * mask, mode='same')
def pixel_probability(model, image, position):
"""
For pixel at `position`, return a 256-way probability vector.
The model only sees pixels before `position` in raster order.
"""
features = model.forward(image) # masked conv stack
logits = features[position] # 256 raw scores
return softmax(logits) # 256 probabilities
# Generation: sample one pixel at a time, left→right, top→bottom
# Training: all pixels computed in parallel (teacher forcing)Training vs generation: parallel learning, sequential painting
A crucial asymmetry exists between training and generation:
During training, the target image is fully known. Because masked convolutions prevent each position from seeing future pixels, all pixel predictions can be computed in a single forward pass — full parallelism. The loss is the sum of cross-entropy losses at every pixel position.
During generation, each pixel must be sampled before the next pixel can be predicted. For a 32×32 image, this means 3,072 sequential forward passes (32 × 32 × 3 channels). This makes generation slow — but training remains efficient.
This training-generation asymmetry is a defining characteristic of autoregressive models: they are easy to train (parallel) but slow to sample from (sequential).
Results and significance
The paper set new state-of-the-art log-likelihood scores on every benchmark tested:
- MNIST (binary): 79.20 nats (Diagonal BiLSTM), compared to 84.55 for the previous best (DRAW).
- CIFAR-10: 3.00 bits/dim (Diagonal BiLSTM), compared to 3.24 for the previous best (NICE).
- ImageNet 32×32: First reported likelihood benchmarks on ImageNet for generative models.
The ranking across architectures was consistent: Diagonal BiLSTM > Row LSTM > PixelCNN. This directly correlated with receptive field size, confirming that capturing more context improves generation quality. The generated samples appeared crisp, varied, and globally coherent — a notable improvement over the blurriness typical of VAEs at the time.
Why it mattered — the autoregressive revolution
PixelRNN demonstrated that exact likelihood and high-quality generation are not mutually exclusive. This was a paradigm shift. Before this paper, the generative modeling landscape was dominated by GANs (sharp but no likelihood) and VAEs (likelihood bound but blurry). PixelRNN carved out a third path: autoregressive models with tractable, exact likelihoods.
The ideas in this paper propagated far beyond images. The same autoregressive decomposition and masked architecture were adapted for:
- Audio: WaveNet applied the same pixel-by-pixel principle to audio waveforms, sample by sample, producing speech of unprecedented quality.
- Discrete latent codes: VQ-VAE used PixelCNN as a powerful prior over its discrete codebook, combining the strengths of autoencoders and autoregressive models.
- Language models: the raster-scan autoregressive principle is conceptually identical to left-to-right language modeling, linking this work to the GPT family.
2016
PixelRNN & PixelCNN
Established autoregressive image modeling with exact likelihood. Introduced masked convolutions and 2D LSTMs for pixel-by-pixel generation.
2016
WaveNet
Applied the PixelRNN autoregressive principle to raw audio waveforms, generating speech sample by sample with remarkable naturalness.
2016
Gated PixelCNN
Added gated activations and fixed the blind spot in PixelCNN's receptive field, matching PixelRNN quality at CNN speed.
2017
PixelCNN++
Replaced the 256-way softmax with a mixture of logistics, capturing pixel value continuity. Improved both efficiency and quality.
2017
VQ-VAE
Combined vector quantization with an autoencoder, using PixelCNN as the prior over discrete latent codes — marrying autoencoders and autoregressive models.
The PixelRNN algorithm step by step
Citationvan den Oord, Kalchbrenner, Kavukcuoglu. Pixel Recurrent Neural Networks. ICML, 2016.
Terms in this paper
- Autoregressive Modelالنموذج التوليدي التراجعي
- Likelihoodالأرجحية
- Pixelالبكسل (عنصر الصورة)
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Masked Convolutionالالتفاف المُقنَّع
- Residual Connectionالوصلة التجاوزية
- Softmaxسوفت ماكس
- Raster Scanالمسح السطري
- Receptive Fieldالحقل الاستقبالي للعصبون
- Generative Modelالنموذج التوليدي