Generative Models2022advanced11 min read
High-Resolution Image Synthesis with Latent Diffusion Models
توليد صور عالية الدقة باستخدام نماذج الانتشار الكامنة
Rombach, R. · Blattmann, A. · Lorenz, D. · Esser, P. · Ommer, B. — CVPR
The problem
By 2021, diffusion models had matched or beaten GANs on image quality — but at a devastating computational cost. a powerful directly in pixel space consumed hundreds of GPU days, and generating a single image required hundreds of sequential steps at full resolution. This made diffusion models impractical for high-resolution synthesis and inaccessible to most researchers. The question was: can we keep diffusion's quality and flexibility while slashing its compute by orders of magnitude?
The contribution
Latent Diffusion Models (LDMs): a two-stage approach that decouples from generative learning. Stage 1 trains an that compresses images into a compact — shrinking a 512×512×3 image to 64×64×4, a 48× compression. Stage 2 trains a diffusion model (a with and ) entirely in that latent space. Cross-attention layers accept from text encoders, semantic maps, or bounding boxes — making LDMs flexible general-purpose generators. The result: state-of-the-art and competitive unconditional generation, semantic synthesis, and — at a fraction of pixel-space compute. This architecture became Stable Diffusion.
The impact
LDMs democratized high-quality image generation. Stable Diffusion — built directly on this paper — became the first open-source, high-resolution model that anyone could run on a consumer GPU. It spawned an entire ecosystem: ControlNet, LoRA fine-tuning, img2img, SDXL, and video generation (Stable Video Diffusion). By proving that diffusion can work in latent space, LDMs also influenced audio (Stable Audio), 3D (Zero-1-to-3), and molecular generation. The paper's architectural pattern — autoencoder + latent diffusion + cross-attention conditioning — is now the standard template for controllable generative models.
Imagine you're an artist who paints by removing splotches of random paint from a canvas until a masterpiece emerges. That's a diffusion model. The problem? Working on a full-size wall mural — every brushstroke at full scale takes forever and costs a fortune in paint.
LDMs give the artist a sketchpad. First, a clever camera (the autoencoder) photographs the mural and shrinks it to a pocket-sized sketch that keeps all the important shapes and colors but discards imperceptible grain. The artist now removes splotches on the tiny sketch — 48 times less surface to work on. When the sketch is done, a projector (the ) blows it back up to mural size, filling in the fine grain automatically. Same masterpiece, a fraction of the effort.
The problem: diffusion in pixel space is brutally expensive
Diffusion models generate images by learning to reverse a -adding process. Starting from pure noise, a iteratively removes a little noise at each step until a clean image emerges. By 2021, this approach achieved state-of-the-art results — but with a painful price tag:
-
Training cost. The denoising network must process the full-resolution image at every step. For a 512×512 RGB image, that's 786,432 values per step, hundreds of steps per sample, millions of samples. Training consumed hundreds of GPU days.
-
cost. Generating one image required 250–1000 sequential denoising steps at full resolution. No parallelism across steps — each step depends on the previous one.
-
Wasted effort. Most of the computation went into modeling imperceptible high-frequency details — the exact grain of a texture, the pixel-level noise pattern — not the semantic content that makes an image look like a cat, a sunset, or a face.
The insight: separate compression from generation
The paper's central insight draws from a fact well known in information theory: most of an image's bits encode imperceptible details — high-frequency textures, sub-pixel noise, compression artifacts. JPEG exploits this by discarding 80%+ of the data with negligible visible loss. LDMs take this further:
Stage 1 — Perceptual compression. Train an autoencoder (based on VQGAN) that learns to compress images into a compact latent space. A 512×512×3 image becomes a 64×64×4 — 48× fewer values. The autoencoder is trained once with a (so it preserves what humans see) and an (so the reconstructions look sharp).
Stage 2 — via diffusion. Train a diffusion model entirely within this latent space. The diffusion model only needs to learn the semantic structure of images — objects, layouts, styles — because the autoencoder already handles pixel-level fidelity.
This separation is powerful: the autoencoder's compression rate can be tuned independently, and the diffusion model operates on a manageable representation where every dimension carries semantic meaning.
Stage 1: the autoencoder — learning to compress
The autoencoder is the foundation of the entire system. It has two halves:
-
The takes an image and maps it to a latent representation , where , , and is the number of latent channels (typically 4).
-
The decoder reconstructs the image: .
Training uses a combination of losses to ensure the latent space is useful for downstream diffusion: a perceptual loss (comparing high-level features via a pretrained VGG network, so the reconstruction looks right to humans), a patch-based adversarial loss (a ensures the output looks realistic at the patch level), and a mild or (to prevent the latent space from blowing up to arbitrarily high variance — keeping it well-behaved for the diffusion model).
Stage 2: diffusion in the latent space
With the autoencoder trained and frozen, the diffusion model operates entirely on latent representations. The process follows the standard DDPM recipe, but in latent space:
(adding noise). Given a clean latent , gradually add Gaussian noise over timesteps to produce increasingly noisy versions , where is nearly pure noise.
(denoising). A neural network learns to predict the noise added at each step. Starting from random noise , it removes noise step by step until it recovers a clean latent . The final image is obtained by passing through the decoder: .
The denoising network is a U-Net — an encoder-decoder architecture with skip connections that's perfectly suited for this task. It processes spatial features at multiple resolutions, and the skip connections preserve fine-grained detail across the .
The flexibility engine: cross-attention conditioning
One of LDMs' most important contributions is a general-purpose conditioning mechanism. To guide the diffusion model with text, semantic maps, or any other signal, the paper introduces cross-attention layers into the U-Net backbone.
Here is the intuition: imagine the U-Net's intermediate features as a room full of image patches, each trying to decide what it should depict. The conditioning signal (say, the text "a castle on a hill at sunset") enters the room as a set of advisors. Each image patch queries the advisors: "what should I look like?" The advisors (text tokens) respond based on how relevant they are to that patch's spatial position and current content.
Formally, a domain-specific encoder maps the conditioning input (text, layout, ) into an intermediate representation . Then, at selected layers of the U-Net, cross-attention is computed: the U-Net features produce queries, while the conditioning produces keys and values.
The full LDM architecture
Putting it all together, the LDM pipeline has three main components:
-
Autoencoder (trained once, then frozen). The encoder compresses images to latent space; the decoder reconstructs them. Trained with perceptual + adversarial + regularization losses.
-
U-Net denoiser (the diffusion model). Operates entirely in latent space. Its backbone is a U-Net augmented with self-attention layers (for spatial reasoning within the image) and cross-attention layers (for conditioning on external signals). The timestep is injected via sinusoidal embeddings.
-
Conditioning encoder . A domain-specific encoder that translates the conditioning input into a sequence of vectors. For text, this is a Transformer language model (BERT or CLIP). For spatial conditions like semantic maps, it can be a simple convolutional encoder.
During training, an image is encoded to , noise is added to get , and the U-Net predicts the noise given , , and the conditioning . During inference, you sample from pure noise and run the reverse process to get , then decode with .
Results: same quality, fraction of the cost
LDMs achieved state-of-the-art or highly competitive results across multiple tasks:
-
Unconditional image generation on CelebA-HQ (256×256) and LSUN (256×256) — scores matching or beating pixel-space diffusion and GANs, with significantly less compute.
-
Image inpainting — new state of the art, outperforming prior methods on both fidelity and semantic coherence of filled regions.
-
Text-to-image synthesis — competitive with DALL·E and GLIDE while requiring only a single consumer GPU for inference (vs. large GPU clusters).
-
Super-resolution — 4× upscaling with strong perceptual quality.
-
Semantic scene synthesis — generating photorealistic images from semantic layouts.
The efficiency gains were dramatic: LDM-4 (f=4) achieved similar quality to pixel-space models at roughly 2.7× faster training. LDM-8 (f=8) was even faster with negligible quality loss. Inference on a single GPU became practical — something unthinkable for pixel-space diffusion at these resolutions.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
# ── Stage 1: Autoencoder (train once, then freeze) ──────────────
def encode(image, encoder):
"""Compress 512×512×3 image to 64×64×4 latent."""
return encoder(image) # (H, W, 3) → (H/f, W/f, c)
def decode(z, decoder):
"""Reconstruct full image from latent."""
return decoder(z) # (H/f, W/f, c) → (H, W, 3)
# ── Stage 2: Diffusion in latent space ──────────────────────────
def add_noise(z0, t, noise):
"""Forward process: add noise at timestep t."""
alpha = noise_schedule(t)
return np.sqrt(alpha) * z0 + np.sqrt(1 - alpha) * noise
def train_step(image, text, encoder, unet, text_encoder):
"""One training step of the latent diffusion model."""
z0 = encode(image, encoder) # compress to latent
t = np.random.randint(0, T) # random timestep
noise = np.random.randn(*z0.shape) # sample noise
zt = add_noise(z0, t, noise) # noisy latent
cond = text_encoder(text) # encode text condition
predicted_noise = unet(zt, t, cond) # predict the noise
loss = np.mean((noise - predicted_noise)**2) # simple MSE
return loss
def generate(text, unet, text_encoder, decoder, steps=50):
"""Generate an image from text."""
zt = np.random.randn(64, 64, 4) # start from noise
cond = text_encoder(text) # encode prompt
for t in reversed(range(steps)):
predicted_noise = unet(zt, t, cond) # predict noise
zt = denoise_step(zt, predicted_noise, t) # remove one step
return decode(zt, decoder) # latent → pixels
# That's the core idea. Stable Diffusion is this loop at scale.What LDMs unlocked
2022
Latent Diffusion Models (LDMs)
Proved diffusion can work in compressed latent space with no quality loss. Slashed training and inference cost by orders of magnitude. Introduced cross-attention conditioning for flexible text-to-image generation.
2022
Stable Diffusion
The open-source implementation of LDMs. First high-quality text-to-image model that ran on consumer GPUs. Millions of users within weeks of release.
2023
ControlNet
Added precise spatial control to LDMs — generate images guided by edges, depth maps, poses, or segmentation masks. Built directly on the U-Net + cross-attention backbone.
2023
SDXL
Scaled LDMs to higher resolution (1024×1024) with a larger U-Net and a refinement model. Significantly improved text rendering and compositional understanding.
2023
Stable Video Diffusion
Extended LDMs to video by adding temporal attention layers, generating coherent video clips from text or image prompts.
2024
Stable Diffusion 3 / FLUX
Replaced the U-Net backbone with a Diffusion Transformer (DiT), but kept the latent-space paradigm and cross-attention conditioning from the original LDM paper.
LDMs' lasting contribution is architectural: the insight that perceptual compression and generative modeling are best solved separately. Train an autoencoder once to handle pixels, then let the focus purely on semantics in a compact, well-structured latent space. This pattern now underlies almost every state-of-the-art image, video, and audio generation system.
CitationRombach, Blattmann, Lorenz, Esser, Ommer. High-Resolution Image Synthesis with Latent Diffusion Models. CVPR, 2022.
Terms in this paper
- Latent Diffusion Modelنموذج الانتشار الكامن
- Perceptual Compressionالضغط الإدراكي
- Semantic Compressionالضغط الدلالي
- Downsampling Factorمعامل تقليص الأبعاد
- Classifier-Free Guidanceالتوجيه بلا مُصنِّف