Generative Models2022advanced11 min read

High-Resolution Image Synthesis with Latent Diffusion Models

توليد صور عالية الدقة باستخدام نماذج الانتشار الكامنة

Rombach, R. · Blattmann, A. · Lorenz, D. · Esser, P. · Ommer, B. — CVPR

The problem

By 2021, diffusion models had matched or beaten GANs on image quality — but at a devastating computational cost. a powerful directly in pixel space consumed hundreds of GPU days, and generating a single image required hundreds of sequential steps at full resolution. This made diffusion models impractical for high-resolution synthesis and inaccessible to most researchers. The question was: can we keep diffusion's quality and flexibility while slashing its compute by orders of magnitude?

The contribution

Latent Diffusion Models (LDMs): a two-stage approach that decouples from generative learning. Stage 1 trains an that compresses images into a compact — shrinking a 512×512×3 image to 64×64×4, a 48× compression. Stage 2 trains a diffusion model (a with and ) entirely in that latent space. Cross-attention layers accept from text encoders, semantic maps, or bounding boxes — making LDMs flexible general-purpose generators. The result: state-of-the-art and competitive unconditional generation, semantic synthesis, and — at a fraction of pixel-space compute. This architecture became Stable Diffusion.

The impact

LDMs democratized high-quality image generation. Stable Diffusion — built directly on this paper — became the first open-source, high-resolution model that anyone could run on a consumer GPU. It spawned an entire ecosystem: ControlNet, LoRA fine-tuning, img2img, SDXL, and video generation (Stable Video Diffusion). By proving that diffusion can work in latent space, LDMs also influenced audio (Stable Audio), 3D (Zero-1-to-3), and molecular generation. The paper's architectural pattern — autoencoder + latent diffusion + cross-attention conditioning — is now the standard template for controllable generative models.

Imagine you're an artist who paints by removing splotches of random paint from a canvas until a masterpiece emerges. That's a diffusion model. The problem? Working on a full-size wall mural — every brushstroke at full scale takes forever and costs a fortune in paint.

LDMs give the artist a sketchpad. First, a clever camera (the autoencoder) photographs the mural and shrinks it to a pocket-sized sketch that keeps all the important shapes and colors but discards imperceptible grain. The artist now removes splotches on the tiny sketch — 48 times less surface to work on. When the sketch is done, a projector (the ) blows it back up to mural size, filling in the fine grain automatically. Same masterpiece, a fraction of the effort.

The problem: diffusion in pixel space is brutally expensive

Diffusion models generate images by learning to reverse a -adding process. Starting from pure noise, a iteratively removes a little noise at each step until a clean image emerges. By 2021, this approach achieved state-of-the-art results — but with a painful price tag:

  • Training cost. The denoising network must process the full-resolution image at every step. For a 512×512 RGB image, that's 786,432 values per step, hundreds of steps per sample, millions of samples. Training consumed hundreds of GPU days.

  • cost. Generating one image required 250–1000 sequential denoising steps at full resolution. No parallelism across steps — each step depends on the previous one.

  • Wasted effort. Most of the computation went into modeling imperceptible high-frequency details — the exact grain of a texture, the pixel-level noise pattern — not the semantic content that makes an image look like a cat, a sunset, or a face.

Open in Lab
Compare pixel-space diffusion (top) vs latent-space diffusion (bottom). Notice the dramatic difference in computation.
The demo wakes as you arrive…

The insight: separate compression from generation

The paper's central insight draws from a fact well known in information theory: most of an image's bits encode imperceptible details — high-frequency textures, sub-pixel noise, compression artifacts. JPEG exploits this by discarding 80%+ of the data with negligible visible loss. LDMs take this further:

Stage 1 — Perceptual compression. Train an autoencoder (based on VQGAN) that learns to compress images into a compact latent space. A 512×512×3 image becomes a 64×64×4 — 48× fewer values. The autoencoder is trained once with a (so it preserves what humans see) and an (so the reconstructions look sharp).

Stage 2 — via diffusion. Train a diffusion model entirely within this latent space. The diffusion model only needs to learn the semantic structure of images — objects, layouts, styles — because the autoencoder already handles pixel-level fidelity.

This separation is powerful: the autoencoder's compression rate can be tuned independently, and the diffusion model operates on a manageable representation where every dimension carries semantic meaning.

Open in Lab
Drag the downsampling factor and watch how compression affects quality and cost.
The demo wakes as you arrive…

Stage 1: the autoencoder — learning to compress

The autoencoder is the foundation of the entire system. It has two halves:

  • The E\mathcal{E} takes an image x∈RH×W×3x \in \mathbb{R}^{H \times W \times 3} and maps it to a latent representation z=E(x)∈Rh×w×cz = \mathcal{E}(x) \in \mathbb{R}^{h \times w \times c}, where h=H/fh = H/f, w=W/fw = W/f, and cc is the number of latent channels (typically 4).

  • The decoder D\mathcal{D} reconstructs the image: x~=D(z)≈x\tilde{x} = \mathcal{D}(z) \approx x.

Training uses a combination of losses to ensure the latent space is useful for downstream diffusion: a perceptual loss (comparing high-level features via a pretrained VGG network, so the reconstruction looks right to humans), a patch-based adversarial loss (a ensures the output looks realistic at the patch level), and a mild or (to prevent the latent space from blowing up to arbitrarily high variance — keeping it well-behaved for the diffusion model).

z=E(x),x~=D(z),z∈Rh×w×cz = \mathcal{E}(x), \quad \tilde{x} = \mathcal{D}(z), \quad z \in \mathbb{R}^{h \times w \times c}
Autoencoder encoding and decoding — 𝓔 compresses image x to latent z · 𝓓 reconstructs x̃ from z · dimensions shrink from H×W×3 to h×w×c with h = H/f
Open in Lab
Watch a 512×512 image get compressed to a 64×64 latent, then reconstructed. The difference is nearly invisible.
The demo wakes as you arrive…

Stage 2: diffusion in the latent space

With the autoencoder trained and frozen, the diffusion model operates entirely on latent representations. The process follows the standard DDPM recipe, but in latent space:

(adding noise). Given a clean latent z0=E(x)z_0 = \mathcal{E}(x), gradually add Gaussian noise over TT timesteps to produce increasingly noisy versions z1,z2,…,zTz_1, z_2, \ldots, z_T, where zTz_T is nearly pure noise.

(denoising). A neural network ϵθ\epsilon_\theta learns to predict the noise ϵ\epsilon added at each step. Starting from random noise zTz_T, it removes noise step by step until it recovers a clean latent z0z_0. The final image is obtained by passing z0z_0 through the decoder: x~=D(z0)\tilde{x} = \mathcal{D}(z_0).

The denoising network is a U-Net — an encoder-decoder architecture with skip connections that's perfectly suited for this task. It processes spatial features at multiple resolutions, and the skip connections preserve fine-grained detail across the .

LLDM=EE(x), ϵ∼N(0,1), t[∥ϵ−ϵθ(zt,t)∥22]L_{LDM} = \mathbb{E}_{\mathcal{E}(x),\, \epsilon \sim \mathcal{N}(0,1),\, t} \left[ \| \epsilon - \epsilon_\theta(z_t, t) \|_2^2 \right]
LDM training objective — predict the noise — ε = the noise actually added · εθ(zₜ, t) = the network's noise prediction at timestep t · the model learns to predict exactly what noise was added, at every noise level
Open in Lab
Step through the denoising process — watch noise transform into structure in latent space.
The demo wakes as you arrive…

The flexibility engine: cross-attention conditioning

One of LDMs' most important contributions is a general-purpose conditioning mechanism. To guide the diffusion model with text, semantic maps, or any other signal, the paper introduces cross-attention layers into the U-Net backbone.

Here is the intuition: imagine the U-Net's intermediate features as a room full of image patches, each trying to decide what it should depict. The conditioning signal (say, the text "a castle on a hill at sunset") enters the room as a set of advisors. Each image patch queries the advisors: "what should I look like?" The advisors (text tokens) respond based on how relevant they are to that patch's spatial position and current content.

Formally, a domain-specific encoder τθ\tau_\theta maps the conditioning input yy (text, layout, ) into an intermediate representation τθ(y)\tau_\theta(y). Then, at selected layers of the U-Net, cross-attention is computed: the U-Net features produce queries, while the conditioning produces keys and values.

Attention(Q,K,V)=softmax ⁣(QK⊤d)Vwhere Q=WQ(i)⋅φi(zt),  K=WK(i)⋅τθ(y),  V=WV(i)⋅τθ(y)\text{Attention}(Q, K, V) = \text{softmax}\!\left(\frac{Q K^\top}{\sqrt{d}}\right) V \quad \text{where } Q = W_Q^{(i)} \cdot \varphi_i(z_t),\; K = W_K^{(i)} \cdot \tau_\theta(y),\; V = W_V^{(i)} \cdot \tau_\theta(y)
Cross-attention — how text guides the image — Q comes from U-Net features (the image asks) · K and V come from the text encoder (the text answers) · each spatial location learns which text tokens matter to it
Open in Lab
See how text tokens attend to different spatial regions of the image during generation.
The demo wakes as you arrive…

The full LDM architecture

Putting it all together, the LDM pipeline has three main components:

  1. Autoencoder (trained once, then frozen). The encoder E\mathcal{E} compresses images to latent space; the decoder D\mathcal{D} reconstructs them. Trained with perceptual + adversarial + regularization losses.

  2. U-Net denoiser ϵθ\epsilon_\theta (the diffusion model). Operates entirely in latent space. Its backbone is a U-Net augmented with self-attention layers (for spatial reasoning within the image) and cross-attention layers (for conditioning on external signals). The timestep tt is injected via sinusoidal embeddings.

  3. Conditioning encoder τθ\tau_\theta. A domain-specific encoder that translates the conditioning input into a sequence of vectors. For text, this is a Transformer language model (BERT or CLIP). For spatial conditions like semantic maps, it can be a simple convolutional encoder.

During training, an image xx is encoded to z0z_0, noise is added to get ztz_t, and the U-Net predicts the noise given ztz_t, tt, and the conditioning τθ(y)\tau_\theta(y). During inference, you sample zTz_T from pure noise and run the reverse process to get z0z_0, then decode with D(z0)\mathcal{D}(z_0).

Open in Lab
Click any component to see what it does in the LDM pipeline.
The demo wakes as you arrive…

Results: same quality, fraction of the cost

LDMs achieved state-of-the-art or highly competitive results across multiple tasks:

  • Unconditional image generation on CelebA-HQ (256×256) and LSUN (256×256) — scores matching or beating pixel-space diffusion and GANs, with significantly less compute.

  • Image inpainting — new state of the art, outperforming prior methods on both fidelity and semantic coherence of filled regions.

  • Text-to-image synthesis — competitive with DALL·E and GLIDE while requiring only a single consumer GPU for inference (vs. large GPU clusters).

  • Super-resolution — 4× upscaling with strong perceptual quality.

  • Semantic scene synthesis — generating photorealistic images from semantic layouts.

The efficiency gains were dramatic: LDM-4 (f=4) achieved similar quality to pixel-space models at roughly 2.7× faster training. LDM-8 (f=8) was even faster with negligible quality loss. Inference on a single GPU became practical — something unthinkable for pixel-space diffusion at these resolutions.

The same idea in code

LDM training and inference loop, simplifiedpython

Simplified to show the idea — not the real implementation.

import numpy as np

# ── Stage 1: Autoencoder (train once, then freeze) ──────────────
def encode(image, encoder):
    """Compress 512×512×3 image to 64×64×4 latent."""
    return encoder(image)   # (H, W, 3) → (H/f, W/f, c)

def decode(z, decoder):
    """Reconstruct full image from latent."""
    return decoder(z)       # (H/f, W/f, c) → (H, W, 3)

# ── Stage 2: Diffusion in latent space ──────────────────────────
def add_noise(z0, t, noise):
    """Forward process: add noise at timestep t."""
    alpha = noise_schedule(t)
    return np.sqrt(alpha) * z0 + np.sqrt(1 - alpha) * noise

def train_step(image, text, encoder, unet, text_encoder):
    """One training step of the latent diffusion model."""
    z0 = encode(image, encoder)               # compress to latent
    t = np.random.randint(0, T)                # random timestep
    noise = np.random.randn(*z0.shape)         # sample noise
    zt = add_noise(z0, t, noise)               # noisy latent

    cond = text_encoder(text)                  # encode text condition
    predicted_noise = unet(zt, t, cond)        # predict the noise
    loss = np.mean((noise - predicted_noise)**2)  # simple MSE
    return loss

def generate(text, unet, text_encoder, decoder, steps=50):
    """Generate an image from text."""
    zt = np.random.randn(64, 64, 4)            # start from noise
    cond = text_encoder(text)                   # encode prompt
    for t in reversed(range(steps)):
        predicted_noise = unet(zt, t, cond)     # predict noise
        zt = denoise_step(zt, predicted_noise, t) # remove one step
    return decode(zt, decoder)                  # latent → pixels

# That's the core idea. Stable Diffusion is this loop at scale.

What LDMs unlocked

  1. 2022

    Latent Diffusion Models (LDMs)

    Proved diffusion can work in compressed latent space with no quality loss. Slashed training and inference cost by orders of magnitude. Introduced cross-attention conditioning for flexible text-to-image generation.

  2. 2022

    Stable Diffusion

    The open-source implementation of LDMs. First high-quality text-to-image model that ran on consumer GPUs. Millions of users within weeks of release.

  3. 2023

    ControlNet

    Added precise spatial control to LDMs — generate images guided by edges, depth maps, poses, or segmentation masks. Built directly on the U-Net + cross-attention backbone.

  4. 2023

    SDXL

    Scaled LDMs to higher resolution (1024×1024) with a larger U-Net and a refinement model. Significantly improved text rendering and compositional understanding.

  5. 2023

    Stable Video Diffusion

    Extended LDMs to video by adding temporal attention layers, generating coherent video clips from text or image prompts.

  6. 2024

    Stable Diffusion 3 / FLUX

    Replaced the U-Net backbone with a Diffusion Transformer (DiT), but kept the latent-space paradigm and cross-attention conditioning from the original LDM paper.

LDMs' lasting contribution is architectural: the insight that perceptual compression and generative modeling are best solved separately. Train an autoencoder once to handle pixels, then let the focus purely on semantics in a compact, well-structured latent space. This pattern now underlies almost every state-of-the-art image, video, and audio generation system.

CitationRombach, Blattmann, Lorenz, Esser, Ommer. High-Resolution Image Synthesis with Latent Diffusion Models. CVPR, 2022.

Terms in this paper