Generative Models2021intermediate10 min read

GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

GLIDE: توليد صور واقعية وتحريرها باستخدام نماذج انتشار موجَّهة بالنص

Nichol, A. · Dhariwal, P. · Ramesh, A. · Shyam, P. · Mishkin, P. · McGrew, B. · Sutskever, I. · Chen, M. — ICML

The problem

By late 2021, generation had made exciting progress. DALL-E showed that autoregressive models could compose concepts from text, and diffusion models had surpassed GANs on unconditional image quality. But no system could do both: generate photorealistic images from arbitrary text prompts. DALL-E's images were often blurry, and diffusion models lacked text conditioning. The missing piece was a way to inject free-form text into a powerful and guide it to produce images that both look real and match the description.

The contribution

A 3.5 billion parameter text-conditional diffusion model that generates photorealistic images from text using . The paper compares two guidance strategies — guidance and classifier-free guidance — and shows through human evaluation that classifier-free guidance produces images preferred 87% of the time for photorealism over DALL-E. The model can also be fine-tuned for , enabling text-driven editing with realistic shadows, reflections, and style matching.

The impact

GLIDE proved that diffusion models could surpass autoregressive models at text-to-image synthesis, establishing the diffusion paradigm that would dominate generative AI. It directly inspired DALL-E 2 (which combined CLIP with diffusion), Imagen (which used large language model encoders with classifier-free guidance), and Stable Diffusion (which moved diffusion to latent space). Classifier-free guidance became the default steering mechanism for virtually all subsequent text-to-image systems.

Imagine a sculptor working with a block of marble that's been shattered into random dust. A client whispers a description: "a corgi wearing a birthday hat." The sculptor cannot see the final statue — but they have two inner voices. One voice says: "here's what any random image looks like at this stage." The other says: "here's what an image matching that description looks like at this stage." By always chiseling toward the described voice and away from the generic one, the sculptor gradually reveals a photorealistic corgi in a party hat — carved from pure noise.

GLIDE is that sculptor. The two voices are classifier-free guidance, and the marble dust is Gaussian noise.

The problem: realistic images or text control — pick one

Before GLIDE, two families of models had made exciting but separate progress. On one side, diffusion models (like those in "Diffusion Models Beat GANs") had achieved photorealistic image quality on ImageNet — producing faces, animals, and scenes that humans could barely distinguish from real photographs. But they were class-conditional: you could ask for "dog" or "car," not "a corgi wearing a red bowtie and a purple party hat."

On the other side, DALL-E showed that an trained on text-image pairs could compose complex scenes from arbitrary text. But the images were often blurry and lacked the photorealism of diffusion models. The best DALL-E samples required generating hundreds of candidates and using CLIP to rerank them — an expensive post-processing step.

The question was straightforward: could you marry diffusion's image quality with DALL-E's text flexibility?

Open in Lab
The generative landscape before GLIDE: diffusion models excelled at photorealism but lacked text control, while DALL-E had text flexibility but blurry outputs.
The demo wakes as you arrive…

Quick recap: how diffusion models generate images

A diffusion model works in two phases. The forward process gradually destroys a real image by adding Gaussian noise over TT timesteps until it becomes pure static. The reverse process learns to undo this destruction: a neural network ϵθ\epsilon_\theta predicts the noise that was added at each step, and subtracts it to recover the image.

Think of it as learning to unscramble: if you know exactly what noise was stirred in, you can stir it back out. The model trains by seeing many noised images and learning to predict the noise component ϵ\epsilon from the noisy version xtx_t.

Lsimple=Et,x0,ϵ[∥ϵ−ϵθ(xt,t)∥2]L_{\text{simple}} = \mathbb{E}_{t, x_0, \epsilon}\left[\|\epsilon - \epsilon_\theta(x_t, t)\|^2\right]
Simple diffusion training objective — The model ϵθ\epsilon_\theta is trained to predict the noise ϵ\epsilon added to a clean image x0x_0 at timestep tt. The loss is simply the mean squared error between the true noise and the predicted noise. By minimizing this, the model learns to denoise at every noise level.
Open in Lab
Watch the forward process add noise step by step, then see the reverse process denoise back to a clean image. Drag the slider to explore any timestep.
The demo wakes as you arrive…

The GLIDE architecture: injecting text into diffusion

GLIDE's core idea is simple: take a powerful diffusion model and give it text understanding. The architecture has three components working together:

1. A — a with 24 layers and width 2048, roughly 1.2 billion parameters. It converts a text caption into a sequence of KK token embeddings.

2. A visual diffusion model — based on the ADM (Ablated Diffusion Model) architecture, scaled to 512 channels and roughly 2.3 billion parameters. This is the workhorse that denoises images.

3. Two injection points — the text enters the visual model in two ways. First, the final replaces the class label (where a class-conditional model would use "dog" or "car"). Second, the full sequence of KK token embeddings is projected and concatenated to the attention context at every attention layer. This gives each layer of the U-Net access to the full text description via .

Together, the model has 3.5 billion parameters — the base model — plus a separate 1.5 billion parameter diffusion model that takes the 64×64 output and increases it to 256×256 resolution.

Open in Lab
Explore GLIDE's architecture: see how the text encoder feeds into the U-Net via two injection paths — final token embedding and cross-attention at every layer.
The demo wakes as you arrive…

Classifier-free guidance: the model steers itself

The most impactful contribution of GLIDE is demonstrating that classifier-free guidance dramatically outperforms CLIP guidance for text-to-image diffusion. But what is classifier-free guidance, and why does it work so well?

The intuition begins with a mental model. During sampling, the model produces two predictions at each step: one conditioned on the text caption cc, and one unconditional (as if no caption were provided, using an empty prompt ∅\varnothing). The difference between these two predictions tells the model: "this is the direction that makes the image more like what the text describes." By amplifying that difference, you push the image harder toward the text.

Think of it as navigation with two compasses. One compass points toward "any image" (the unconditional prediction). The other points toward "an image matching this description" (the conditional prediction). The vector from the first to the second is the "text direction." Multiplying that vector by a s>1s > 1 makes the model follow the text more aggressively — at the cost of some diversity.

ϵ^θ(xt∣c)=ϵθ(xt∣∅)+s⋅(ϵθ(xt∣c)−ϵθ(xt∣∅))\hat{\epsilon}_\theta(x_t | c) = \epsilon_\theta(x_t | \varnothing) + s \cdot (\epsilon_\theta(x_t | c) - \epsilon_\theta(x_t | \varnothing))
Classifier-free guidance formula — The guided noise prediction ϵ^\hat{\epsilon} extrapolates from the unconditional prediction ϵθ(xt∣∅)\epsilon_\theta(x_t|\varnothing) in the direction of the conditional prediction ϵθ(xt∣c)\epsilon_\theta(x_t|c). When s=1s = 1, this reduces to standard conditional sampling. When s>1s > 1, the model overshoots — amplifying text-relevant features at the cost of diversity. GLIDE uses s=3.0s = 3.0.
Open in Lab
Drag the guidance scale to see how classifier-free guidance steers the denoising direction. At s=1 the model samples normally; at higher scales it overshoots toward the text.
The demo wakes as you arrive…

CLIP guidance: the external judge

The alternative guidance approach uses CLIP as an external evaluator. CLIP has two encoders — one for images f(x)f(x) and one for text g(c)g(c) — and outputs a similarity score. Higher dot product f(x)⋅g(c)f(x) \cdot g(c) means the image better matches the caption.

To use CLIP for guidance, at each denoising step the model computes the gradient of the CLIP similarity with respect to the current noisy image xtx_t. This gradient tells the model: "adjusting the image in this direction will increase its similarity to the text." The denoising mean is then shifted by this gradient, scaled by a guidance coefficient.

One critical detail: standard CLIP models were trained on clean images, not noisy ones. The noisy intermediate images during diffusion sampling are out-of-distribution for standard CLIP. GLIDE addresses this by training a noised CLIP model that explicitly processes noisy images f(xt,t)f(x_t, t), making the gradients more accurate.

μ^θ(xt∣c)=μθ(xt∣c)+s⋅Σθ(xt∣c) ∇xt ⁣(f(xt)⋅g(c))\hat{\mu}_\theta(x_t | c) = \mu_\theta(x_t | c) + s \cdot \Sigma_\theta(x_t | c)\,\nabla_{x_t}\!\bigl(f(x_t) \cdot g(c)\bigr)
CLIP guidance formula — The denoising mean μθ\mu_\theta is shifted by the gradient of CLIP similarity, scaled by the variance Σθ\Sigma_\theta and guidance scale ss. This pushes each denoising step toward images that CLIP considers more similar to the caption.

The verdict: classifier-free guidance wins on all fronts

GLIDE ran a systematic comparison between the two guidance strategies. On automated metrics, CLIP guidance achieves higher CLIP scores — but this is misleading. The authors hypothesized that CLIP guidance finds adversarial examples for the CLIP evaluator: images that score high on CLIP similarity without actually looking better to humans.

The definitive answer came from human evaluation. Evaluators were shown pairs of images and asked to choose which was more photorealistic or better matched the caption. The results were decisive: classifier-free guidance won in both photorealism and caption similarity. When comparing GLIDE with classifier-free guidance against DALL-E (even with DALL-E using expensive CLIP reranking of 512 candidates), human evaluators preferred GLIDE 87% of the time for photorealism and 69% for caption similarity.

On MS-COCO, GLIDE achieved a zero-shot FID of 12.24 — competitive with models explicitly trained on COCO — using a model roughly 3× smaller than DALL-E (3.5B vs. 12B parameters).

Open in Lab
Compare classifier-free guidance vs CLIP guidance: see human evaluation results and how each method trades off diversity for fidelity.
The demo wakes as you arrive…

Beyond generation: text-driven image editing

GLIDE's second major capability is image inpainting — filling in erased regions of an existing image guided by text. The user masks a region (e.g., a blank area above a couch), provides a text prompt ("a painting of a corgi on the wall"), and the model fills in that region with content matching the prompt while respecting the surrounding context.

To achieve this, the model is fine-tuned with four additional input channels: the masked RGB image (3 channels) plus a binary mask channel. Random regions of training images are erased, and the model learns to reconstruct them conditioned on the surrounding pixels and the text caption. The additional input weights are initialized to zero so the model starts from its pre-trained capabilities.

The results are remarkable: the model produces shadows, reflections, and lighting that match the surrounding scene. It can even match artistic styles when editing paintings. Users can iteratively build complex scenes — starting with a generated room, then adding furniture, decorations, and details one text prompt at a time.

Open in Lab
See how GLIDE's inpainting pipeline works: the model receives the masked image plus text, and fills the erased region while matching the surrounding context.
The demo wakes as you arrive…

Safety considerations: power demands responsibility

The authors recognized that GLIDE could generate convincing fake images and enable deepfakes. Their response was multi-layered: they trained a smaller 300M parameter model (GLIDE filtered) on a dataset filtered to remove people, violence, and hate symbols. They released this smaller model publicly while retaining the full model.

Notably, the filtering process revealed biases: the model produced stereotypically pink toys for "toys for girls" versus gender-neutral results for "toys for boys," and tended to generate church-like buildings for "a religious place." These biases were amplified by classifier-free guidance. The authors were transparent about these limitations, setting an important precedent for responsible AI disclosure.

What GLIDE set in motion

  1. 2020

    DDPM — Diffusion enters the scene

    Ho et al. show that denoising diffusion probabilistic models can generate high-quality images, reviving interest in diffusion-based generation.

  2. 2021

    Diffusion beats GANs

    Dhariwal & Nichol introduce classifier guidance and show diffusion models surpass GANs on ImageNet — but only with class labels, not text.

  3. 2021

    DALL-E — Text to image via autoregression

    Ramesh et al. train a 12B parameter autoregressive model on text-image pairs, showing that free-form text prompts can compose complex scenes — but with blurry results.

  4. 2021

    GLIDE — Text-guided diffusion with classifier-free guidance

    Nichol et al. inject text into a 3.5B diffusion model, compare CLIP and classifier-free guidance, and show the latter produces photorealistic images preferred 87% over DALL-E.

  5. 2022

    DALL-E 2 — CLIP + diffusion

    Ramesh et al. combine CLIP embeddings with a diffusion decoder, building directly on GLIDE's insight that diffusion outperforms autoregression for image generation.

  6. 2022

    Imagen — Large language models meet diffusion

    Saharia et al. replace the text encoder with a frozen T5-XXL, showing that stronger text understanding dramatically improves image quality — using classifier-free guidance throughout.

  7. 2022

    Stable Diffusion — Diffusion goes latent and open

    Rombach et al. move diffusion to latent space for efficiency, release open weights, and democratize text-to-image generation. Classifier-free guidance remains central.

CitationNichol, Dhariwal, Ramesh, Shyam, Mishkin, McGrew, Sutskever, Chen. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. ICML, 2022.

Terms in this paper