Generative Models2021intermediate10 min read
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
GLIDE: توليد صور واقعية وتحريرها باستخدام نماذج انتشار موجَّهة بالنص
Nichol, A. · Dhariwal, P. · Ramesh, A. · Shyam, P. · Mishkin, P. · McGrew, B. · Sutskever, I. · Chen, M. — ICML
The problem
By late 2021, generation had made exciting progress. DALL-E showed that autoregressive models could compose concepts from text, and diffusion models had surpassed GANs on unconditional image quality. But no system could do both: generate photorealistic images from arbitrary text prompts. DALL-E's images were often blurry, and diffusion models lacked text conditioning. The missing piece was a way to inject free-form text into a powerful and guide it to produce images that both look real and match the description.
The contribution
A 3.5 billion parameter text-conditional diffusion model that generates photorealistic images from text using . The paper compares two guidance strategies — guidance and classifier-free guidance — and shows through human evaluation that classifier-free guidance produces images preferred 87% of the time for photorealism over DALL-E. The model can also be fine-tuned for , enabling text-driven editing with realistic shadows, reflections, and style matching.
The impact
GLIDE proved that diffusion models could surpass autoregressive models at text-to-image synthesis, establishing the diffusion paradigm that would dominate generative AI. It directly inspired DALL-E 2 (which combined CLIP with diffusion), Imagen (which used large language model encoders with classifier-free guidance), and Stable Diffusion (which moved diffusion to latent space). Classifier-free guidance became the default steering mechanism for virtually all subsequent text-to-image systems.
Imagine a sculptor working with a block of marble that's been shattered into random dust. A client whispers a description: "a corgi wearing a birthday hat." The sculptor cannot see the final statue — but they have two inner voices. One voice says: "here's what any random image looks like at this stage." The other says: "here's what an image matching that description looks like at this stage." By always chiseling toward the described voice and away from the generic one, the sculptor gradually reveals a photorealistic corgi in a party hat — carved from pure noise.
GLIDE is that sculptor. The two voices are classifier-free guidance, and the marble dust is Gaussian noise.
The problem: realistic images or text control — pick one
Before GLIDE, two families of models had made exciting but separate progress. On one side, diffusion models (like those in "Diffusion Models Beat GANs") had achieved photorealistic image quality on ImageNet — producing faces, animals, and scenes that humans could barely distinguish from real photographs. But they were class-conditional: you could ask for "dog" or "car," not "a corgi wearing a red bowtie and a purple party hat."
On the other side, DALL-E showed that an trained on text-image pairs could compose complex scenes from arbitrary text. But the images were often blurry and lacked the photorealism of diffusion models. The best DALL-E samples required generating hundreds of candidates and using CLIP to rerank them — an expensive post-processing step.
The question was straightforward: could you marry diffusion's image quality with DALL-E's text flexibility?
Quick recap: how diffusion models generate images
A diffusion model works in two phases. The forward process gradually destroys a real image by adding Gaussian noise over timesteps until it becomes pure static. The reverse process learns to undo this destruction: a neural network predicts the noise that was added at each step, and subtracts it to recover the image.
Think of it as learning to unscramble: if you know exactly what noise was stirred in, you can stir it back out. The model trains by seeing many noised images and learning to predict the noise component from the noisy version .
The GLIDE architecture: injecting text into diffusion
GLIDE's core idea is simple: take a powerful diffusion model and give it text understanding. The architecture has three components working together:
1. A — a with 24 layers and width 2048, roughly 1.2 billion parameters. It converts a text caption into a sequence of token embeddings.
2. A visual diffusion model — based on the ADM (Ablated Diffusion Model) architecture, scaled to 512 channels and roughly 2.3 billion parameters. This is the workhorse that denoises images.
3. Two injection points — the text enters the visual model in two ways. First, the final replaces the class label (where a class-conditional model would use "dog" or "car"). Second, the full sequence of token embeddings is projected and concatenated to the attention context at every attention layer. This gives each layer of the U-Net access to the full text description via .
Together, the model has 3.5 billion parameters — the base model — plus a separate 1.5 billion parameter diffusion model that takes the 64×64 output and increases it to 256×256 resolution.
Classifier-free guidance: the model steers itself
The most impactful contribution of GLIDE is demonstrating that classifier-free guidance dramatically outperforms CLIP guidance for text-to-image diffusion. But what is classifier-free guidance, and why does it work so well?
The intuition begins with a mental model. During sampling, the model produces two predictions at each step: one conditioned on the text caption , and one unconditional (as if no caption were provided, using an empty prompt ). The difference between these two predictions tells the model: "this is the direction that makes the image more like what the text describes." By amplifying that difference, you push the image harder toward the text.
Think of it as navigation with two compasses. One compass points toward "any image" (the unconditional prediction). The other points toward "an image matching this description" (the conditional prediction). The vector from the first to the second is the "text direction." Multiplying that vector by a makes the model follow the text more aggressively — at the cost of some diversity.
CLIP guidance: the external judge
The alternative guidance approach uses CLIP as an external evaluator. CLIP has two encoders — one for images and one for text — and outputs a similarity score. Higher dot product means the image better matches the caption.
To use CLIP for guidance, at each denoising step the model computes the gradient of the CLIP similarity with respect to the current noisy image . This gradient tells the model: "adjusting the image in this direction will increase its similarity to the text." The denoising mean is then shifted by this gradient, scaled by a guidance coefficient.
One critical detail: standard CLIP models were trained on clean images, not noisy ones. The noisy intermediate images during diffusion sampling are out-of-distribution for standard CLIP. GLIDE addresses this by training a noised CLIP model that explicitly processes noisy images , making the gradients more accurate.
The verdict: classifier-free guidance wins on all fronts
GLIDE ran a systematic comparison between the two guidance strategies. On automated metrics, CLIP guidance achieves higher CLIP scores — but this is misleading. The authors hypothesized that CLIP guidance finds adversarial examples for the CLIP evaluator: images that score high on CLIP similarity without actually looking better to humans.
The definitive answer came from human evaluation. Evaluators were shown pairs of images and asked to choose which was more photorealistic or better matched the caption. The results were decisive: classifier-free guidance won in both photorealism and caption similarity. When comparing GLIDE with classifier-free guidance against DALL-E (even with DALL-E using expensive CLIP reranking of 512 candidates), human evaluators preferred GLIDE 87% of the time for photorealism and 69% for caption similarity.
On MS-COCO, GLIDE achieved a zero-shot FID of 12.24 — competitive with models explicitly trained on COCO — using a model roughly 3× smaller than DALL-E (3.5B vs. 12B parameters).
Beyond generation: text-driven image editing
GLIDE's second major capability is image inpainting — filling in erased regions of an existing image guided by text. The user masks a region (e.g., a blank area above a couch), provides a text prompt ("a painting of a corgi on the wall"), and the model fills in that region with content matching the prompt while respecting the surrounding context.
To achieve this, the model is fine-tuned with four additional input channels: the masked RGB image (3 channels) plus a binary mask channel. Random regions of training images are erased, and the model learns to reconstruct them conditioned on the surrounding pixels and the text caption. The additional input weights are initialized to zero so the model starts from its pre-trained capabilities.
The results are remarkable: the model produces shadows, reflections, and lighting that match the surrounding scene. It can even match artistic styles when editing paintings. Users can iteratively build complex scenes — starting with a generated room, then adding furniture, decorations, and details one text prompt at a time.
Safety considerations: power demands responsibility
The authors recognized that GLIDE could generate convincing fake images and enable deepfakes. Their response was multi-layered: they trained a smaller 300M parameter model (GLIDE filtered) on a dataset filtered to remove people, violence, and hate symbols. They released this smaller model publicly while retaining the full model.
Notably, the filtering process revealed biases: the model produced stereotypically pink toys for "toys for girls" versus gender-neutral results for "toys for boys," and tended to generate church-like buildings for "a religious place." These biases were amplified by classifier-free guidance. The authors were transparent about these limitations, setting an important precedent for responsible AI disclosure.
What GLIDE set in motion
2020
DDPM — Diffusion enters the scene
Ho et al. show that denoising diffusion probabilistic models can generate high-quality images, reviving interest in diffusion-based generation.
2021
Diffusion beats GANs
Dhariwal & Nichol introduce classifier guidance and show diffusion models surpass GANs on ImageNet — but only with class labels, not text.
2021
DALL-E — Text to image via autoregression
Ramesh et al. train a 12B parameter autoregressive model on text-image pairs, showing that free-form text prompts can compose complex scenes — but with blurry results.
2021
GLIDE — Text-guided diffusion with classifier-free guidance
Nichol et al. inject text into a 3.5B diffusion model, compare CLIP and classifier-free guidance, and show the latter produces photorealistic images preferred 87% over DALL-E.
2022
DALL-E 2 — CLIP + diffusion
Ramesh et al. combine CLIP embeddings with a diffusion decoder, building directly on GLIDE's insight that diffusion outperforms autoregression for image generation.
2022
Imagen — Large language models meet diffusion
Saharia et al. replace the text encoder with a frozen T5-XXL, showing that stronger text understanding dramatically improves image quality — using classifier-free guidance throughout.
2022
Stable Diffusion — Diffusion goes latent and open
Rombach et al. move diffusion to latent space for efficiency, release open weights, and democratize text-to-image generation. Classifier-free guidance remains central.
CitationNichol, Dhariwal, Ramesh, Shyam, Mishkin, McGrew, Sutskever, Chen. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. ICML, 2022.
Terms in this paper
- Classifier-Free Guidanceالتوجيه بلا مُصنِّف
- CLIPبنية كليب للربط التقابلي
- Diffusion Modelنموذج الانتشار
- Text-to-Imageتحويل النص إلى صورة
- Image Inpaintingملء فراغات الصور
- Guidance Scaleمقياس التوجيه
- Noise Scheduleجدول الضوضاء
- Denoisingإزالة الضوضاء
- U-Netشبكة U-Net
- Cross-Attentionالانتباه التبادلي
- Transformerالمحوِّل
- FID Scoreمقياس FID
- Photorealistic Renderingالتصيير الواقعي
- Autoregressive Modelالنموذج التوليدي التراجعي
- Generative Modelالنموذج التوليدي