Generative Models2021advanced11 min read

DALL·E: Zero-Shot Text-to-Image Generation

DALL·E: توليد الصور من النصوص بدون أمثلة

Ramesh, A. · Pavlov, M. · Goh, G. · Gray, S. · Voss, C. · Radford, A. · Chen, M. · Sutskever, I. — ICML

The problem

Before DALL·E, generation relied on domain-specific architectures, auxiliary losses, and side information like object part labels or segmentation masks. These models worked well on narrow datasets (birds, flowers, faces) but could not generalize to arbitrary text prompts. Nobody had shown that a single model, trained at scale on text–image pairs without task-specific engineering, could generate coherent images for prompts it had never seen — from "a pentagon-shaped clock" to "a snail made of a harp."

The contribution

A two-stage approach that treats image generation as a sequence-prediction problem. Stage 1 trains a discrete () that compresses each 256×256 image into a 32×32 grid of tokens from a of 8,192 entries. Stage 2 trains a 12-billion-parameter that models 256 text tokens and 1,024 image tokens as a single stream, learning the joint distribution of text and images. At , the model generates 512 candidate images per prompt, and reranks them for the best text–image alignment. Without any task-specific fine-tuning, DALL·E achieves image generation competitive with domain-specific models on MS-COCO.

The impact

DALL·E proved that text-to-image generation could be solved with scale and simplicity — a large transformer, a large dataset, and no domain-specific tricks. It demonstrated that autoregressive models trained on internet-scale data can compose concepts zero-shot: combining objects, attributes, and spatial relationships in ways never seen during . This insight directly inspired GLIDE (which replaced the autoregressive with diffusion), DALL·E 2, Imagen, and the latent diffusion framework behind Stable Diffusion. DALL·E opened the door to the modern era of text-to-image generation.

Imagine an artist who only knows how to paint by numbers. First, an assistant studies thousands of paintings and creates a numbered palette of 8,192 colors — each number maps to a small patch of texture and color. Then a storyteller reads a description aloud — "a stained-glass window of a butterfly" — and the artist, who has memorized how stories map to number sequences, places color-patches one by one on a 32×32 canvas. The storyteller never shows the artist a butterfly; the artist reconstructs it purely from the rhythm of language.

DALL·E works the same way: a discrete VAE builds the palette, a giant transformer is the artist, and language is the only instruction.

The big idea: images are just long sentences

The core insight behind DALL·E is surprisingly simple: if we can turn an image into a sequence of discrete tokens — much like words in a sentence — then we can train a language model to predict those tokens given a text prompt. The model never "sees" pixels directly. Instead, it learns the joint probability of text tokens followed by image tokens, and generates images by sampling that distribution autoregressively — one at a time, left to right.

This unifies text understanding and image generation into a single architecture: the transformer. No GANs, no separate text feeding into a convolutional generator, no object detectors. Just a single stream of tokens and the same next-token prediction objective that powers GPT-3.

But there is a catch: a 256×256 color image has 196,608 pixel values. Modeling each pixel as a token would require a sequence length no transformer could handle. The solution is compression — and that is where the discrete VAE comes in.

Open in Lab
The full DALL·E pipeline: text prompt → BPE tokens + dVAE image tokens → transformer predicts the sequence → dVAE decoder reconstructs the image → CLIP reranks 512 candidates.
The demo wakes as you arrive…

Stage 1: learning the visual vocabulary (dVAE)

The first stage of DALL·E trains a discrete Variational Autoencoder — the dVAE — on images alone, without any text. Its job is to learn a visual codebook: a dictionary of 8,192 visual "words," each representing a small patch of texture, color, and shape.

The dVAE encoder takes a 256×256 RGB image and compresses it into a 32×32 grid. At each of the 1,024 grid positions, the encoder outputs a probability distribution over the 8,192 codebook entries, asking: "which visual word best describes this patch?" The decoder then takes those 1,024 chosen codebook vectors and reconstructs the full 256×256 image.

Think of it like a mosaic artist who has 8,192 tiles to choose from. For each small region of the image, the artist picks the tile that best matches that region. The dVAE learns both which tiles to offer (the codebook) and how to read them back into a coherent image (the decoder).

Open in Lab
See how the dVAE compresses a 256×256 image into 1,024 codebook tokens. Click a grid cell to see which codebook entry was chosen and its reconstruction.
The demo wakes as you arrive…

But there is a fundamental problem: picking the "best" codebook entry at each position requires an argmax operation, which is not differentiable — gradients cannot flow through a hard selection. Without gradients, we cannot train the encoder end-to-end.

The solution is the relaxation. Instead of picking one entry, the encoder outputs a soft probability distribution over all 8,192 entries. During training, the decoder receives a weighted average of codebook vectors rather than a single vector. A parameter τ\tau controls how "soft" this averaging is: high temperature gives a smooth blend; as temperature is annealed toward zero, the distribution sharpens toward a one-hot selection — approaching argmax without ever being non-differentiable.

At inference time, the temperature is set to zero (true argmax), so each position maps to exactly one codebook entry — producing the clean, discrete tokens that the transformer will later consume.

Open in Lab
Drag the temperature slider to see how Gumbel Softmax transitions from a soft blend of codebook vectors to a hard one-hot selection.
The demo wakes as you arrive…

The dVAE is trained by maximizing the . The ELBO has two parts: a reconstruction term that rewards the decoder for faithfully recreating the original image, and a term that keeps the encoder's distribution close to a uniform over the codebook. The balance ensures the model uses the codebook efficiently without collapsing to a few entries.

LdVAE=Eqϕ(z∣x)[ln⁡pθ(x∣z)]−KL(qϕ(z∣x) ∥ p(z))\mathcal{L}_{\text{dVAE}} = \mathbb{E}_{q_\phi(z|x)} \big[\ln p_\theta(x|z)\big] - \text{KL}\big(q_\phi(z|x) \,\|\, p(z)\big)
dVAE training objective — the Evidence Lower Bound (ELBO) — The first term measures how well the decoder pθp_\theta reconstructs the image xx from the latent code zz. The second term penalizes the encoder qϕq_\phi for deviating from a uniform prior, encouraging diverse codebook usage. The Gumbel Softmax relaxation makes the expectation differentiable.

Stage 2: the autoregressive transformer

Once the dVAE is trained and frozen, every image in the dataset can be encoded as 1,024 discrete tokens. Each text caption is encoded as up to 256 tokens with a of 16,384. The text and image tokens are concatenated into a single sequence of up to 1,280 tokens, and a 12-billion-parameter decoder-only transformer is trained to model their joint distribution autoregressively.

The transformer has 64 layers, each with 62 attention heads. It uses three types of attention masks merged into a single operation. Text tokens attend to all preceding text tokens (standard causal mask). Image tokens attend to all text tokens (full ) and to preceding image tokens using a pattern that alternates between row, column, and convolutional masks. This sparse pattern drastically reduces memory while preserving spatial coherence.

The loss function is a weighted cross-entropy: 1/8 weight on text tokens and 7/8 weight on image tokens. This reflects the priority — the model should focus on generating accurate images, using text comprehension as a means to that end.

Open in Lab
Explore the three attention masks: text-to-text (causal), image-to-text (full), and image-to-image (sparse row/column/conv). Toggle between them to see the pattern.
The demo wakes as you arrive…
Ltransformer=−18∑i=1256ln⁡p(yi∣y<i)−78∑j=11024ln⁡p(zj∣y,z<j)\mathcal{L}_{\text{transformer}} = - \frac{1}{8} \sum_{i=1}^{256} \ln p(y_i \mid y_{<i}) - \frac{7}{8} \sum_{j=1}^{1024} \ln p(z_j \mid y, z_{<j})
Transformer training loss — weighted cross-entropy over text and image tokens — yiy_i are text tokens and zjz_j are image tokens. The 7/8 weight on image tokens emphasizes visual generation quality. The model learns to read the text and continue with the correct image tokens — the same next-token objective as GPT-3.

Inference: generating and reranking with CLIP

At inference time, the model generates images by autoregressively sampling image tokens conditioned on the text prompt. But autoregressive sampling is stochastic — each run produces a different image, and quality varies. Some samples nail the prompt, others miss entirely.

The authors' solution is to generate many candidates and let a judge pick the best one. They sample 512 images for each prompt, then use CLIP — a model trained to understand text–image alignment — to score each candidate against the original text. The top-ranked images are selected as the final output.

This is like holding 512 auditions for a role and letting a casting director (CLIP) pick the best performance. The transformer is the actor; CLIP is the critic. Together, they produce results far better than any single sample.

Open in Lab
See how CLIP reranking works: 512 samples are scored and the top results are selected. Toggle to see the difference between random samples and CLIP-ranked output.
The demo wakes as you arrive…

Zero-shot generalization: composing unseen concepts

The most remarkable property of DALL·E is its ability to compose concepts it has never seen together during training. Given "an armchair in the shape of an avocado," it generates plausible images even though no such object exists in any training set. Given "a store front that has the word 'openai' written on it," it renders readable text in context.

The model can also perform zero-shot image-to-image translation. By prefixing the prompt with "the same [object] in the style of [X]," DALL·E can change textures, materials, and artistic styles without ever being trained on paired style-transfer datasets. It handles geographic knowledge ("a photo of the food of China"), temporal knowledge ("a photo of a phone from the 1920s"), and even cross-domain analogies ("a snail made of a harp").

This zero-shot compositional ability comes entirely from scale: a sufficiently large model trained on sufficiently diverse text–image pairs learns to bind adjectives to nouns, spatial prepositions to objects, and stylistic modifiers to scenes — all through next-token prediction alone.

Open in Lab
Compose concepts interactively: pick an object, a material, and a style to see how DALL·E would combine them — a demonstration of zero-shot compositional generalization.
The demo wakes as you arrive…

Putting it all together: the two-stage pipeline

Let us walk through the complete DALL·E pipeline from prompt to final image:

Stage 1 (offline, once): Train the dVAE on images alone. The encoder learns to compress 256×256 images to 32×32 grids of discrete tokens (codebook size 8,192). The decoder learns to reconstruct images from these tokens. After training, the dVAE is frozen.

Stage 2 (offline, once): Encode every training image into 1,024 tokens using the frozen dVAE encoder. Encode every caption into up to 256 BPE tokens. Concatenate text + image tokens and train the 12B transformer to model the joint distribution autoregressively.

Inference: Given a new text prompt, BPE-encode it to ≤256 tokens. Sample 1,024 image tokens autoregressively from the transformer, conditioned on the text. Decode the tokens to a 256×256 image using the frozen dVAE decoder. Repeat 512 times. Use CLIP to rank all 512 candidates and return the best ones.

Pseudocode: DALL·E inference pipelinepython

Simplified to show the idea — not the real implementation.

# DALL-E Inference Pipeline (Simplified)
def generate_image(text_prompt, n_candidates=512):
    # Step 1: Encode text using BPE
    text_tokens = bpe_encode(text_prompt, max_len=256, vocab_size=16384)

    # Step 2: Sample image tokens autoregressively
    candidates = []
    for _ in range(n_candidates):
        img_tokens = []
        for j in range(1024):  # 32x32 grid
            logits = transformer(text_tokens + img_tokens)
            next_token = sample(logits, temperature=1.0)
            img_tokens.append(next_token)

        # Step 3: Decode tokens to image via frozen dVAE
        image = dvae_decoder(img_tokens)  # 1024 tokens -> 256x256 image
        candidates.append(image)

    # Step 4: Rerank with CLIP
    scores = [clip_score(text_prompt, img) for img in candidates]
    best = sorted(zip(scores, candidates), reverse=True)
    return [img for _, img in best[:32]]

Legacy: from tokens to diffusion

  1. 2017

    VQ-VAE (van den Oord et al.)

    Introduced vector quantization for learning discrete latent representations of images. Showed that images can be faithfully reconstructed from a small number of discrete codes — the foundation for DALL·E's visual codebook.

  2. 2020

    GPT-3 (Brown et al.)

    Demonstrated that scaling autoregressive transformers to 175 billion parameters enables remarkable few-shot and zero-shot generalization in language. DALL·E directly applied this insight to the visual domain.

  3. 2021

    DALL·E + CLIP (Ramesh et al. / Radford et al.)

    DALL·E showed that autoregressive token prediction can generate coherent images from text. CLIP, released simultaneously, provided the reranking mechanism that made zero-shot generation practical.

  4. 2022

    GLIDE, DALL·E 2, Imagen, Stable Diffusion

    The next generation replaced autoregressive decoding with diffusion models, achieving higher resolution and better coherence. But the core insight — treat image generation as a text-conditioned sequence problem at scale — came directly from DALL·E.

DALL·E's lasting contribution is not the specific architecture — autoregressive image generation has largely been replaced by diffusion. Its contribution is the paradigm: that scale and simplicity beat domain-specific engineering. A sufficiently large transformer, trained on enough text–image pairs with a standard next-token objective, can learn to compose visual concepts it has never seen together. Every text-to-image model that followed — GLIDE, DALL·E 2, Imagen, Stable Diffusion — builds on this foundation.

CitationRamesh, Pavlov, Goh, Gray, Voss, Radford, Chen, Sutskever. Zero-Shot Text-to-Image Generation. ICML, 2021.

Terms in this paper