Multimodal AI2015intermediate11 min read

Show and Tell: A Neural Image Caption Generator

أرِني وأخبرني: مولّد وصف الصور بالشبكات العصبية

Vinyals, O. · Toshev, A. · Bengio, S. · Erhan, D. — CVPR

The problem

Before 2015, image understanding and language generation were separate worlds. Image classifiers could label a photo "cat" or "dog," but could not say "A cat sitting on a windowsill next to a potted plant." Earlier captioning systems glued hand-crafted visual detectors to rigid sentence templates, producing stiff, formulaic descriptions that broke on anything the templates didn't anticipate.

The contribution

A single end-to-end neural network that reads an image and writes a natural-language caption. A () encodes the image into a fixed-length vector, which is fed once as the first "word" to an that generates the caption word by word. The entire pipeline — vision backbone plus language model — is trained jointly by maximizing the likelihood of correct captions, with no hand-designed features or rules. at produces fluent, accurate descriptions.

The impact

Show and Tell was the first model to convincingly demonstrate that a single neural network could bridge vision and language end-to-end, establishing the –decoder paradigm for . It won the 2015 MSCOCO Image Captioning Challenge and inspired a wave of vision-language models — from Show, Attend and Tell (adding ) to modern VQA systems and multimodal large language models. The CNN-as-encoder, RNN-as-decoder blueprint became the template for an entire field.

Imagine visiting a museum with a friend who speaks a different language. You look at a painting and absorb the whole scene — colors, shapes, composition — then compress that experience into a single mental snapshot. You hand that snapshot to your friend the storyteller, who unrolls it into a vivid sentence in their language, one word at a time: "A girl in a red dress is standing on a bridge over a river."

That is exactly what Show and Tell does. The CNN is the viewer: it stares at the image and distills everything into a compact code. The LSTM is the storyteller: it receives that code and narrates what it "saw," choosing each word based on the code plus every word it has already written.

The problem: images need sentences, not just labels

By 2014, CNNs like GoogLeNet could classify images into 1,000 categories with near-human accuracy. But only answers "what is it?" — it cannot describe relationships, actions, or spatial arrangements. "Dog" tells you nothing about whether the dog is running on a beach, sleeping on a couch, or catching a frisbee.

Earlier captioning approaches took a pipeline strategy: run an object detector, an attribute classifier, and a spatial-relation extractor, then stitch the outputs into a sentence using templates or a language model. Every stage introduced errors that cascaded downstream, and the templates could only produce sentences the engineers had anticipated. Truly novel scenes — a cat riding a surfboard — broke the system entirely.

What was needed was a model that learns to go from raw pixels to complete sentences in a single, trainable pipeline — the way models were beginning to do for machine translation.

Open in Lab
Left: the old pipeline approach — detect, classify, template. Right: Show and Tell's end-to-end approach — one network, pixels to sentence.
The demo wakes as you arrive…

The idea: treat an image as the first word of a sentence

The key insight is borrowed directly from machine translation. In seq2seq, an encoder reads a source sentence and compresses it into a fixed-length vector, then a decoder generates the target sentence from that vector. Show and Tell replaces the text encoder with an image encoder:

  • Encoder (CNN): The last hidden layer of GoogLeNet — a deep convolutional network pre-trained on ImageNet — produces a fixed-length vector that summarizes the image content. Think of it as a 1024-dimensional "description" of the scene that captures objects, their properties, and their spatial relationships, all encoded as numbers.

  • Decoder (LSTM): A Long Short-Term Memory network receives the image vector as its initial input (at time step -1) and generates the caption one word at a time. At each step, it takes the previous word's , updates its , and predicts the next word until it produces a special .

The image is fed once at the beginning, not at every time step. The authors found that feeding the image repeatedly caused the LSTM to rely too heavily on it and produce less fluent language. A single injection forces the LSTM to internalize the visual information in its memory cells and recall it as needed — like a storyteller who glances at a scene once and then speaks from memory.

Open in Lab
Click on any stage to see what happens at that point in the pipeline.
The demo wakes as you arrive…

The training objective: maximize the caption probability

The model is trained to maximize the probability of the correct caption given the image. This is a standard maximum likelihood objective. For an image II and its correct caption S=(S0,S1,…,SN)S = (S_0, S_1, \ldots, S_N), the model maximizes:

θ∗=arg⁡max⁡θ∑(I,S)log⁡p(S∣I; θ)\theta^* = \arg\max_\theta \sum_{(I,S)} \log p(S \mid I;\, \theta)
Training objective — maximize log-probability of correct captions — θ* = the best parameters · (I, S) = image-caption pairs from the dataset · p(S|I; θ) = the probability the model assigns to caption S for image I

Since a caption is a sequence of words, the joint probability factorizes using the — each word depends on the image and on all previous words:

log⁡p(S∣I)=∑t=0Nlog⁡p(St∣I,S0,…,St−1)\log p(S \mid I) = \sum_{t=0}^{N} \log p(S_t \mid I, S_0, \ldots, S_{t-1})
Chain-rule decomposition — language models generate one word at a time — Each term p(Sₜ | I, S₀, …, Sₜ₋₁) is modeled by the LSTM: the hidden state carries all past context, and softmax over the vocabulary gives the next-word distribution.

Inside the decoder: how the LSTM writes

At each time step tt, the LSTM performs three operations:

  1. Embed the previous word: convert the one-hot word vector St−1S_{t-1} into a dense embedding WeSt−1W_e S_{t-1}. At t=0t = 0, this is a special START token; at t=−1t = -1, this is the CNN image vector passed through a learned linear projection.

  2. Update the hidden state: the LSTM gates (input, forget, output) decide what to remember, what to forget, and what to output. The cell state acts as a conveyor belt carrying visual and linguistic information forward.

  3. Predict the next word: a linear layer followed by maps the hidden state to a probability distribution over the entire . The word with the highest probability is selected (or beam search explores multiple candidates).

x−1=Wimg CNN(I),xt=We St    (t≥0),pt+1=LSTM(xt)x_{-1} = W_{img}\, \text{CNN}(I), \quad x_t = W_e\, S_t \;\; (t \ge 0), \quad p_{t+1} = \text{LSTM}(x_t)
The three equations of Show and Tell — x₋₁ = image embedding (CNN output projected to LSTM dimension) · xₜ = word embedding at step t · pₜ₊₁ = probability of next word, produced by LSTM then softmax
Open in Lab
Step through the LSTM decoding process word by word. Watch how the hidden state evolves and how each word probability is computed.
The demo wakes as you arrive…

Training: teacher forcing

During , the model uses : at each time step, it receives the ground-truth previous word rather than its own prediction. Think of a language student being corrected word by word: even if the student says the wrong word, the teacher supplies the right one so the student can continue from the correct context. This accelerates training because the model never spirals into a sequence of errors.

The loss is the sum of cross-entropy losses at each time step — each word prediction is scored against the correct word, and the total loss is backpropagated through both the LSTM and the last layer of the CNN. This joint training means the CNN learns image features that are useful for describing images, not just classifying them.

Open in Lab
Toggle between teacher forcing (training) and free-running (inference) to see how error propagation differs.
The demo wakes as you arrive…

Inference: beam search for better captions

At inference time, there is no teacher — the LSTM must generate words on its own. The simplest strategy is : at each step, pick the single most probable word. But the locally best word at step 3 might lead to a dead end at step 7.

Beam search hedges this bet. It keeps the top kk partial captions (the "beam") at every step, expands each by one word, scores all candidates, and keeps the best kk again. With a beam size of 20, the model explores 20 parallel hypothesis "threads" and returns the highest-scoring complete caption.

This matters because language has long-range structure. The word "are" at position 5 might only make sense if the subject at position 1 was plural — but greedy decoding already committed to a singular subject. Beam search can recover from such mistakes by keeping the plural hypothesis alive.

Open in Lab
Watch beam search explore multiple caption hypotheses in parallel. Adjust the beam size to see how wider beams produce better captions.
The demo wakes as you arrive…

The visual encoder: from pixels to meaning

The CNN used is GoogLeNet ( v1), winner of the 2014 ILSVRC classification challenge. Its key innovation is the Inception module: instead of choosing one filter size, it applies 1×1, 3×3, and 5×5 convolutions in parallel and concatenates the results. This lets the network capture patterns at multiple scales simultaneously.

For captioning, the authors remove GoogLeNet's final classification layer (the 1,000-class softmax) and use the 1024-dimensional activation vector from the layer just below. This vector is a rich, learned summary of the image — it encodes not just "what objects are present" but spatial relationships, textures, and scene context. A learned linear transformation WimgW_{img} then projects this 1024-d vector into the LSTM's embedding space (typically 512-d), so the image occupies the same representational space as words.

Crucially, this last CNN layer is fine-tuned during captioning training. The gradients from the captioning loss flow back into the CNN, adjusting its features to emphasize what matters for description rather than classification — for example, learning to encode spatial relationships ("on top of," "next to") that a classifier would ignore.

Open in Lab
See how GoogLeNet processes an image from pixels through Inception modules to a single 1024-d vector.
The demo wakes as you arrive…

Results: how well does it work?

Show and Tell was evaluated on multiple captioning benchmarks using the BLEU metric, which measures n-gram overlap between the generated caption and reference captions written by humans. Higher BLEU scores indicate closer match to human descriptions:

  • Pascal VOC 2008: BLEU-1 jumped from 25 (previous state of the art) to 59, approaching human performance at 69.
  • Flickr30k: BLEU-1 improved from 56 to 66.
  • SBU: BLEU-1 rose from 19 to 28.
  • MSCOCO: The model achieved BLEU-4 of 27.7, a new state of the art, and won the 2015 MSCOCO captioning challenge.

Perhaps more impressive than the numbers are the qualitative results. The model generates descriptions that are grammatically correct, semantically meaningful, and often surprisingly detailed — capturing actions, attributes, and spatial relationships that no template-based system could produce.

Open in Lab
Compare BLEU scores across datasets. Hover over bars to see the improvement from previous best to Show and Tell.
The demo wakes as you arrive…

The idea in code

Show and Tell — CNN encoder + LSTM decoderpython

Simplified to show the idea — not the real implementation.

import numpy as np

def softmax(x):
    e = np.exp(x - x.max(axis=-1, keepdims=True))
    return e / e.sum(axis=-1, keepdims=True)

def show_and_tell(image, W_img, W_embed, lstm, W_out, vocab, max_len=20):
    """
    image:   CNN feature vector (1024,)
    W_img:   projection matrix (1024 x embed_dim) — maps image to word space
    W_embed: word embedding matrix (vocab_size x embed_dim)
    lstm:    an LSTM cell with .step(input, hidden) -> (output, new_hidden)
    W_out:   output projection (hidden_dim x vocab_size)
    """
    # Step -1: Feed the image as the first "word"
    img_embed = image @ W_img                    # (embed_dim,)
    hidden = lstm.step(img_embed, init_hidden())  # LSTM reads the image

    # Step 0+: Generate caption word by word
    word = START_TOKEN
    caption = []
    for t in range(max_len):
        word_embed = W_embed[word]               # look up embedding
        output, hidden = lstm.step(word_embed, hidden)
        logits = output @ W_out                  # score every vocab word
        probs = softmax(logits)
        word = np.argmax(probs)                  # greedy pick
        if word == STOP_TOKEN:
            break
        caption.append(vocab[word])

    return " ".join(caption)

# That's the whole pipeline:
# CNN(image) → project → LSTM → softmax → words
# All parameters (W_img, W_embed, LSTM, W_out) trained jointly.

Why it mattered

  1. 2014

    Seq2Seq for translation (Sutskever et al.)

    Proved that an encoder-decoder LSTM can translate between languages, establishing the paradigm Show and Tell borrows from.

  2. 2015

    Show and Tell (this paper)

    Replaced the text encoder with a CNN image encoder, creating the first end-to-end neural image captioner. Won MSCOCO 2015 challenge.

  3. 2015

    Show, Attend and Tell (Xu et al.)

    Added spatial attention — the LSTM "looks at" different image regions at each word, rather than relying on a single fixed vector. Became the dominant captioning architecture.

  4. 2015

    VQA — Visual Question Answering (Antol et al.)

    Extended the vision-language paradigm from captioning to answering questions about images, opening a new research direction.

  5. 2017

    Bottom-Up and Top-Down Attention (Anderson et al.)

    Combined object-level features (from Faster R-CNN) with top-down attention, pushing captioning quality further. Built directly on Show and Tell's encoder-decoder paradigm.

  6. 2021

    CLIP — Contrastive Language-Image Pretraining

    Learned joint image-text representations from 400M web pairs. The vision-language bridge Show and Tell pioneered, scaled to internet size.

  7. 2023

    LLaVA — Large Language-and-Vision Assistant

    Connected a vision encoder to a large language model via a projection layer — the same CNN → project → language model pipeline Show and Tell introduced, now with a billion-parameter LLM as the decoder.

The direct line from Show and Tell to modern multimodal AI is remarkably clear: replace the CNN with a , replace the LSTM with a Transformer decoder (or a full LLM), scale the data from MSCOCO to the entire web — and you get systems like GPT-4 Vision, Gemini, and Claude with vision capabilities. The core idea — encode the image, decode the language — has never changed.

CitationVinyals, Toshev, Bengio, Erhan. Show and Tell: A Neural Image Caption Generator. CVPR, 2015.

Terms in this paper