Multimodal AI2015intermediate11 min read
Show and Tell: A Neural Image Caption Generator
أرِني وأخبرني: مولّد وصف الصور بالشبكات العصبية
Vinyals, O. · Toshev, A. · Bengio, S. · Erhan, D. — CVPR
The problem
Before 2015, image understanding and language generation were separate worlds. Image classifiers could label a photo "cat" or "dog," but could not say "A cat sitting on a windowsill next to a potted plant." Earlier captioning systems glued hand-crafted visual detectors to rigid sentence templates, producing stiff, formulaic descriptions that broke on anything the templates didn't anticipate.
The contribution
A single end-to-end neural network that reads an image and writes a natural-language caption. A () encodes the image into a fixed-length vector, which is fed once as the first "word" to an that generates the caption word by word. The entire pipeline — vision backbone plus language model — is trained jointly by maximizing the likelihood of correct captions, with no hand-designed features or rules. at produces fluent, accurate descriptions.
The impact
Show and Tell was the first model to convincingly demonstrate that a single neural network could bridge vision and language end-to-end, establishing the –decoder paradigm for . It won the 2015 MSCOCO Image Captioning Challenge and inspired a wave of vision-language models — from Show, Attend and Tell (adding ) to modern VQA systems and multimodal large language models. The CNN-as-encoder, RNN-as-decoder blueprint became the template for an entire field.
Imagine visiting a museum with a friend who speaks a different language. You look at a painting and absorb the whole scene — colors, shapes, composition — then compress that experience into a single mental snapshot. You hand that snapshot to your friend the storyteller, who unrolls it into a vivid sentence in their language, one word at a time: "A girl in a red dress is standing on a bridge over a river."
That is exactly what Show and Tell does. The CNN is the viewer: it stares at the image and distills everything into a compact code. The LSTM is the storyteller: it receives that code and narrates what it "saw," choosing each word based on the code plus every word it has already written.
The problem: images need sentences, not just labels
By 2014, CNNs like GoogLeNet could classify images into 1,000 categories with near-human accuracy. But only answers "what is it?" — it cannot describe relationships, actions, or spatial arrangements. "Dog" tells you nothing about whether the dog is running on a beach, sleeping on a couch, or catching a frisbee.
Earlier captioning approaches took a pipeline strategy: run an object detector, an attribute classifier, and a spatial-relation extractor, then stitch the outputs into a sentence using templates or a language model. Every stage introduced errors that cascaded downstream, and the templates could only produce sentences the engineers had anticipated. Truly novel scenes — a cat riding a surfboard — broke the system entirely.
What was needed was a model that learns to go from raw pixels to complete sentences in a single, trainable pipeline — the way models were beginning to do for machine translation.
The idea: treat an image as the first word of a sentence
The key insight is borrowed directly from machine translation. In seq2seq, an encoder reads a source sentence and compresses it into a fixed-length vector, then a decoder generates the target sentence from that vector. Show and Tell replaces the text encoder with an image encoder:
-
Encoder (CNN): The last hidden layer of GoogLeNet — a deep convolutional network pre-trained on ImageNet — produces a fixed-length vector that summarizes the image content. Think of it as a 1024-dimensional "description" of the scene that captures objects, their properties, and their spatial relationships, all encoded as numbers.
-
Decoder (LSTM): A Long Short-Term Memory network receives the image vector as its initial input (at time step -1) and generates the caption one word at a time. At each step, it takes the previous word's , updates its , and predicts the next word until it produces a special .
The image is fed once at the beginning, not at every time step. The authors found that feeding the image repeatedly caused the LSTM to rely too heavily on it and produce less fluent language. A single injection forces the LSTM to internalize the visual information in its memory cells and recall it as needed — like a storyteller who glances at a scene once and then speaks from memory.
The training objective: maximize the caption probability
The model is trained to maximize the probability of the correct caption given the image. This is a standard maximum likelihood objective. For an image and its correct caption , the model maximizes:
Since a caption is a sequence of words, the joint probability factorizes using the — each word depends on the image and on all previous words:
Inside the decoder: how the LSTM writes
At each time step , the LSTM performs three operations:
-
Embed the previous word: convert the one-hot word vector into a dense embedding . At , this is a special START token; at , this is the CNN image vector passed through a learned linear projection.
-
Update the hidden state: the LSTM gates (input, forget, output) decide what to remember, what to forget, and what to output. The cell state acts as a conveyor belt carrying visual and linguistic information forward.
-
Predict the next word: a linear layer followed by maps the hidden state to a probability distribution over the entire . The word with the highest probability is selected (or beam search explores multiple candidates).
Training: teacher forcing
During , the model uses : at each time step, it receives the ground-truth previous word rather than its own prediction. Think of a language student being corrected word by word: even if the student says the wrong word, the teacher supplies the right one so the student can continue from the correct context. This accelerates training because the model never spirals into a sequence of errors.
The loss is the sum of cross-entropy losses at each time step — each word prediction is scored against the correct word, and the total loss is backpropagated through both the LSTM and the last layer of the CNN. This joint training means the CNN learns image features that are useful for describing images, not just classifying them.
Inference: beam search for better captions
At inference time, there is no teacher — the LSTM must generate words on its own. The simplest strategy is : at each step, pick the single most probable word. But the locally best word at step 3 might lead to a dead end at step 7.
Beam search hedges this bet. It keeps the top partial captions (the "beam") at every step, expands each by one word, scores all candidates, and keeps the best again. With a beam size of 20, the model explores 20 parallel hypothesis "threads" and returns the highest-scoring complete caption.
This matters because language has long-range structure. The word "are" at position 5 might only make sense if the subject at position 1 was plural — but greedy decoding already committed to a singular subject. Beam search can recover from such mistakes by keeping the plural hypothesis alive.
The visual encoder: from pixels to meaning
The CNN used is GoogLeNet ( v1), winner of the 2014 ILSVRC classification challenge. Its key innovation is the Inception module: instead of choosing one filter size, it applies 1×1, 3×3, and 5×5 convolutions in parallel and concatenates the results. This lets the network capture patterns at multiple scales simultaneously.
For captioning, the authors remove GoogLeNet's final classification layer (the 1,000-class softmax) and use the 1024-dimensional activation vector from the layer just below. This vector is a rich, learned summary of the image — it encodes not just "what objects are present" but spatial relationships, textures, and scene context. A learned linear transformation then projects this 1024-d vector into the LSTM's embedding space (typically 512-d), so the image occupies the same representational space as words.
Crucially, this last CNN layer is fine-tuned during captioning training. The gradients from the captioning loss flow back into the CNN, adjusting its features to emphasize what matters for description rather than classification — for example, learning to encode spatial relationships ("on top of," "next to") that a classifier would ignore.
Results: how well does it work?
Show and Tell was evaluated on multiple captioning benchmarks using the BLEU metric, which measures n-gram overlap between the generated caption and reference captions written by humans. Higher BLEU scores indicate closer match to human descriptions:
- Pascal VOC 2008: BLEU-1 jumped from 25 (previous state of the art) to 59, approaching human performance at 69.
- Flickr30k: BLEU-1 improved from 56 to 66.
- SBU: BLEU-1 rose from 19 to 28.
- MSCOCO: The model achieved BLEU-4 of 27.7, a new state of the art, and won the 2015 MSCOCO captioning challenge.
Perhaps more impressive than the numbers are the qualitative results. The model generates descriptions that are grammatically correct, semantically meaningful, and often surprisingly detailed — capturing actions, attributes, and spatial relationships that no template-based system could produce.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def softmax(x):
e = np.exp(x - x.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
def show_and_tell(image, W_img, W_embed, lstm, W_out, vocab, max_len=20):
"""
image: CNN feature vector (1024,)
W_img: projection matrix (1024 x embed_dim) — maps image to word space
W_embed: word embedding matrix (vocab_size x embed_dim)
lstm: an LSTM cell with .step(input, hidden) -> (output, new_hidden)
W_out: output projection (hidden_dim x vocab_size)
"""
# Step -1: Feed the image as the first "word"
img_embed = image @ W_img # (embed_dim,)
hidden = lstm.step(img_embed, init_hidden()) # LSTM reads the image
# Step 0+: Generate caption word by word
word = START_TOKEN
caption = []
for t in range(max_len):
word_embed = W_embed[word] # look up embedding
output, hidden = lstm.step(word_embed, hidden)
logits = output @ W_out # score every vocab word
probs = softmax(logits)
word = np.argmax(probs) # greedy pick
if word == STOP_TOKEN:
break
caption.append(vocab[word])
return " ".join(caption)
# That's the whole pipeline:
# CNN(image) → project → LSTM → softmax → words
# All parameters (W_img, W_embed, LSTM, W_out) trained jointly.Why it mattered
2014
Seq2Seq for translation (Sutskever et al.)
Proved that an encoder-decoder LSTM can translate between languages, establishing the paradigm Show and Tell borrows from.
2015
Show and Tell (this paper)
Replaced the text encoder with a CNN image encoder, creating the first end-to-end neural image captioner. Won MSCOCO 2015 challenge.
2015
Show, Attend and Tell (Xu et al.)
Added spatial attention — the LSTM "looks at" different image regions at each word, rather than relying on a single fixed vector. Became the dominant captioning architecture.
2015
VQA — Visual Question Answering (Antol et al.)
Extended the vision-language paradigm from captioning to answering questions about images, opening a new research direction.
2017
Bottom-Up and Top-Down Attention (Anderson et al.)
Combined object-level features (from Faster R-CNN) with top-down attention, pushing captioning quality further. Built directly on Show and Tell's encoder-decoder paradigm.
2021
CLIP — Contrastive Language-Image Pretraining
Learned joint image-text representations from 400M web pairs. The vision-language bridge Show and Tell pioneered, scaled to internet size.
2023
LLaVA — Large Language-and-Vision Assistant
Connected a vision encoder to a large language model via a projection layer — the same CNN → project → language model pipeline Show and Tell introduced, now with a billion-parameter LLM as the decoder.
The direct line from Show and Tell to modern multimodal AI is remarkably clear: replace the CNN with a , replace the LSTM with a Transformer decoder (or a full LLM), scale the data from MSCOCO to the entire web — and you get systems like GPT-4 Vision, Gemini, and Claude with vision capabilities. The core idea — encode the image, decode the language — has never changed.
CitationVinyals, Toshev, Bengio, Erhan. Show and Tell: A Neural Image Caption Generator. CVPR, 2015.
Terms in this paper
- Convolutional Neural Network (CNN)الشبكة العصبية الالتفافية
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Beam Searchبحث الحزمة
- Image Captioningوصف الصور
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Teacher Forcingالتوجيه بالمرجع
- BLEU Scoreمعيار بلو الإحصائي لتقييم الترجمة
- Word Embeddingتضمين الكلمة
- GoogLeNetغوغل نت
- Softmaxسوفت ماكس
- Sequence-to-Sequenceتسلسل إلى تسلسل
- Maximum Likelihood Estimationتقدير الأرجحية القصوى