Computer Vision2015intermediate12 min read
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
أَرِني، انتبِه وأخبرني: توليد وصف الصور بالانتباه البصري
Xu, K. · Ba, J. · Kiros, R. · Cho, K. · Courville, A. · Salakhutdinov, R. · Zemel, R. · Bengio, Y. — ICML
The problem
Early models (like Show and Tell) encoded the entire into a single fixed-length vector using a , then fed that vector to an to generate a caption. This forced all visual information — objects, relationships, background — through one narrow channel. The LSTM had to remember everything from that single snapshot, making it difficult to generate accurate descriptions of complex scenes with multiple objects.
The contribution
An mechanism for image captioning that lets the LSTM "look back" at the image at every time step. Instead of a single vector, a CNN produces a grid of spatial feature vectors (annotation vectors). At each word, an attention module scores every grid location, and the decoder receives a weighted combination of those features — focusing on the relevant image region. Two variants are proposed: deterministic (differentiable, trained with ) and stochastic hard attention (samples one region, trained with REINFORCE). A encourages the model to attend to the full image over the course of a caption.
The impact
This paper was the first to bring to image captioning, demonstrating that neural networks can learn where to look — and show us where they looked. The attention maps provided an unprecedented window into the model's reasoning, making it a landmark for interpretable AI. It directly inspired attention mechanisms across computer vision, from VQA to dense captioning, and laid the conceptual groundwork for vision-language models like CLIP.
Imagine you're a tour guide describing a painting to a visitor over the phone. The old approach was like glancing at the painting once, memorizing everything, then talking for two minutes from memory. You'd get the gist right but mix up details — was the dog on the left or the right? Was there a boat in the background?
Show, Attend and Tell replaces that with a pointer and a magnifying glass: as you say each word, you slide the magnifying glass to the part of the painting that word describes. Saying "dog"? You zoom into the dog. Saying "frisbee"? The glass slides to the frisbee. You never have to hold the whole painting in your head — you just look at what matters right now.
The bottleneck: one vector for an entire image
Before this paper, image captioning followed a simple two-stage pipeline:
-
A convolutional neural network (typically VGGNet or GoogLeNet) processes the image and outputs a single — a fixed-length numerical summary of the entire scene.
-
An LSTM decoder receives that vector once, at the first time step, then generates the caption word by word purely from its .
The problem is a bottleneck: thousands of pixels, dozens of objects, and all their spatial relationships must compress into one vector of 512 or 4096 numbers. Early words get the freshest signal; later words rely on the LSTM's fading memory of that initial snapshot. As scenes grow more complex, the descriptions grow more generic — the model says "a man on a field" instead of "a football player catching a ball near the goal post".
The idea: let the decoder look back at the image
The key insight is borrowed from Bahdanau's attention for : instead of encoding the image into a single vector, keep a grid of feature vectors and let the decoder choose which grid cells to focus on at each time step. The architecture has three parts:
Encoder (CNN): Extract features not from the final fully connected layer, but from a lower convolutional layer — specifically the fourth convolutional layer of VGGNet. This produces annotation vectors, each of dimension . Each vector encodes information about a particular spatial region of the image. Think of it as a 14×14 grid overlay on the image, where each cell contains a rich description of what's in that cell.
Attention mechanism: At each time step, a small neural network scores how relevant each of the 196 locations is, given the decoder's current state. The scores become weights (via ), and the weighted sum of annotation vectors produces a — a dynamic summary that highlights the part of the image relevant to the current word.
Decoder (LSTM): Generates one word at a time, conditioned on the previous word, the previous hidden state, and the context vector from the attention mechanism. This means the decoder's input changes at every time step — it sees a different "view" of the image depending on what word it's about to generate.
How attention works: scoring, weighting, and focusing
At each time step , the model must decide which of the image regions to attend to. It does this in three stages:
Stage 1 — Energy score: A small alignment network takes each and the previous hidden state , and produces a scalar energy measuring how relevant region is to the current decoding step. This is — the same mechanism Bahdanau used for translation.
Stage 2 — Attention weights: The energies are normalized via softmax so they sum to 1: is the probability that location is the right place to focus on when generating word .
Stage 3 — Context vector: The context vector is a weighted combination of annotation vectors. This single vector — a dynamic, word-specific summary of the image — is fed to the LSTM along with the previous word .
Two flavors: soft attention vs. hard attention
The paper proposes two variants of the attention mechanism, each with different tradeoffs:
Soft (deterministic) attention takes the weighted average of all annotation vectors. Every region contributes, but regions with higher weights contribute more. This is fully differentiable — you can train it end-to-end with standard backpropagation. Think of it as looking at the whole painting through a filter that makes relevant areas brighter and irrelevant areas dimmer.
Hard (stochastic) attention samples one region according to the distribution. At each time step, the model picks exactly one grid location and uses only as the context. This is like putting a spotlight on one part of the painting and ignoring everything else. Because sampling is not differentiable, hard attention must be trained with the REINFORCE algorithm — a method from that treats the attention decisions as a sequence of actions.
The soft variant is simpler to train and more stable; the hard variant can sometimes achieve slightly higher performance because it forces sharper focus, but at the cost of higher during .
Doubly stochastic attention: don't ignore any part of the image
By construction, the attention weights at each time step sum to 1 over locations (each word's focus is a valid probability distribution). But nothing prevents the model from always attending to the same few regions and ignoring the rest of the image.
To fix this, the authors introduce doubly stochastic regularization: in addition to the row constraint (), they add a penalty encouraging the column sums to also be approximately 1 (). This means every region of the image should be attended to roughly equally over the full caption — no region should be completely ignored.
Intuitively, think of a conference where every speaker (word) must pick audience members (regions) to look at. The first constraint says each speaker's attention must be a valid distribution. The second says no audience member should be totally ignored over the whole conference — everyone gets looked at by someone.
The gating scalar β: how much should the model look at the image?
Not every word in a caption needs visual information. Function words like "a", "the", "of" are determined by grammar, not by the image. The paper introduces a learned gating scalar that modulates the context vector: . When is close to 1, the LSTM relies heavily on the image; when close to 0, it relies on language context alone.
This is a simple but powerful idea: the model learns when to look at the image and when to rely on language patterns. When generating "a", drops — there's nothing to look at. When generating "frisbee", rises — the model needs to see the object.
Training the model
Soft attention training is straightforward: minimize the of the correct caption words, plus the doubly stochastic penalty. Since soft attention is a differentiable weighted average, gradients flow through the attention weights normally.
Hard attention training is more complex. Because sampling a single region is non-differentiable, the authors use the REINFORCE algorithm. The objective becomes a on the , and the gradient is estimated using Monte Carlo sampling. To reduce variance, they use a moving average baseline and multiply the REINFORCE gradient with the reward, then add an term to encourage exploration.
Both models use the same CNN encoder (VGGNet pre-trained on ImageNet), and the decoder is an LSTM with 512 hidden units. The initial hidden and cell states of the LSTM are computed from the mean of the annotation vectors through learned linear projections.
The attention mechanism in code
Simplified to show the idea — not the real implementation.
import numpy as np
def softmax(x):
e = np.exp(x - x.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
def attention(annotations, h_prev, W_a, U_a, v_a):
"""
annotations: (L, D) — 196 region vectors, each 512-dim
h_prev: (H,) — decoder's previous hidden state
W_a: (D, A) — project annotation to alignment dim
U_a: (H, A) — project hidden state to alignment dim
v_a: (A,) — alignment vector
"""
# Stage 1: energy scores — how relevant is each region?
energy = np.tanh(annotations @ W_a + h_prev @ U_a) # (L, A)
energy = energy @ v_a # (L,)
# Stage 2: normalize into attention weights
alpha = softmax(energy) # (L,) sums to 1
# Stage 3: weighted sum = context vector
context = alpha @ annotations # (D,)
return context, alpha
# At each time step, the LSTM receives:
# input = [word_embedding(prev_word); context_vector]
# The context vector changes every step — a new "view" of the image.
# That's the key difference from Show and Tell's fixed single vector.Results and interpretability
The model was evaluated on three standard image captioning benchmarks: Flickr8k, Flickr30k, and MS COCO, using BLEU and METEOR metrics. Both soft and hard attention models outperformed the baseline Show and Tell model that used no attention, validating that attention provides a genuine improvement rather than just added complexity.
But the most striking result wasn't in the numbers — it was in the attention maps. By visualizing weights at each time step, the authors showed that the model genuinely learns to fixate on the correct objects: when generating "bird", the attention highlights the bird; when generating "water", it shifts to the water. This provided the first compelling evidence that attention in neural networks corresponds to something humanly interpretable.
Why it mattered
This paper occupies a pivotal position in the history of visual attention. It took an idea proven in NLP (Bahdanau attention for translation) and showed it works across modalities — from sequences of words to grids of image patches. It demonstrated that attention can serve as an interpretability tool, not just a performance booster. And it established the -with-attention template that would dominate image captioning, VQA, and eventually multi-modal models.
2014
Bahdanau Attention (NLP)
Introduced additive attention for machine translation. The decoder attends to different source words at each step. Directly inspired Show, Attend and Tell.
2015
Show and Tell (no attention)
CNN encoder to single vector, LSTM decoder. No attention — the decoder sees the image only once. Show, Attend and Tell improved on this.
2015
Show, Attend and Tell
This paper. First visual attention for captioning — soft and hard variants, doubly stochastic regularization, interpretable attention maps.
2016
Visual Question Answering (VQA)
Attention extended to answering questions about images. The model attends to image regions relevant to the question, building on the same attention framework.
2017
Transformer (self-attention)
Self-attention replaced recurrence entirely. The conceptual leap from static encoding to dynamic querying — pioneered by visual attention — became the foundation of modern AI.
2021
CLIP (vision + language at scale)
Contrastive learning between images and text at massive scale. Inherited the vision of connecting visual and linguistic representations that Show, Attend and Tell pioneered.
CitationXu, Ba, Kiros, Cho, Courville, Salakhutdinov, Zemel, Bengio. Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. ICML, 2015.
Terms in this paper
- Soft Attentionالانتباه المَرِن
- Attentionآلية الانتباه
- Convolutional Neural Network (CNN)الشبكة العصبية الالتفافية
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Context Vectorمتجه السياق
- Imageالصورة الرقمية
- Feature Mapخريطة السمات
- Attention Weightوزن الانتباه البيني
- Additive Attentionالانتباه الجمعي