Computer Vision2015intermediate12 min read

Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

أَرِني، انتبِه وأخبرني: توليد وصف الصور بالانتباه البصري

Xu, K. · Ba, J. · Kiros, R. · Cho, K. · Courville, A. · Salakhutdinov, R. · Zemel, R. · Bengio, Y. — ICML

The problem

Early models (like Show and Tell) encoded the entire into a single fixed-length vector using a , then fed that vector to an to generate a caption. This forced all visual information — objects, relationships, background — through one narrow channel. The LSTM had to remember everything from that single snapshot, making it difficult to generate accurate descriptions of complex scenes with multiple objects.

The contribution

An mechanism for image captioning that lets the LSTM "look back" at the image at every time step. Instead of a single vector, a CNN produces a grid of spatial feature vectors (annotation vectors). At each word, an attention module scores every grid location, and the decoder receives a weighted combination of those features — focusing on the relevant image region. Two variants are proposed: deterministic (differentiable, trained with ) and stochastic hard attention (samples one region, trained with REINFORCE). A encourages the model to attend to the full image over the course of a caption.

The impact

This paper was the first to bring to image captioning, demonstrating that neural networks can learn where to look — and show us where they looked. The attention maps provided an unprecedented window into the model's reasoning, making it a landmark for interpretable AI. It directly inspired attention mechanisms across computer vision, from VQA to dense captioning, and laid the conceptual groundwork for vision-language models like CLIP.

Imagine you're a tour guide describing a painting to a visitor over the phone. The old approach was like glancing at the painting once, memorizing everything, then talking for two minutes from memory. You'd get the gist right but mix up details — was the dog on the left or the right? Was there a boat in the background?

Show, Attend and Tell replaces that with a pointer and a magnifying glass: as you say each word, you slide the magnifying glass to the part of the painting that word describes. Saying "dog"? You zoom into the dog. Saying "frisbee"? The glass slides to the frisbee. You never have to hold the whole painting in your head — you just look at what matters right now.

The bottleneck: one vector for an entire image

Before this paper, image captioning followed a simple two-stage pipeline:

  1. A convolutional neural network (typically VGGNet or GoogLeNet) processes the image and outputs a single — a fixed-length numerical summary of the entire scene.

  2. An LSTM decoder receives that vector once, at the first time step, then generates the caption word by word purely from its .

The problem is a bottleneck: thousands of pixels, dozens of objects, and all their spatial relationships must compress into one vector of 512 or 4096 numbers. Early words get the freshest signal; later words rely on the LSTM's fading memory of that initial snapshot. As scenes grow more complex, the descriptions grow more generic — the model says "a man on a field" instead of "a football player catching a ball near the goal post".

Open in Lab
Left: the old pipeline squeezes the image through one vector. Right: attention lets the decoder look back at different regions for each word.
The demo wakes as you arrive…

The idea: let the decoder look back at the image

The key insight is borrowed from Bahdanau's attention for : instead of encoding the image into a single vector, keep a grid of feature vectors and let the decoder choose which grid cells to focus on at each time step. The architecture has three parts:

Encoder (CNN): Extract features not from the final fully connected layer, but from a lower convolutional layer — specifically the fourth convolutional layer of VGGNet. This produces L=14×14=196L = 14 \times 14 = 196 annotation vectors, each of dimension D=512D = 512. Each vector encodes information about a particular spatial region of the image. Think of it as a 14×14 grid overlay on the image, where each cell contains a rich description of what's in that cell.

Attention mechanism: At each time step, a small neural network scores how relevant each of the 196 locations is, given the decoder's current state. The scores become weights (via ), and the weighted sum of annotation vectors produces a — a dynamic summary that highlights the part of the image relevant to the current word.

Decoder (LSTM): Generates one word at a time, conditioned on the previous word, the previous hidden state, and the context vector from the attention mechanism. This means the decoder's input changes at every time step — it sees a different "view" of the image depending on what word it's about to generate.

Open in Lab
Click on any component to see its role in the pipeline.
The demo wakes as you arrive…

How attention works: scoring, weighting, and focusing

At each time step tt, the model must decide which of the L=196L = 196 image regions to attend to. It does this in three stages:

Stage 1 — Energy score: A small alignment network takes each aia_i and the previous hidden state ht−1h_{t-1}, and produces a scalar energy etie_{ti} measuring how relevant region ii is to the current decoding step. This is — the same mechanism Bahdanau used for translation.

Stage 2 — Attention weights: The energies are normalized via softmax so they sum to 1: αti\alpha_{ti} is the probability that location ii is the right place to focus on when generating word tt.

Stage 3 — Context vector: The context vector z^t\hat{z}_t is a weighted combination of annotation vectors. This single vector — a dynamic, word-specific summary of the image — is fed to the LSTM along with the previous word .

eti=fatt(ai,ht−1)=va⊤tanh⁡(Waai+Uaht−1)e_{ti} = f_{att}(a_i, h_{t-1}) = v_a^\top \tanh(W_a a_i + U_a h_{t-1})
Energy score — how relevant is region i to the current word? — aia_i = annotation vector for region i · ht−1h_{t-1} = decoder's previous hidden state · WaW_a, UaU_a, vav_a = learned parameters of the alignment network
αti=exp⁡(eti)∑k=1Lexp⁡(etk)\alpha_{ti} = \frac{\exp(e_{ti})}{\sum_{k=1}^{L} \exp(e_{tk})}
Attention weight — normalized focus probability — Softmax over all L=196 locations · each αti\alpha_{ti} sums to 1 over locations · higher weight = more focus on that region
z^t=ϕ({ai},{αi})=∑i=1Lαtiai\hat{z}_t = \phi(\{a_i\}, \{\alpha_i\}) = \sum_{i=1}^{L} \alpha_{ti} a_i
Context vector — the word-specific image summary — A weighted average of all annotation vectors · changes at every time step · the decoder sees a new "view" of the image for each word it generates
Open in Lab
Step through the attention mechanism: watch how different regions light up as the model generates each word.
The demo wakes as you arrive…

Two flavors: soft attention vs. hard attention

The paper proposes two variants of the attention mechanism, each with different tradeoffs:

Soft (deterministic) attention takes the weighted average of all annotation vectors. Every region contributes, but regions with higher α\alpha weights contribute more. This is fully differentiable — you can train it end-to-end with standard backpropagation. Think of it as looking at the whole painting through a filter that makes relevant areas brighter and irrelevant areas dimmer.

Hard (stochastic) attention samples one region according to the α\alpha distribution. At each time step, the model picks exactly one grid location sts_t and uses only asta_{s_t} as the context. This is like putting a spotlight on one part of the painting and ignoring everything else. Because sampling is not differentiable, hard attention must be trained with the REINFORCE algorithm — a method from that treats the attention decisions as a sequence of actions.

The soft variant is simpler to train and more stable; the hard variant can sometimes achieve slightly higher performance because it forces sharper focus, but at the cost of higher during .

Open in Lab
Toggle between soft and hard attention to see how each selects image regions. Soft blends all regions; hard picks one.
The demo wakes as you arrive…

Doubly stochastic attention: don't ignore any part of the image

By construction, the attention weights αti\alpha_{ti} at each time step sum to 1 over locations (each word's focus is a valid probability distribution). But nothing prevents the model from always attending to the same few regions and ignoring the rest of the image.

To fix this, the authors introduce doubly stochastic regularization: in addition to the row constraint (∑iαti=1\sum_i \alpha_{ti} = 1), they add a penalty encouraging the column sums to also be approximately 1 (∑tαti≈1\sum_t \alpha_{ti} \approx 1). This means every region of the image should be attended to roughly equally over the full caption — no region should be completely ignored.

Intuitively, think of a conference where every speaker (word) must pick audience members (regions) to look at. The first constraint says each speaker's attention must be a valid distribution. The second says no audience member should be totally ignored over the whole conference — everyone gets looked at by someone.

Ld=λ∑i=1L(1−∑t=1Cαti)2L_d = \lambda \sum_{i=1}^{L} \left(1 - \sum_{t=1}^{C} \alpha_{ti}\right)^2
Doubly stochastic penalty — encourages full coverage — Added to the main loss · CC = caption length · LL = number of image regions · penalizes regions whose total attention over the caption deviates from 1
Open in Lab
Watch how doubly stochastic regularization distributes attention across all image regions over the course of a caption.
The demo wakes as you arrive…

The gating scalar β: how much should the model look at the image?

Not every word in a caption needs visual information. Function words like "a", "the", "of" are determined by grammar, not by the image. The paper introduces a learned gating scalar βt=σ(fβ(ht−1))\beta_t = \sigma(f_\beta(h_{t-1})) that modulates the context vector: z^t=βt∑iαtiai\hat{z}_t = \beta_t \sum_i \alpha_{ti} a_i. When β\beta is close to 1, the LSTM relies heavily on the image; when close to 0, it relies on language context alone.

This is a simple but powerful idea: the model learns when to look at the image and when to rely on language patterns. When generating "a", β\beta drops — there's nothing to look at. When generating "frisbee", β\beta rises — the model needs to see the object.

Training the model

Soft attention training is straightforward: minimize the of the correct caption words, plus the doubly stochastic penalty. Since soft attention is a differentiable weighted average, gradients flow through the attention weights normally.

Hard attention training is more complex. Because sampling a single region is non-differentiable, the authors use the REINFORCE algorithm. The objective becomes a on the , and the gradient is estimated using Monte Carlo sampling. To reduce variance, they use a moving average baseline and multiply the REINFORCE gradient with the reward, then add an term to encourage exploration.

Both models use the same CNN encoder (VGGNet pre-trained on ImageNet), and the decoder is an LSTM with 512 hidden units. The initial hidden and cell states of the LSTM are computed from the mean of the annotation vectors through learned linear projections.

The attention mechanism in code

Soft attention for image captioning, step by steppython

Simplified to show the idea — not the real implementation.

import numpy as np

def softmax(x):
    e = np.exp(x - x.max(axis=-1, keepdims=True))
    return e / e.sum(axis=-1, keepdims=True)

def attention(annotations, h_prev, W_a, U_a, v_a):
    """
    annotations: (L, D) — 196 region vectors, each 512-dim
    h_prev:      (H,)   — decoder's previous hidden state
    W_a:         (D, A) — project annotation to alignment dim
    U_a:         (H, A) — project hidden state to alignment dim
    v_a:         (A,)   — alignment vector
    """
    # Stage 1: energy scores — how relevant is each region?
    energy = np.tanh(annotations @ W_a + h_prev @ U_a)  # (L, A)
    energy = energy @ v_a                                 # (L,)

    # Stage 2: normalize into attention weights
    alpha = softmax(energy)                               # (L,) sums to 1

    # Stage 3: weighted sum = context vector
    context = alpha @ annotations                         # (D,)

    return context, alpha

# At each time step, the LSTM receives:
#   input = [word_embedding(prev_word); context_vector]
# The context vector changes every step — a new "view" of the image.
# That's the key difference from Show and Tell's fixed single vector.

Results and interpretability

The model was evaluated on three standard image captioning benchmarks: Flickr8k, Flickr30k, and MS COCO, using BLEU and METEOR metrics. Both soft and hard attention models outperformed the baseline Show and Tell model that used no attention, validating that attention provides a genuine improvement rather than just added complexity.

But the most striking result wasn't in the numbers — it was in the attention maps. By visualizing α\alpha weights at each time step, the authors showed that the model genuinely learns to fixate on the correct objects: when generating "bird", the attention highlights the bird; when generating "water", it shifts to the water. This provided the first compelling evidence that attention in neural networks corresponds to something humanly interpretable.

Open in Lab
Click through each word to see where the model focuses its attention on the image grid.
The demo wakes as you arrive…

Why it mattered

This paper occupies a pivotal position in the history of visual attention. It took an idea proven in NLP (Bahdanau attention for translation) and showed it works across modalities — from sequences of words to grids of image patches. It demonstrated that attention can serve as an interpretability tool, not just a performance booster. And it established the -with-attention template that would dominate image captioning, VQA, and eventually multi-modal models.

  1. 2014

    Bahdanau Attention (NLP)

    Introduced additive attention for machine translation. The decoder attends to different source words at each step. Directly inspired Show, Attend and Tell.

  2. 2015

    Show and Tell (no attention)

    CNN encoder to single vector, LSTM decoder. No attention — the decoder sees the image only once. Show, Attend and Tell improved on this.

  3. 2015

    Show, Attend and Tell

    This paper. First visual attention for captioning — soft and hard variants, doubly stochastic regularization, interpretable attention maps.

  4. 2016

    Visual Question Answering (VQA)

    Attention extended to answering questions about images. The model attends to image regions relevant to the question, building on the same attention framework.

  5. 2017

    Transformer (self-attention)

    Self-attention replaced recurrence entirely. The conceptual leap from static encoding to dynamic querying — pioneered by visual attention — became the foundation of modern AI.

  6. 2021

    CLIP (vision + language at scale)

    Contrastive learning between images and text at massive scale. Inherited the vision of connecting visual and linguistic representations that Show, Attend and Tell pioneered.

CitationXu, Ba, Kiros, Cho, Courville, Salakhutdinov, Zemel, Bengio. Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. ICML, 2015.

Terms in this paper