RNNs & Sequence Models2014intermediate8 min read

Neural Machine Translation by Jointly Learning to Align and Translate

الترجمة الآلية العصبية عبر التعلم المشترك للمحاذاة والترجمة

Bahdanau, D. · Cho, K. · Bengio, Y. — ICLR

The problem

The standard approach for neural machine translation compresses an entire variable-length source sentence into a single . As sentences grow longer, this forces the to discard information, causing translation quality to collapse on sentences beyond 20–30 words.

The contribution

An mechanism that lets the dynamically look back at all encoder states when generating each target word. A learned scores how relevant each source position is to the current decoding step, producing a soft that replaces the fixed bottleneck with a position-aware recomputed at every step. The encoder is also upgraded to a , creating richer annotations that capture both left and right context.

The impact

The foundational paper that introduced the attention mechanism to sequence-to-sequence models. It directly inspired the Transformer's , Show-Attend-and-Tell for image captioning, and every modern attention variant. Without this paper, there would be no Transformer, no GPT, and no Claude.

Imagine a simultaneous interpreter at a conference. The old approach is like listening to an entire 10-minute speech, summarizing it in one breath, then translating from memory. Details inevitably get lost.

Bahdanau's attention is like giving the interpreter a glass wall to the speaker's notes: while translating each phrase, the interpreter glances back at the relevant part of the original text, focusing on whatever is most helpful right now — a date here, a name there, a technical term further back — instead of relying on a single compressed memory.

The bottleneck: one vector to remember everything

In the Encoder-Decoder model introduced by Cho et al. (2014), an encoder RNN reads a variable-length source sentence and compresses it into a single fixed-length vector cc. A decoder RNN then generates the translation one word at a time, conditioned on cc.

The problem becomes vivid with long sentences. Think of cc as a suitcase: no matter how long your trip, you must fit everything into the same bag. A 5-word sentence fits comfortably; a 50-word sentence overflows. Empirically, the Cho et al. model's BLEU score drops sharply on sentences longer than about 20 words.

Open in Lab
Watch information get squeezed into a fixed-length vector. Then toggle to see how attention preserves access to every source word.
The demo wakes as you arrive…

The idea: let the decoder look back

Bahdanau, Cho, and Bengio proposed a simple but transformative change. Instead of forcing the entire source sentence through a single vector, let the decoder re-examine the entire source at each generation step — but with a learned preference for the positions that matter most right now.

The mechanism has three stages, like asking a librarian for help:

  • Annotate: the encoder produces one () per source word — like placing each word's meaning on an index card.
  • Score: when the decoder is about to generate word ii, an alignment model scores every annotation — like the librarian ranking which cards are relevant to your current question.
  • Attend: the scores become weights (via ), and the context vector is the of all annotations — like the librarian handing you a custom summary made from the most relevant cards.

Step 1 — Annotate: the bidirectional encoder

A standard RNN reads left-to-right: the annotation for word jj summarizes only words 1,…,j1, \ldots, j. But to translate a word, you often need context from both sides — the word "bank" means something different depending on whether "river" or "account" follows it.

The paper uses a bidirectional RNN (BiRNN). A forward RNN reads the sentence left-to-right, producing hidden states h1→,…,hT→\overrightarrow{h_1}, \ldots, \overrightarrow{h_T}. A backward RNN reads right-to-left, producing h1←,…,hT←\overleftarrow{h_1}, \ldots, \overleftarrow{h_T}. The two are concatenated into a single annotation per word:

hj=[hj→⊤;hj←⊤]⊤h_j = \left[\overrightarrow{h_j}^{\top} ; \overleftarrow{h_j}^{\top}\right]^{\top}
Bidirectional annotation — each word knows its full neighborhood — Concatenating forward and backward hidden states gives each annotation a summary of the words surrounding position j, not just the words before it.
Open in Lab
Watch the forward (→) and backward (←) passes build annotations. Each annotation captures full sentence context.
The demo wakes as you arrive…

Step 2 — Score: the alignment model

Now we need a way to decide: when generating the ii-th target word, how much should each source annotation hjh_j contribute? The paper introduces a small feedforward network — called the alignment model — that takes two inputs:

  • si−1s_{i-1}: the decoder's previous hidden state (what the decoder is "thinking about")
  • hjh_j: the encoder's annotation at source position jj (what that source word "contains")

The alignment model produces a scalar energy eije_{ij} that measures how well hjh_j matches the current decoding context si−1s_{i-1}. Think of it as measuring how loudly each source word "answers" when the decoder asks "who is relevant to what I'm about to say?"

eij=va⊤tanh⁡(Wa si−1+Ua hj)e_{ij} = v_a^{\top} \tanh(W_a \, s_{i-1} + U_a \, h_j)
Additive attention energy — the heart of the paper — This score measures how relevant a particular input position is to the current decoding step. The attention mechanism combines information from the decoder's current state and the encoded representation of an input token, then produces a single compatibility score. Higher scores indicate that the decoder should pay more attention to that part of the input when generating the next output token. All components of this scoring function are learned automatically during training.
Open in Lab
Step through the alignment model: see how W_a, U_a, and v_a transform two vectors into one relevance score.
The demo wakes as you arrive…

Step 3 — Attend: building the context vector

The energies eije_{ij} are raw scores — they could be any real number. To turn them into a proper probability distribution (non-negative, summing to 1), we apply softmax:

αij=exp⁡(eij)∑k=1Texp⁡(eik)\alpha_{ij} = \frac{\exp(e_{ij})}{\sum_{k=1}^{T} \exp(e_{ik})}
Attention weights — a soft selection over source positions — These weights determine how much attention the model gives to each input position when generating the current output token. The attention mechanism converts relevance scores into a probability distribution, assigning larger weights to the most relevant parts of the input while still allowing other positions to contribute. All attention weights for a given decoding step add up to one, making them easy to interpret as a distribution of focus.

The context vector cic_i is then the weighted sum of all annotations — each source word contributes its information proportional to its :

ci=∑j=1Tαij hjc_i = \sum_{j=1}^{T} \alpha_{ij} \, h_j
Context vector — a custom summary for each target word — Earlier encoder-decoder models relied on a single fixed summary of the entire input sequence. Attention replaces this with a dynamic context that is recomputed at every decoding step. As a result, different output words can focus on different parts of the input, allowing the model to access the most relevant information whenever it is needed.
Open in Lab
Click a target word to see how its context vector is assembled from weighted source annotations.
The demo wakes as you arrive…

The complete architecture

Putting it all together, the decoder generates each target word yiy_i conditioned on three things:

  • yi−1y_{i-1}: the previous target word (what was just generated)
  • si−1s_{i-1}: the decoder's hidden state (accumulated decoding context)
  • cic_i: the attention-derived context vector (what the source says is relevant now)

The decoder then updates its hidden state and produces an output probability:

si=f(si−1,  yi−1,  ci)p(yi∣y<i,x)=g(yi−1,  si,  ci)\begin{aligned} s_i &= f(s_{i-1},\; y_{i-1},\; c_i) \\[4pt] p(y_i \mid y_{<i}, \mathbf{x}) &= g(y_{i-1},\; s_i,\; c_i) \end{aligned}
Decoder with attention — the full generation step — At each generation step, the decoder updates its internal state using the previously generated output and the current attention-based context. It then uses this updated state together with the context information to predict the next word. By repeatedly updating its state and refreshing its focus on the input sequence, the model can generate coherent outputs while remaining aligned with the relevant parts of the source sentence.
Open in Lab
Click any component to understand its role in the encoder-decoder-attention pipeline.
The demo wakes as you arrive…

Seeing alignment in action

One of the most striking results in the paper is the alignment visualization. When the model translates English to French, the attention weights form a near-diagonal pattern — because the two languages have roughly similar word order. But at interesting points, the pattern deviates: the French adjective follows the noun, so the model learns to attend to a word after the current diagonal position. The model discovers grammar on its own.

Open in Lab
Hover over the heatmap to see attention weights. Notice the near-diagonal pattern and where it breaks due to word reordering.
The demo wakes as you arrive…

The same idea in code

Bahdanau additive attention, completepython

Simplified to show the idea — not the real implementation.

import numpy as np

def tanh(x):
    return np.tanh(x)

def softmax(x):
    e = np.exp(x - x.max(axis=-1, keepdims=True))
    return e / e.sum(axis=-1, keepdims=True)

def bahdanau_attention(s_prev, annotations, W_a, U_a, v_a):
    """
    s_prev:      decoder hidden state at step i-1   (d_dec,)
    annotations: all encoder annotations h_1..h_T   (T, d_enc)
    W_a:         learned weight for decoder state    (d_attn, d_dec)
    U_a:         learned weight for annotations      (d_attn, d_enc)
    v_a:         learned weight to collapse to score (d_attn,)
    """
    T = annotations.shape[0]

    # 1. Project decoder state and each annotation into alignment space
    dec_proj = W_a @ s_prev                   # (d_attn,)
    enc_proj = (U_a @ annotations.T).T        # (T, d_attn)

    # 2. Combine with tanh → energy scores
    energies = np.array([
        v_a @ tanh(dec_proj + enc_proj[j])    # scalar per source word
        for j in range(T)
    ])                                        # (T,)

    # 3. Softmax → attention weights
    weights = softmax(energies)               # (T,) sums to 1

    # 4. Weighted sum → context vector
    context = weights @ annotations           # (d_enc,)

    return context, weights

Results and impact

On English-to-French translation (WMT'14), the attention-based model (RNNsearch-50) achieved a BLEU score of 28.45, approaching the phrase-based statistical machine translation system Moses at 33.30 — despite being a purely neural, end-to-end model with no hand-crafted features.

Crucially, the attention model maintained translation quality on long sentences where the baseline RNN Encoder-Decoder collapsed. The gap widened with sentence length: the longer the sentence, the bigger the advantage of attention.

Historical evolution

  1. 2014

    Seq2Seq (Sutskever et al.)

    Stacked LSTMs in an encoder-decoder configuration for machine translation. Used a single fixed-length context vector — the bottleneck this paper addresses.

  2. 2014

    Bahdanau Attention (this paper)

    Introduced the attention mechanism: a learned alignment model that lets the decoder dynamically focus on different source positions. The first soft attention for NMT.

  3. 2015

    Luong Attention

    Simplified Bahdanau's additive scoring to a multiplicative (dot-product) form. Introduced local vs global attention and input-feeding. Faster and equally effective.

  4. 2015

    Show, Attend and Tell

    Applied attention to image captioning — the decoder attends to different regions of a CNN feature map when generating each word. Proved attention generalizes beyond text.

  5. 2016

    Google Neural Machine Translation (GNMT)

    Scaled attention-based NMT to production at Google Translate. Used 8-layer encoder-decoder with attention and achieved near-human quality on several language pairs.

  6. 2017

    Transformer — Attention Is All You Need

    Removed recurrence entirely, replacing it with self-attention. Every position attends to every other position in parallel — enabling the GPU parallelism that made modern LLMs possible.

CitationBahdanau, D., Cho, K., Bengio, Y.. Neural Machine Translation by Jointly Learning to Align and Translate. ICLR 2015, 2014.

Terms in this paper