RNNs & Sequence Models2014intermediate8 min read
Neural Machine Translation by Jointly Learning to Align and Translate
الترجمة الآلية العصبية عبر التعلم المشترك للمحاذاة والترجمة
Bahdanau, D. · Cho, K. · Bengio, Y. — ICLR
The problem
The standard approach for neural machine translation compresses an entire variable-length source sentence into a single . As sentences grow longer, this forces the to discard information, causing translation quality to collapse on sentences beyond 20–30 words.
The contribution
An mechanism that lets the dynamically look back at all encoder states when generating each target word. A learned scores how relevant each source position is to the current decoding step, producing a soft that replaces the fixed bottleneck with a position-aware recomputed at every step. The encoder is also upgraded to a , creating richer annotations that capture both left and right context.
The impact
The foundational paper that introduced the attention mechanism to sequence-to-sequence models. It directly inspired the Transformer's , Show-Attend-and-Tell for image captioning, and every modern attention variant. Without this paper, there would be no Transformer, no GPT, and no Claude.
Imagine a simultaneous interpreter at a conference. The old approach is like listening to an entire 10-minute speech, summarizing it in one breath, then translating from memory. Details inevitably get lost.
Bahdanau's attention is like giving the interpreter a glass wall to the speaker's notes: while translating each phrase, the interpreter glances back at the relevant part of the original text, focusing on whatever is most helpful right now — a date here, a name there, a technical term further back — instead of relying on a single compressed memory.
The bottleneck: one vector to remember everything
In the Encoder-Decoder model introduced by Cho et al. (2014), an encoder RNN reads a variable-length source sentence and compresses it into a single fixed-length vector . A decoder RNN then generates the translation one word at a time, conditioned on .
The problem becomes vivid with long sentences. Think of as a suitcase: no matter how long your trip, you must fit everything into the same bag. A 5-word sentence fits comfortably; a 50-word sentence overflows. Empirically, the Cho et al. model's BLEU score drops sharply on sentences longer than about 20 words.
The idea: let the decoder look back
Bahdanau, Cho, and Bengio proposed a simple but transformative change. Instead of forcing the entire source sentence through a single vector, let the decoder re-examine the entire source at each generation step — but with a learned preference for the positions that matter most right now.
The mechanism has three stages, like asking a librarian for help:
- Annotate: the encoder produces one () per source word — like placing each word's meaning on an index card.
- Score: when the decoder is about to generate word , an alignment model scores every annotation — like the librarian ranking which cards are relevant to your current question.
- Attend: the scores become weights (via ), and the context vector is the of all annotations — like the librarian handing you a custom summary made from the most relevant cards.
Step 1 — Annotate: the bidirectional encoder
A standard RNN reads left-to-right: the annotation for word summarizes only words . But to translate a word, you often need context from both sides — the word "bank" means something different depending on whether "river" or "account" follows it.
The paper uses a bidirectional RNN (BiRNN). A forward RNN reads the sentence left-to-right, producing hidden states . A backward RNN reads right-to-left, producing . The two are concatenated into a single annotation per word:
Step 2 — Score: the alignment model
Now we need a way to decide: when generating the -th target word, how much should each source annotation contribute? The paper introduces a small feedforward network — called the alignment model — that takes two inputs:
- : the decoder's previous hidden state (what the decoder is "thinking about")
- : the encoder's annotation at source position (what that source word "contains")
The alignment model produces a scalar energy that measures how well matches the current decoding context . Think of it as measuring how loudly each source word "answers" when the decoder asks "who is relevant to what I'm about to say?"
Step 3 — Attend: building the context vector
The energies are raw scores — they could be any real number. To turn them into a proper probability distribution (non-negative, summing to 1), we apply softmax:
The context vector is then the weighted sum of all annotations — each source word contributes its information proportional to its :
The complete architecture
Putting it all together, the decoder generates each target word conditioned on three things:
- : the previous target word (what was just generated)
- : the decoder's hidden state (accumulated decoding context)
- : the attention-derived context vector (what the source says is relevant now)
The decoder then updates its hidden state and produces an output probability:
Seeing alignment in action
One of the most striking results in the paper is the alignment visualization. When the model translates English to French, the attention weights form a near-diagonal pattern — because the two languages have roughly similar word order. But at interesting points, the pattern deviates: the French adjective follows the noun, so the model learns to attend to a word after the current diagonal position. The model discovers grammar on its own.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def tanh(x):
return np.tanh(x)
def softmax(x):
e = np.exp(x - x.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
def bahdanau_attention(s_prev, annotations, W_a, U_a, v_a):
"""
s_prev: decoder hidden state at step i-1 (d_dec,)
annotations: all encoder annotations h_1..h_T (T, d_enc)
W_a: learned weight for decoder state (d_attn, d_dec)
U_a: learned weight for annotations (d_attn, d_enc)
v_a: learned weight to collapse to score (d_attn,)
"""
T = annotations.shape[0]
# 1. Project decoder state and each annotation into alignment space
dec_proj = W_a @ s_prev # (d_attn,)
enc_proj = (U_a @ annotations.T).T # (T, d_attn)
# 2. Combine with tanh → energy scores
energies = np.array([
v_a @ tanh(dec_proj + enc_proj[j]) # scalar per source word
for j in range(T)
]) # (T,)
# 3. Softmax → attention weights
weights = softmax(energies) # (T,) sums to 1
# 4. Weighted sum → context vector
context = weights @ annotations # (d_enc,)
return context, weightsResults and impact
On English-to-French translation (WMT'14), the attention-based model (RNNsearch-50) achieved a BLEU score of 28.45, approaching the phrase-based statistical machine translation system Moses at 33.30 — despite being a purely neural, end-to-end model with no hand-crafted features.
Crucially, the attention model maintained translation quality on long sentences where the baseline RNN Encoder-Decoder collapsed. The gap widened with sentence length: the longer the sentence, the bigger the advantage of attention.
Historical evolution
2014
Seq2Seq (Sutskever et al.)
Stacked LSTMs in an encoder-decoder configuration for machine translation. Used a single fixed-length context vector — the bottleneck this paper addresses.
2014
Bahdanau Attention (this paper)
Introduced the attention mechanism: a learned alignment model that lets the decoder dynamically focus on different source positions. The first soft attention for NMT.
2015
Luong Attention
Simplified Bahdanau's additive scoring to a multiplicative (dot-product) form. Introduced local vs global attention and input-feeding. Faster and equally effective.
2015
Show, Attend and Tell
Applied attention to image captioning — the decoder attends to different regions of a CNN feature map when generating each word. Proved attention generalizes beyond text.
2016
Google Neural Machine Translation (GNMT)
Scaled attention-based NMT to production at Google Translate. Used 8-layer encoder-decoder with attention and achieved near-human quality on several language pairs.
2017
Transformer — Attention Is All You Need
Removed recurrence entirely, replacing it with self-attention. Every position attends to every other position in parallel — enabling the GPU parallelism that made modern LLMs possible.
CitationBahdanau, D., Cho, K., Bengio, Y.. Neural Machine Translation by Jointly Learning to Align and Translate. ICLR 2015, 2014.
Terms in this paper
- Additive Attentionالانتباه الجمعي
- Alignment Modelنموذج المحاذاة
- Soft Attentionالانتباه المَرِن
- Fixed-Length Context Vectorمتجه السياق ثابت الحجم
- Annotationالتمثيل التعليقي
- Energy Scoreدرجة الطاقة
- Bidirectional RNNالشبكة التكرارية ثنائية الاتجاه