Language Models2014intermediate11 min read

Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation

تعلُّم تمثيلات العبارات باستخدام المُرمِّز-فاكّ الترميز التكراري للترجمة الآلية الإحصائية

Cho, K. · van Merriënboer, B. · Gulcehre, C. · Bahdanau, D. · Bougares, F. · Schwenk, H. · Bengio, Y. — EMNLP

The problem

In 2014, (SMT) systems like Moses translated by looking up phrase tables — manually engineered dictionaries of phrase pairs with scores. Neural networks had started showing promise for language, but there was no general-purpose neural architecture that could map a variable-length input sequence to a variable-length output sequence. Standard RNNs suffered from vanishing gradients and couldn't hold onto meaning across long sentences. The field needed both a flexible framework and a recurrent unit that could actually remember.

The contribution

Two contributions in one paper. First, the architecture: one RNN reads an input sequence and compresses it into a fixed-length , a second RNN generates the output sequence from that vector. Both are trained jointly to maximize the of the target given the source. Second, the — a new recurrent cell with two gates (reset and update) that learns when to forget old information and when to pass it through unchanged. The matches LSTM performance with fewer parameters and simpler computation.

The impact

This paper laid the architectural foundation for neural . The - pattern became the blueprint for seq2seq models everywhere — from chatbots to summarization to speech recognition. The GRU became a standard building block, used alongside LSTM in countless systems. Most importantly, this work — together with Sutskever et al. (2014) — proved that pure neural approaches could genuinely compete with decades of hand-engineered SMT pipelines, setting the stage for the mechanism (Bahdanau et al., 2015) and ultimately the .

Imagine a simultaneous interpreter at a conference. She listens to the entire English sentence, holding its meaning in her mind — not a word-for-word transcript, but the gist compressed into a single mental snapshot. Then she speaks the same idea in French, generating each word based on that mental snapshot plus what she's already said.

That mental snapshot is the context vector — the bridge between understanding and generating. The listener side is the encoder; the speaker side is the decoder. And the interpreter's ability to hold important details while discarding noise? That's the GRU — a memory cell with selective gates that decide what to keep and what to let go.

The problem: phrase tables don't generalize

Before neural machine translation, systems like Moses relied on phrase-based statistical machine translation (SMT). The pipeline had three hand-crafted stages: a phrase table mapping source phrases to target phrases with scores, a ensuring fluent output, and a decoder combining both with tuned feature weights.

This worked, but had a fundamental limitation: the phrase table only knew phrases it had seen in data. It couldn't generalize to new word combinations, couldn't capture subtle semantic similarities between phrases, and couldn't learn that "big" and "large" should translate similarly. Every phrase was an isolated lookup — no shared understanding.

Meanwhile, standard recurrent neural networks offered continuous representations but hit their own wall: vanishing gradients. As the RNN processed each word, error signals shrank exponentially when backpropagated through many time steps, making it nearly impossible to learn dependencies between words separated by more than a few positions. An RNN reading "The cat that sat on the mat was happy" would struggle to connect "was" back to "cat."

The big idea: compress, then generate

The RNN Encoder-Decoder has a beautifully simple structure. Two RNNs work in tandem:

The encoder reads the input sequence one at a time, updating its at each step. After the last token, the final hidden state — the context vector c — holds a compressed summary of the entire input. Think of it as squeezing a sentence through a funnel into a single point in continuous space.

The decoder takes that context vector as its initial state and generates the output sequence one token at a time. At each step, it uses its current hidden state, the previously generated word, and the context vector to predict the next word. It keeps generating until it produces a special end-of-sequence token.

Open in Lab
Click "Encode" to watch the encoder compress the input into a context vector, then click "Decode" to see the decoder generate the output one word at a time.
The demo wakes as you arrive…
p(y1,…,yT∣x1,…,xT′)=∏t=1Tp(yt∣y1,…,yt−1,c)p(y_1, \ldots, y_T \mid x_1, \ldots, x_{T'}) = \prod_{t=1}^{T} p(y_t \mid y_1, \ldots, y_{t-1}, \mathbf{c})
The encoder-decoder objective — factored conditional probability — The probability of the entire output sequence given the input is the product of each output word's probability, conditioned on all previously generated words and the context vector c. Training maximizes this joint probability.

Read the formula as a conveyor belt: at each time step tt, the decoder looks at what it has produced so far (y1y_1 through yt−1y_{t-1}), glances at the compressed input (c\mathbf{c}), and picks the most likely next word. The encoder and decoder are trained jointly — gradients flow from the decoder's word predictions all the way back through the context vector into the encoder, teaching the encoder what kind of summary the decoder actually needs.

The GRU: a simpler way to remember

The paper's second contribution is the Gated Recurrent Unit (GRU) — a new type of recurrent cell designed to capture long-term dependencies without the complexity of LSTM.

The key insight: instead of LSTM's three gates (input, forget, output) and a separate cell state, the GRU uses just two gates and merges the cell state into the hidden state. Fewer moving parts, fewer parameters, faster training — yet comparable performance.

Think of the GRU as a smart notebook with two controls. The decides how much of the previous notes to erase before writing new ones — like a "clean slate" dial. The decides how much of the old page to keep versus how much to replace with fresh content — like a "keep vs. replace" slider.

Open in Lab
Drag the sliders to change the reset and update gate values and see how they control memory flow through the GRU cell.
The demo wakes as you arrive…

Here's how the GRU computes its new hidden state step by step:

Step 1 — Reset gate rtr_t: Look at the current input xtx_t and the previous hidden state ht−1h_{t-1}. Produce a value between 0 and 1 for each dimension. Where rt≈0r_t \approx 0, the previous state is erased; where rt≈1r_t \approx 1, it's fully visible.

Step 2 — Candidate state h~t\tilde{h}_t: Compute a fresh proposal for the new state, but using the reset-filtered version of the previous state. This lets the GRU "start fresh" when the reset gate is closed.

Step 3 — Update gate ztz_t: Another 0-to-1 value per dimension. This is the dial: the final hidden state is a weighted mix of the old state and the candidate. Where zt≈1z_t \approx 1, keep the old state (skip the current input); where zt≈0z_t \approx 0, fully replace with the candidate.

rt=σ(Wrxt+Urht−1)r_t = \sigma(W_r x_t + U_r h_{t-1})
Reset gate — controls erasure of previous state — σ is the sigmoid function (output between 0 and 1). W_r and U_r are learned weight matrices. When r_t → 0, the previous hidden state is erased before computing the candidate.
h~t=tanh⁡(Wxt+U(rt⊙ht−1))\tilde{h}_t = \tanh(W x_t + U(r_t \odot h_{t-1}))
Candidate hidden state — a fresh proposal filtered by the reset gate — ⊙ means element-wise multiplication. The reset gate masks certain dimensions of the old state, letting the candidate be computed from a "cleaned" version. tanh squashes the result to [-1, 1].
zt=σ(Wzxt+Uzht−1)z_t = \sigma(W_z x_t + U_z h_{t-1})
Update gate — the keep-vs-replace dial — Same structure as the reset gate but with its own weights. High z means "keep the old state" — the unit acts as a long-term memory. Low z means "use the new candidate."
ht=zt⊙ht−1+(1−zt)⊙h~th_t = z_t \odot h_{t-1} + (1 - z_t) \odot \tilde{h}_t
Final hidden state — interpolation between old and new — This is the core of the GRU. The update gate smoothly blends the previous state with the candidate. No separate cell state needed — the hidden state IS the memory.

GRU vs LSTM: same goal, different mechanics

Open in Lab
Compare the internal structure of GRU and LSTM side by side. Click each gate to see its role.
The demo wakes as you arrive…

The fundamental difference: LSTM maintains a separate cell state ctc_t that runs like a conveyor belt alongside the hidden state, with gates controlling what goes on and off the belt. GRU merges cell state and hidden state into one, using the update gate as the only mixing control. This makes GRU's information flow simpler — fewer pathways for gradients to traverse — which is why it trains faster on many tasks.

Neither is universally better. LSTM tends to excel on tasks requiring very long memory (thousands of steps) thanks to its dedicated cell state. GRU tends to shine on tasks with moderate sequence lengths and smaller datasets, where its efficiency matters more.

How the encoder-decoder boosted translation

Cho et al. didn't replace SMT entirely — they made it smarter. The RNN Encoder-Decoder was trained on English-French phrase pairs, learning to assign a conditional probability p(target phrase∣source phrase)p(\text{target phrase} \mid \text{source phrase}) to each pair. These probabilities were then added as extra features in the existing SMT system's log-linear model.

The result: the SMT system with RNN Encoder-Decoder scores outperformed the on BLEU by a significant margin. The neural model captured semantic similarities that the phrase table couldn't — phrases with similar meanings received similar scores even if they'd never appeared together in training data.

Open in Lab
Explore how the encoder-decoder maps phrase pairs into continuous space. Similar phrases cluster together even if they never co-occurred in training.
The demo wakes as you arrive…

The bottleneck problem: one vector to rule them all?

The encoder-decoder's elegant simplicity is also its Achilles' heel. The entire input — whether 5 words or 50 — must be compressed into a single fixed-length vector. As sentences grow longer, this bottleneck loses information. The decoder has no way to "look back" at individual input words; it only sees the compressed summary.

Cho et al. themselves noted this limitation: performance degraded significantly on sentences longer than about 20-30 words. This compression problem was the direct motivation for Bahdanau et al.'s attention mechanism in 2015, which lets the decoder attend to different parts of the input at each decoding step rather than relying on a single static summary.

Open in Lab
Increase the sentence length and watch the context vector struggle to hold all the information — the reconstruction quality drops with length.
The demo wakes as you arrive…

The idea in code

GRU cell and Encoder-Decoder, from scratchpython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(x):
    return 1 / (1 + np.exp(-np.clip(x, -15, 15)))

def gru_cell(x_t, h_prev, W_r, U_r, W_z, U_z, W_h, U_h):
    """One GRU step: takes input x_t and previous state h_prev."""
    r_t = sigmoid(W_r @ x_t + U_r @ h_prev)       # reset gate
    z_t = sigmoid(W_z @ x_t + U_z @ h_prev)       # update gate
    h_tilde = np.tanh(W_h @ x_t + U_h @ (r_t * h_prev))  # candidate
    h_t = z_t * h_prev + (1 - z_t) * h_tilde      # final state
    return h_t

def encode(sequence, W_r, U_r, W_z, U_z, W_h, U_h):
    """Read a sequence and return the final hidden state (context vector)."""
    h = np.zeros(W_r.shape[0])            # initial hidden state
    for x_t in sequence:
        h = gru_cell(x_t, h, W_r, U_r, W_z, U_z, W_h, U_h)
    return h  # this IS the context vector c

def decode_step(y_prev, h_prev, c, W_r, U_r, W_z, U_z, W_h, U_h):
    """One decoding step: input is [y_prev; c], updating hidden state."""
    x_t = np.concatenate([y_prev, c])     # decoder input = prev word + context
    h_t = gru_cell(x_t, h_prev, W_r, U_r, W_z, U_z, W_h, U_h)
    return h_t

# The full pipeline:
# 1. Encoder reads source → context vector c
# 2. Decoder generates target words one at a time from c
# 3. Both trained jointly via backpropagation through time

What the model learned

Beyond translation scores, Cho et al. showed something remarkable: the encoder-decoder learned meaningful representations of phrases. When they visualized the hidden representations of phrase pairs using 2D projections, semantically similar phrases clustered together naturally.

Phrases like "was given" and "was awarded" ended up near each other in the learned space, even though they're different words. Phrases describing temporal events grouped together. The model discovered the structure of language without being told what structure to look for — an early glimpse of that would later become the foundation of embeddings and pre-trained models.

Why it changed everything

  1. 2014

    This paper (Cho et al.)

    Introduced the RNN Encoder-Decoder framework and the GRU. Used neural phrase scores to boost SMT — proved neural models add value to translation.

  2. 2014

    Sequence to Sequence (Sutskever et al.)

    Scaled the encoder-decoder idea with deep LSTMs and showed end-to-end neural translation could rival SMT — no phrase table needed at all.

  3. 2015

    Attention Mechanism (Bahdanau et al.)

    Solved the bottleneck problem by letting the decoder attend to all encoder states — not just the final one. This was the direct successor of this paper.

  4. 2017

    Transformer (Vaswani et al.)

    Replaced recurrence entirely with self-attention, but kept the encoder-decoder structure. The architectural DNA of this paper lives on in every modern LLM.

CitationCho, van Merriënboer, Gulcehre, Bahdanau, Bougares, Schwenk, Bengio. Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. EMNLP, 2014.

Terms in this paper