RNNs & Sequence Models2014intermediate12 min read
Sequence to Sequence Learning with Neural Networks
التعلُّم من تسلسل إلى تسلسل باستخدام الشبكات العصبية
Sutskever, I. · Vinyals, O. · Le, Q.V. — NeurIPS
The problem
Deep Neural Networks (DNNs) excel at fixed-size inputs and outputs, but most real-world problems — translation, summarization, dialogue — involve variable-length sequences whose lengths are not known in advance. Traditional phrase-based statistical (SMT) systems were complex pipelines of hand-engineered components. There was no simple, general neural approach that maps one sequence directly to another.
The contribution
A general end-to-end sequence learning architecture using two multilayer LSTMs: an that reads the source sequence and compresses it into a single fixed-dimensional (the last ), and a that generates the target sequence one token at a time from that vector. A key trick: reversing the source sentence order, which places corresponding words closer together and dramatically improves learning. On English-to-French translation, the system achieved 34.81 BLEU — surpassing a phrase-based SMT baseline and approaching state-of-the-art when used to re-rank n-best lists.
The impact
The first convincing demonstration that a pure neural network — with no hand-engineered features — could rival statistical machine translation. It introduced the paradigm that became the backbone of all modern systems. Every subsequent milestone — mechanisms, the , GPT, BART — built directly on this architecture. It also inspired applications far beyond translation: image captioning (Show and Tell), speech synthesis (Tacotron), and conversational AI.
Imagine a simultaneous interpreter at a UN conference. She listens to the entire speech in French — absorbing meaning, intent, nuance — until the speaker finishes. Only then does she begin rendering it in English, one sentence at a time, from the compressed understanding she built in her mind.
She cannot go back and re-listen. Everything she needs must be in that mental summary. If the speech was short, her summary is vivid. If it was long, details start to blur — this is the of a fixed-size mental representation.
The Seq2Seq model works the same way: one listens to the full input and compresses it into a single vector, then a second LSTM speaks the output from that vector alone.
The problem: DNNs cannot handle variable-length sequences
By 2014, deep neural networks had conquered image classification, speech recognition, and many fixed-size tasks. But the standard DNN takes a fixed-size input and produces a fixed-size output. Real language doesn't work that way: a five-word English sentence might translate into eight French words.
The dominant solution — phrase-based statistical machine translation — was a pipeline of separately engineered pieces: an alignment model, a language model, a reordering model, phrase tables, and a decoder. Each piece was designed by hand. Each needed its own data and tuning. There was no simple way to train the whole thing end-to-end.
What if a single neural network could read an arbitrary-length input and write an arbitrary-length output?
The idea: two LSTMs — one listens, one speaks
Sutskever, Vinyals, and Le proposed a remarkably simple architecture. Split the problem into two halves:
The Encoder — a multilayer LSTM that reads the source sentence one word at a time. At each step it updates its hidden state, building an increasingly rich internal summary. After the last word, its final hidden state becomes a single fixed-dimensional vector — the context vector — that encodes the meaning of the entire input sentence.
The Decoder — a second multilayer LSTM, initialized with the context vector as its starting hidden state. It generates the target sentence one word at a time: at each step it receives the previously generated word and its own hidden state, predicts a probability distribution over the using , and emits the most likely next word. A special end-of-sentence token signals the decoder to stop.
The entire system is trained end-to-end to maximize the probability of the correct translation given the source sentence.
The formal objective: maximize conditional probability
The goal is to estimate the conditional probability of a target sequence given a source sequence , where the two lengths and may differ. The encoder reads the source and produces a context vector , then the decoder models:
In practice, training maximizes the log-probability of the correct target sequence, which is equivalent to minimizing the cross-entropy loss. Each word prediction is a classification over the entire vocabulary (80,000 words for French in this paper).
The bottleneck: cramming meaning into a fixed-size vector
The context vector is the only link between encoder and decoder. Think of it as a narrow bridge between two islands: every bit of meaning from the source sentence must cross this bridge. The bridge has a fixed width (1,000 dimensions in this paper), whether the sentence is 5 words or 50.
For short sentences, the bridge is wide enough. For long sentences, the model must squeeze more meaning through the same opening — and inevitably loses detail. This bottleneck is the architecture's greatest limitation, and is precisely what motivated the attention mechanism a year later.
The trick that made it work: reverse the source
One of the paper's most surprising findings: reversing the order of the source sentence (so "A B C" becomes "C B A" before encoding) improved the by nearly 5 points.
Why does this help? Consider translating "The cat sat" → "Le chat assis". Without reversal, the encoder's first word ("The") is steps away from the decoder's first word ("Le") in the computational graph — a long chain where gradients decay. With reversal, the encoder's last word processed ("The", now at the end) sits right next to the decoder's first word ("Le") — a short-term dependency that LSTMs handle well.
The reversal doesn't help every word pair — the end of the sentence gets pushed farther away — but it ensures the beginning of the translation has strong, undecayed signal. In practice, this gave the model a reliable "running start" and learning improved dramatically.
Architecture choices that matter
The authors made several deliberate design decisions:
Deep LSTMs (4 layers) — Shallow LSTMs (1 ) performed noticeably worse. Each additional layer adds representational capacity, letting the network build increasingly abstract features. Four layers gave the best balance of depth and trainability.
Separate encoder and decoder parameters — The two LSTMs do not share weights. The encoder learns to compress; the decoder learns to generate. Untying them gives each network the freedom to specialize.
1,000-dimensional hidden states — Each LSTM layer carries a 1,000-dimensional hidden state vector. With 4 layers, the model has 8,000 real-valued parameters carrying forward at each time step (1,000 × 4 layers × 2 for cell state and hidden state).
160k source vocabulary, 80k target vocabulary — Words not in these vocabularies are replaced with a special unknown token. The vocabulary size directly determines the softmax output layer's size, impacting both memory and computation.
Decoding: beam search finds better translations
At test time, the decoder must choose words one at a time. The simplest approach — — picks the highest-probability word at each step. But greedy decoding can't undo a bad early choice: if it picks a mediocre first word, the rest of the sentence is built on a shaky foundation.
keeps the top candidate sentences at each step (the "beam"). At each position, it expands every candidate by every possible next word, scores all expansions, and keeps only the best . It's like exploring multiple paths through a maze simultaneously instead of always committing to one turn.
The paper found that a beam size of just 2 gave substantial improvement over greedy search, while larger beams gave diminishing returns.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def lstm_step(x, h_prev, c_prev, W, U, b):
"""One step of LSTM: input x, previous hidden h, previous cell c."""
z = W @ x + U @ h_prev + b # combine input + hidden
i, f, o, g = np.split(z, 4) # input, forget, output gates + candidate
i, f, o = sigmoid(i), sigmoid(f), sigmoid(o)
g = np.tanh(g)
c = f * c_prev + i * g # cell update: forget old + write new
h = o * np.tanh(c) # hidden state: gated cell
return h, c
def encode(source_words, params):
"""Read the reversed source sentence, return final hidden state = context vector."""
h, c = np.zeros(D), np.zeros(D)
for word in reversed(source_words): # KEY: reverse the input
x = params['embed_src'][word]
h, c = lstm_step(x, h, c, *params['enc'])
return h # <-- this single vector must encode the ENTIRE input
def decode(context_vector, params, max_len=50):
"""Generate target words one at a time, starting from the context vector."""
h = context_vector # decoder starts where encoder ended
c = np.zeros(D)
word = '`<SOS>`' # start-of-sentence token
output = []
for _ in range(max_len):
x = params['embed_tgt'][word]
h, c = lstm_step(x, h, c, *params['dec'])
logits = params['W_out'] @ h + params['b_out'] # score every word
probs = softmax(logits) # probabilities over vocab
word = vocab[np.argmax(probs)] # greedy pick (or beam search)
if word == '`<EOS>`': break
output.append(word)
return output
# Training: maximize log p(target | reversed_source)
# At test time: beam search replaces greedy argmax for better resultsResults that changed the field
On the WMT'14 English-to-French translation benchmark:
-
A single Seq2Seq model scored 34.81 BLEU, surpassing the best phrase-based SMT baseline (33.30 BLEU) despite having no hand-engineered linguistic features.
-
An ensemble of 5 LSTMs with different random initializations and beam search reached 34.81 BLEU, showing that diversity helps even within the same architecture.
-
When the Seq2Seq model was used to re-rank the top 1,000 translations from the SMT system, the combined system achieved 36.5 BLEU — close to the best result at the time (37.0 BLEU).
-
The model handled long sentences remarkably well — contrary to expectations about the fixed-size bottleneck — though performance did degrade on sentences beyond ~35 words.
What the model learned
A remarkable finding emerged when the authors visualized the hidden representations. They projected the context vectors of many sentences into 2D using and discovered that sentences with similar meanings clustered together — even when their surface word order was different.
For instance, "John admires Mary" and "Mary is admired by John" mapped to nearly the same point, while "John admires Mary" and "John criticizes Mary" were far apart. The LSTM had learned to encode meaning, not just word sequence. This was early evidence that sequence models develop genuine semantic representations — not mere pattern memorization.
Training at scale (2014 edition)
Training the model was a feat of engineering for its time:
- 12 million sentence pairs from the WMT'14 English-French dataset - 8 GPUs running in parallel for roughly 10 days - 384 million parameters — enormous by 2014 standards - to prevent exploding gradients - Mini-batches of 128 sentences, grouped by length for efficiency - initialized at 0.7, halved every half-epoch after 5 epochs
The model processed about 6,300 words per second across the 8 GPUs. Sentences were bucketed by length so that mini-batches contained sentences of similar size, reducing wasted padding computation.
Limitations and what came next
The Seq2Seq paper was honest about its limitations:
Fixed-size bottleneck — compressing a 50-word sentence into one 1,000-dimensional vector necessarily loses information. The model's performance degraded on sentences longer than ~35 words, exactly where the bottleneck bites hardest.
No alignment mechanism — the decoder has no way to "look back" at specific source words. When translating word 15 of the output, it relies entirely on whatever the context vector preserved about source word 3. This motivated Bahdanau et al.'s attention mechanism (2015), which lets the decoder attend to different source words at each generation step.
Unknown words — words outside the fixed vocabulary become a generic unknown token, losing meaning. Subsequent work on subword tokenization () addressed this.
Sequential processing — LSTMs process tokens one at a time, limiting parallelism and training speed. The Transformer (2017) would solve this entirely.
The lineage: from Seq2Seq to modern NLP
2014
Seq2Seq (this paper)
Two LSTMs — one encodes, one decodes. Proved that end-to-end neural sequence learning can rival hand-engineered translation pipelines.
2014
RNN Encoder-Decoder (Cho et al.)
A concurrent, closely related architecture that also introduced the GRU cell. Together with Seq2Seq, established the encoder-decoder paradigm.
2015
Attention Mechanism (Bahdanau et al.)
Instead of one context vector, let the decoder attend to every encoder hidden state at each step. Solved the fixed-size bottleneck problem.
2015
Show and Tell (Vinyals et al.)
Applied Seq2Seq to image captioning — a CNN encoder replaces the LSTM encoder, and the decoder generates a natural language description of the image.
2016
Google Neural Machine Translation (GNMT)
Scaled Seq2Seq + attention to production quality. 8-layer encoder, 8-layer decoder, and replaced Google Translate's phrase-based system entirely.
2017
Transformer (Vaswani et al.)
Replaced recurrence with self-attention, solving the sequential processing bottleneck. The encoder-decoder structure is inherited directly from Seq2Seq.
2019
BART (Lewis et al.)
A denoising autoencoder using the Seq2Seq Transformer architecture. Pre-trained by corrupting text and learning to reconstruct it — a direct descendant of the encoder-decoder idea.
CitationSutskever, Vinyals, Le. Sequence to Sequence Learning with Neural Networks. NeurIPS, 2014.
Terms in this paper
- Sequence-to-Sequenceتسلسل إلى تسلسل
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Context Vectorمتجه السياق
- Beam Searchبحث الحزمة
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Hidden Stateالحالة المخفية
- end-to-endمن طرف إلى طرف
- BLEU Scoreمعيار بلو الإحصائي لتقييم الترجمة