RNNs & Sequence Models2014intermediate12 min read

Sequence to Sequence Learning with Neural Networks

التعلُّم من تسلسل إلى تسلسل باستخدام الشبكات العصبية

Sutskever, I. · Vinyals, O. · Le, Q.V. — NeurIPS

The problem

Deep Neural Networks (DNNs) excel at fixed-size inputs and outputs, but most real-world problems — translation, summarization, dialogue — involve variable-length sequences whose lengths are not known in advance. Traditional phrase-based statistical (SMT) systems were complex pipelines of hand-engineered components. There was no simple, general neural approach that maps one sequence directly to another.

The contribution

A general end-to-end sequence learning architecture using two multilayer LSTMs: an that reads the source sequence and compresses it into a single fixed-dimensional (the last ), and a that generates the target sequence one token at a time from that vector. A key trick: reversing the source sentence order, which places corresponding words closer together and dramatically improves learning. On English-to-French translation, the system achieved 34.81 BLEU — surpassing a phrase-based SMT baseline and approaching state-of-the-art when used to re-rank n-best lists.

The impact

The first convincing demonstration that a pure neural network — with no hand-engineered features — could rival statistical machine translation. It introduced the paradigm that became the backbone of all modern systems. Every subsequent milestone — mechanisms, the , GPT, BART — built directly on this architecture. It also inspired applications far beyond translation: image captioning (Show and Tell), speech synthesis (Tacotron), and conversational AI.

Imagine a simultaneous interpreter at a UN conference. She listens to the entire speech in French — absorbing meaning, intent, nuance — until the speaker finishes. Only then does she begin rendering it in English, one sentence at a time, from the compressed understanding she built in her mind.

She cannot go back and re-listen. Everything she needs must be in that mental summary. If the speech was short, her summary is vivid. If it was long, details start to blur — this is the of a fixed-size mental representation.

The Seq2Seq model works the same way: one listens to the full input and compresses it into a single vector, then a second LSTM speaks the output from that vector alone.

The problem: DNNs cannot handle variable-length sequences

By 2014, deep neural networks had conquered image classification, speech recognition, and many fixed-size tasks. But the standard DNN takes a fixed-size input and produces a fixed-size output. Real language doesn't work that way: a five-word English sentence might translate into eight French words.

The dominant solution — phrase-based statistical machine translation — was a pipeline of separately engineered pieces: an alignment model, a language model, a reordering model, phrase tables, and a decoder. Each piece was designed by hand. Each needed its own data and tuning. There was no simple way to train the whole thing end-to-end.

What if a single neural network could read an arbitrary-length input and write an arbitrary-length output?

The idea: two LSTMs — one listens, one speaks

Sutskever, Vinyals, and Le proposed a remarkably simple architecture. Split the problem into two halves:

The Encoder — a multilayer LSTM that reads the source sentence one word at a time. At each step it updates its hidden state, building an increasingly rich internal summary. After the last word, its final hidden state becomes a single fixed-dimensional vector — the context vector — that encodes the meaning of the entire input sentence.

The Decoder — a second multilayer LSTM, initialized with the context vector as its starting hidden state. It generates the target sentence one word at a time: at each step it receives the previously generated word and its own hidden state, predicts a probability distribution over the using , and emits the most likely next word. A special end-of-sentence token signals the decoder to stop.

The entire system is trained end-to-end to maximize the probability of the correct translation given the source sentence.

Open in Lab
Click any layer to explore the encoder-decoder pipeline step by step.
The demo wakes as you arrive…

The formal objective: maximize conditional probability

The goal is to estimate the conditional probability of a target sequence y1,…,yT′y_1, \dots, y_{T'} given a source sequence x1,…,xTx_1, \dots, x_T, where the two lengths TT and T′T' may differ. The encoder reads the source and produces a context vector vv, then the decoder models:

p(y1,…,yT′∣x1,…,xT)=∏t=1T′p(yt∣v,y1,…,yt−1)p(y_1, \dots, y_{T'} \mid x_1, \dots, x_T) = \prod_{t=1}^{T'} p(y_t \mid v, y_1, \dots, y_{t-1})
The chain-rule decomposition — the entire training objective — The model generates the output sequence one word at a time. At each step, it uses both the information extracted from the input sequence and all words generated so far to predict the next word. By repeatedly applying this process, the decoder constructs a complete output sentence while maintaining consistency with the input and previously generated content.

In practice, training maximizes the log-probability of the correct target sequence, which is equivalent to minimizing the cross-entropy loss. Each word prediction is a classification over the entire vocabulary (80,000 words for French in this paper).

The bottleneck: cramming meaning into a fixed-size vector

The context vector vv is the only link between encoder and decoder. Think of it as a narrow bridge between two islands: every bit of meaning from the source sentence must cross this bridge. The bridge has a fixed width (1,000 dimensions in this paper), whether the sentence is 5 words or 50.

For short sentences, the bridge is wide enough. For long sentences, the model must squeeze more meaning through the same opening — and inevitably loses detail. This bottleneck is the architecture's greatest limitation, and is precisely what motivated the attention mechanism a year later.

Open in Lab
Drag the sentence length slider to see how the fixed-size vector struggles with longer inputs.
The demo wakes as you arrive…

The trick that made it work: reverse the source

One of the paper's most surprising findings: reversing the order of the source sentence (so "A B C" becomes "C B A" before encoding) improved the by nearly 5 points.

Why does this help? Consider translating "The cat sat" → "Le chat assis". Without reversal, the encoder's first word ("The") is TT steps away from the decoder's first word ("Le") in the computational graph — a long chain where gradients decay. With reversal, the encoder's last word processed ("The", now at the end) sits right next to the decoder's first word ("Le") — a short-term dependency that LSTMs handle well.

The reversal doesn't help every word pair — the end of the sentence gets pushed farther away — but it ensures the beginning of the translation has strong, undecayed signal. In practice, this gave the model a reliable "running start" and learning improved dramatically.

Open in Lab
Toggle reversal on/off to see how gradient path lengths change between aligned word pairs.
The demo wakes as you arrive…

Architecture choices that matter

The authors made several deliberate design decisions:

Deep LSTMs (4 layers) — Shallow LSTMs (1 ) performed noticeably worse. Each additional layer adds representational capacity, letting the network build increasingly abstract features. Four layers gave the best balance of depth and trainability.

Separate encoder and decoder parameters — The two LSTMs do not share weights. The encoder learns to compress; the decoder learns to generate. Untying them gives each network the freedom to specialize.

1,000-dimensional hidden states — Each LSTM layer carries a 1,000-dimensional hidden state vector. With 4 layers, the model has 8,000 real-valued parameters carrying forward at each time step (1,000 × 4 layers × 2 for cell state and hidden state).

160k source vocabulary, 80k target vocabulary — Words not in these vocabularies are replaced with a special unknown token. The vocabulary size directly determines the softmax output layer's size, impacting both memory and computation.

Decoding: beam search finds better translations

At test time, the decoder must choose words one at a time. The simplest approach — — picks the highest-probability word at each step. But greedy decoding can't undo a bad early choice: if it picks a mediocre first word, the rest of the sentence is built on a shaky foundation.

keeps the top BB candidate sentences at each step (the "beam"). At each position, it expands every candidate by every possible next word, scores all expansions, and keeps only the best BB. It's like exploring multiple paths through a maze simultaneously instead of always committing to one turn.

The paper found that a beam size of just 2 gave substantial improvement over greedy search, while larger beams gave diminishing returns.

Open in Lab
Watch beam search (B=3) explore multiple candidate translations simultaneously. Greedy search commits to one path; beam search hedges its bets.
The demo wakes as you arrive…

The same idea in code

A minimal seq2seq encoder-decoder with LSTMpython

Simplified to show the idea — not the real implementation.

import numpy as np

def lstm_step(x, h_prev, c_prev, W, U, b):
    """One step of LSTM: input x, previous hidden h, previous cell c."""
    z = W @ x + U @ h_prev + b      # combine input + hidden
    i, f, o, g = np.split(z, 4)      # input, forget, output gates + candidate
    i, f, o = sigmoid(i), sigmoid(f), sigmoid(o)
    g = np.tanh(g)
    c = f * c_prev + i * g           # cell update: forget old + write new
    h = o * np.tanh(c)               # hidden state: gated cell
    return h, c

def encode(source_words, params):
    """Read the reversed source sentence, return final hidden state = context vector."""
    h, c = np.zeros(D), np.zeros(D)
    for word in reversed(source_words):          # KEY: reverse the input
        x = params['embed_src'][word]
        h, c = lstm_step(x, h, c, *params['enc'])
    return h  # <-- this single vector must encode the ENTIRE input

def decode(context_vector, params, max_len=50):
    """Generate target words one at a time, starting from the context vector."""
    h = context_vector       # decoder starts where encoder ended
    c = np.zeros(D)
    word = '`<SOS>`'           # start-of-sentence token
    output = []
    for _ in range(max_len):
        x = params['embed_tgt'][word]
        h, c = lstm_step(x, h, c, *params['dec'])
        logits = params['W_out'] @ h + params['b_out']   # score every word
        probs = softmax(logits)                           # probabilities over vocab
        word = vocab[np.argmax(probs)]                    # greedy pick (or beam search)
        if word == '`<EOS>`': break
        output.append(word)
    return output

# Training: maximize log p(target | reversed_source)
# At test time: beam search replaces greedy argmax for better results

Results that changed the field

On the WMT'14 English-to-French translation benchmark:

  • A single Seq2Seq model scored 34.81 BLEU, surpassing the best phrase-based SMT baseline (33.30 BLEU) despite having no hand-engineered linguistic features.

  • An ensemble of 5 LSTMs with different random initializations and beam search reached 34.81 BLEU, showing that diversity helps even within the same architecture.

  • When the Seq2Seq model was used to re-rank the top 1,000 translations from the SMT system, the combined system achieved 36.5 BLEU — close to the best result at the time (37.0 BLEU).

  • The model handled long sentences remarkably well — contrary to expectations about the fixed-size bottleneck — though performance did degrade on sentences beyond ~35 words.

Open in Lab
Compare BLEU scores across approaches. Notice how the neural system matches hand-engineered SMT.
The demo wakes as you arrive…

What the model learned

A remarkable finding emerged when the authors visualized the hidden representations. They projected the context vectors of many sentences into 2D using and discovered that sentences with similar meanings clustered together — even when their surface word order was different.

For instance, "John admires Mary" and "Mary is admired by John" mapped to nearly the same point, while "John admires Mary" and "John criticizes Mary" were far apart. The LSTM had learned to encode meaning, not just word sequence. This was early evidence that sequence models develop genuine semantic representations — not mere pattern memorization.

Training at scale (2014 edition)

Training the model was a feat of engineering for its time:

  • 12 million sentence pairs from the WMT'14 English-French dataset - 8 GPUs running in parallel for roughly 10 days - 384 million parameters — enormous by 2014 standards - to prevent exploding gradients - Mini-batches of 128 sentences, grouped by length for efficiency - initialized at 0.7, halved every half-epoch after 5 epochs

The model processed about 6,300 words per second across the 8 GPUs. Sentences were bucketed by length so that mini-batches contained sentences of similar size, reducing wasted padding computation.

Limitations and what came next

The Seq2Seq paper was honest about its limitations:

Fixed-size bottleneck — compressing a 50-word sentence into one 1,000-dimensional vector necessarily loses information. The model's performance degraded on sentences longer than ~35 words, exactly where the bottleneck bites hardest.

No alignment mechanism — the decoder has no way to "look back" at specific source words. When translating word 15 of the output, it relies entirely on whatever the context vector preserved about source word 3. This motivated Bahdanau et al.'s attention mechanism (2015), which lets the decoder attend to different source words at each generation step.

Unknown words — words outside the fixed vocabulary become a generic unknown token, losing meaning. Subsequent work on subword tokenization () addressed this.

Sequential processing — LSTMs process tokens one at a time, limiting parallelism and training speed. The Transformer (2017) would solve this entirely.

The lineage: from Seq2Seq to modern NLP

  1. 2014

    Seq2Seq (this paper)

    Two LSTMs — one encodes, one decodes. Proved that end-to-end neural sequence learning can rival hand-engineered translation pipelines.

  2. 2014

    RNN Encoder-Decoder (Cho et al.)

    A concurrent, closely related architecture that also introduced the GRU cell. Together with Seq2Seq, established the encoder-decoder paradigm.

  3. 2015

    Attention Mechanism (Bahdanau et al.)

    Instead of one context vector, let the decoder attend to every encoder hidden state at each step. Solved the fixed-size bottleneck problem.

  4. 2015

    Show and Tell (Vinyals et al.)

    Applied Seq2Seq to image captioning — a CNN encoder replaces the LSTM encoder, and the decoder generates a natural language description of the image.

  5. 2016

    Google Neural Machine Translation (GNMT)

    Scaled Seq2Seq + attention to production quality. 8-layer encoder, 8-layer decoder, and replaced Google Translate's phrase-based system entirely.

  6. 2017

    Transformer (Vaswani et al.)

    Replaced recurrence with self-attention, solving the sequential processing bottleneck. The encoder-decoder structure is inherited directly from Seq2Seq.

  7. 2019

    BART (Lewis et al.)

    A denoising autoencoder using the Seq2Seq Transformer architecture. Pre-trained by corrupting text and learning to reconstruct it — a direct descendant of the encoder-decoder idea.

CitationSutskever, Vinyals, Le. Sequence to Sequence Learning with Neural Networks. NeurIPS, 2014.

Terms in this paper