RNNs & Sequence Models1997advanced4 min read

Long Short-Term Memory

الذاكرة الطويلة قصيرة المدى

Hochreiter, S. · Schmidhuber, J. — Neural Computation

The problem

Standard RNNs suffer from vanishing and exploding gradients when backpropagating through time. As error signals are multiplied repeatedly across steps, they either shrink to zero or blow up to infinity, making it impossible to learn long-range dependencies.

The contribution

The (CEC): a memory cell whose internal update is purely additive, keeping the 's derivative at exactly 1.0 as it flows backward. Input and output gates learn via gradient descent when to write new information and when to expose it. The 2000 addition of the completes the modern architecture.

The impact

The dominant sequence model for two decades. LSTMs powered commercial speech recognition, machine translation, and text-to-speech (early Siri, Google Translate, Alexa) before the era. Its bottleneck — — is precisely what motivated the attention mechanism and the Transformer.

A standard trying to retain context across a long text is like carrying a handful of dry sand through a desert wind. Every step you take, a fraction slips through your fingers, and after 50 steps your hand is completely empty. An carries that context inside a secure, locking briefcase. The briefcase stays sealed tightly. If step 5 contains a vital clue, the Input Gate unlocks, slips it inside, and locks back up. It travels safely through 100 steps of irrelevant noise until the Output Gate decides it is time to open and present the contents.

The problem: gradients that vanish

In a standard RNN the at step t is:

ht=tanh⁡(Wht−1+xt)h_t = \tanh(W h_{t-1} + x_t)
Standard RNN step

When we backpropagate from step T to an early step t, the forces a product of repeated Jacobians:

∂ET∂ht  ∝  ∏k=t+1TW⊤⋅diag⁡(1−tanh⁡2(⋅))\frac{\partial E_T}{\partial h_t} \;\propto\; \prod_{k=t+1}^{T} W^\top \cdot \operatorname{diag}(1 - \tanh^2(\cdot))
Why gradients vanish in standard RNNs

If the eigenvalues of W are slightly less than 1, this product vanishes exponentially — reaching near zero within ~10–20 steps. If they are greater than 1 it explodes toward infinity. Either way, the network becomes historically blind: it cannot learn what happened far in the past.

Open in Lab
Drag the weight multiplier — watch RNN gradients vanish or explode while the LSTM's CEC channel stays perfectly flat at 1.
The demo wakes as you arrive…

Hochreiter & Schmidhuber's key insight: design a memory location whose internal derivative is exactly 1.0. Instead of multiplying the past state by a weight matrix, the internal Cₜ updates additively:

Ct=Ct−1+C~tC_t = C_{t-1} + \tilde{C}_t
Core additive update (original 1997 form) — The additive connection means the gradient flowing backward encounters a derivative of exactly 1 — it can travel across 1,000 steps without losing strength.

Adding the Forget Gate (Gers et al., 2000)

The original 1997 LSTM lacked one crucial mechanism: it could never clear its own memory. Processing a continuous stream of text, the cell state would grow without bound, eventually saturating and corrupting new inputs. Felix Gers, Schmidhuber, and Cummins solved this in 2000 by introducing the Forget Gate (fₜ) — a multiplicative scalar that scales the previous cell state before adding new content:

Ct=ft⊙Ct−1+it⊙C~tC_t = f_t \odot C_{t-1} + i_t \odot \tilde{C}_t
Modern cell state update (with forget gate) — When fₜ → 0 the old memory is wiped; when fₜ → 1 it is fully preserved. This lets the network reset at sentence or topic boundaries.

The complete modern LSTM cell

ft=σ(Wf[ht−1,xt]+bf)it=σ(Wi[ht−1,xt]+bi)C~t=tanh⁡(Wc[ht−1,xt]+bc)Ct=ft⊙Ct−1+it⊙C~tot=σ(Wo[ht−1,xt]+bo)ht=ot⊙tanh⁡(Ct)\begin{aligned} f_t &= \sigma(W_f [h_{t-1}, x_t] + b_f) \\ i_t &= \sigma(W_i [h_{t-1}, x_t] + b_i) \\ \tilde{C}_t &= \tanh(W_c [h_{t-1}, x_t] + b_c) \\ C_t &= f_t \odot C_{t-1} + i_t \odot \tilde{C}_t \\ o_t &= \sigma(W_o [h_{t-1}, x_t] + b_o) \\ h_t &= o_t \odot \tanh(C_t) \end{aligned}
Standard LSTM equations (post-2000) — Four parallel linear transformations, all trained by backpropagation through time (BPTT).
Open in Lab
Click any component to understand its role, its equation, and its relationship to the others.
The demo wakes as you arrive…

Watch the gates in action

Open in Lab
Step through one forward pass — watch each gate value determine what gets written, kept, and exposed at each stage.
The demo wakes as you arrive…

The same idea in code

One LSTM step with forget gatepython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(x):
    return 1 / (1 + np.exp(-x))

def lstm_step(x_t, h_prev, C_prev, params):
    """
    x_t:    input vector at current step
    h_prev: previous hidden state
    C_prev: previous cell state (the protected memory)
    """
    combined = np.concatenate([h_prev, x_t])

    # 1. Forget Gate — how much of the old memory to keep?
    f_t = sigmoid(params['W_f'] @ combined + params['b_f'])

    # 2. Input Gate — how much of the new candidate to write?
    i_t = sigmoid(params['W_i'] @ combined + params['b_i'])

    # 3. Candidate — what would we write?
    C_tilde = np.tanh(params['W_c'] @ combined + params['b_c'])

    # 4. Cell update — the whole paper in one line
    C_t = f_t * C_prev + i_t * C_tilde     # additive → gradient stays at 1

    # 5. Output Gate — how much to reveal right now?
    o_t = sigmoid(params['W_o'] @ combined + params['b_o'])

    # 6. Hidden state
    h_t = o_t * np.tanh(C_t)

    return h_t, C_t

Historical evolution

  1. 1997

    Birth of LSTM

    Hochreiter and Schmidhuber introduce the CEC and input/output gating. Vanishing gradients are solved for the first time in a practical recurrent architecture.

  2. 2000

    The Forget Gate

    Felix Gers, Schmidhuber, and Cummins add the forget gate, giving LSTM the ability to reset its memory at sequence boundaries. This completes the standard modern cell.

  3. 2006

    LSTMs + CTC Loss

    Alex Graves combines LSTMs with Connectionist Temporal Classification (CTC), enabling unsegmented speech recognition — the foundation of modern voice assistants.

  4. 2014

    Seq2Seq and machine translation

    Sutskever et al. use stacked LSTMs in an encoder–decoder configuration, revolutionizing machine translation at Google. LSTM becomes the backbone of commercial NLP.

  5. 2015

    Attention over LSTM

    Bahdanau et al. add dynamic alignment over LSTM hidden states — the crucial intermediate step that directly inspired the self-attention mechanism in Transformers.

  6. 2017

    Transformers replace recurrence

    "Attention Is All You Need" shows self-attention alone — no recurrence — matches and then exceeds LSTM on translation, while being fully parallelizable. The LSTM era ends.

CitationHochreiter, S., Schmidhuber, J.. Long Short-Term Memory. Neural Computation 9(8), 1997.

Terms in this paper