RNNs & Sequence Models1997advanced4 min read
Long Short-Term Memory
الذاكرة الطويلة قصيرة المدى
Hochreiter, S. · Schmidhuber, J. — Neural Computation
The problem
Standard RNNs suffer from vanishing and exploding gradients when backpropagating through time. As error signals are multiplied repeatedly across steps, they either shrink to zero or blow up to infinity, making it impossible to learn long-range dependencies.
The contribution
The (CEC): a memory cell whose internal update is purely additive, keeping the 's derivative at exactly 1.0 as it flows backward. Input and output gates learn via gradient descent when to write new information and when to expose it. The 2000 addition of the completes the modern architecture.
The impact
The dominant sequence model for two decades. LSTMs powered commercial speech recognition, machine translation, and text-to-speech (early Siri, Google Translate, Alexa) before the era. Its bottleneck — — is precisely what motivated the attention mechanism and the Transformer.
A standard trying to retain context across a long text is like carrying a handful of dry sand through a desert wind. Every step you take, a fraction slips through your fingers, and after 50 steps your hand is completely empty. An carries that context inside a secure, locking briefcase. The briefcase stays sealed tightly. If step 5 contains a vital clue, the Input Gate unlocks, slips it inside, and locks back up. It travels safely through 100 steps of irrelevant noise until the Output Gate decides it is time to open and present the contents.
The problem: gradients that vanish
In a standard RNN the at step t is:
When we backpropagate from step T to an early step t, the forces a product of repeated Jacobians:
If the eigenvalues of W are slightly less than 1, this product vanishes exponentially — reaching near zero within ~10–20 steps. If they are greater than 1 it explodes toward infinity. Either way, the network becomes historically blind: it cannot learn what happened far in the past.
The solution: Constant Error Carousel
Hochreiter & Schmidhuber's key insight: design a memory location whose internal derivative is exactly 1.0. Instead of multiplying the past state by a weight matrix, the internal Cₜ updates additively:
Adding the Forget Gate (Gers et al., 2000)
The original 1997 LSTM lacked one crucial mechanism: it could never clear its own memory. Processing a continuous stream of text, the cell state would grow without bound, eventually saturating and corrupting new inputs. Felix Gers, Schmidhuber, and Cummins solved this in 2000 by introducing the Forget Gate (fₜ) — a multiplicative scalar that scales the previous cell state before adding new content:
The complete modern LSTM cell
Watch the gates in action
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def sigmoid(x):
return 1 / (1 + np.exp(-x))
def lstm_step(x_t, h_prev, C_prev, params):
"""
x_t: input vector at current step
h_prev: previous hidden state
C_prev: previous cell state (the protected memory)
"""
combined = np.concatenate([h_prev, x_t])
# 1. Forget Gate — how much of the old memory to keep?
f_t = sigmoid(params['W_f'] @ combined + params['b_f'])
# 2. Input Gate — how much of the new candidate to write?
i_t = sigmoid(params['W_i'] @ combined + params['b_i'])
# 3. Candidate — what would we write?
C_tilde = np.tanh(params['W_c'] @ combined + params['b_c'])
# 4. Cell update — the whole paper in one line
C_t = f_t * C_prev + i_t * C_tilde # additive → gradient stays at 1
# 5. Output Gate — how much to reveal right now?
o_t = sigmoid(params['W_o'] @ combined + params['b_o'])
# 6. Hidden state
h_t = o_t * np.tanh(C_t)
return h_t, C_tHistorical evolution
1997
Birth of LSTM
Hochreiter and Schmidhuber introduce the CEC and input/output gating. Vanishing gradients are solved for the first time in a practical recurrent architecture.
2000
The Forget Gate
Felix Gers, Schmidhuber, and Cummins add the forget gate, giving LSTM the ability to reset its memory at sequence boundaries. This completes the standard modern cell.
2006
LSTMs + CTC Loss
Alex Graves combines LSTMs with Connectionist Temporal Classification (CTC), enabling unsegmented speech recognition — the foundation of modern voice assistants.
2014
Seq2Seq and machine translation
Sutskever et al. use stacked LSTMs in an encoder–decoder configuration, revolutionizing machine translation at Google. LSTM becomes the backbone of commercial NLP.
2015
Attention over LSTM
Bahdanau et al. add dynamic alignment over LSTM hidden states — the crucial intermediate step that directly inspired the self-attention mechanism in Transformers.
2017
Transformers replace recurrence
"Attention Is All You Need" shows self-attention alone — no recurrence — matches and then exceeds LSTM on translation, while being fully parallelizable. The LSTM era ends.
CitationHochreiter, S., Schmidhuber, J.. Long Short-Term Memory. Neural Computation 9(8), 1997.
Terms in this paper
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Cell Stateحالة الخلية الحسابية
- Forget Gateبوابة النسيان
- Gating Mechanismآلية البوابات
- Vanishing Gradientاضمحلال متجهات الميل
- Constant Error Carouselدولاب الخطأ الثابت
- Long-Term Dependencyالاعتمادية البعيدة المدى