RNNs & Sequence Models2014intermediate11 min read
Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
تقييم تجريبي للشبكات العصبية التكرارية ذات البوابات في نمذجة التتابعات
Chung, J. · Gulcehre, C. · Cho, K. · Bengio, Y. — NIPS 2014 Workshop
The problem
By 2014, the had become the go-to solution for tasks — music, speech, translation — because vanilla RNNs suffered from vanishing gradients on long sequences. But the LSTM's four-gate architecture was complex, slow to train, and -heavy. Cho et al. had just proposed a simpler gated unit () for , but no one had systematically compared it against LSTM and vanilla tanh RNNs across multiple domains. The field needed an empirical answer: does the simpler design sacrifice quality?
The contribution
A controlled head-to-head comparison of three recurrent units — vanilla tanh, LSTM, and GRU — on polyphonic music modeling and raw speech signal modeling. All models were given the same number of parameters so the comparison is fair. Key finding: gated units (LSTM and GRU) clearly outperform tanh units. GRU is comparable to LSTM, often converging faster. The simpler two-gate design does not sacrifice quality.
The impact
This paper gave the community confidence that GRU is a legitimate alternative to LSTM. It accelerated adoption of GRU in production systems where faster and fewer parameters matter — edge devices, real-time applications, and rapid prototyping. The paper's central insight — that simpler gating can match complex gating — also influenced later work on minimal recurrent architectures and helped shape the design philosophy that led to streamlined sequence models.
Think of a as writing on a whiteboard and erasing everything before writing the next line — by page 100 you've lost page 1 entirely.
An LSTM fixes this by adding a locked filing cabinet (memory cell), a shredder (forget gate), a mail slot (input gate), and a tinted window (output gate). Nothing is lost accidentally, but the machinery is heavy.
A GRU asks: what if we skip the filing cabinet and the window, and instead give the writer just two dials — one that controls how much of the old text to keep () and one that controls how much old context to consider before drafting new text ()? The result: nearly the same recall, half the moving parts.
The problem: vanilla RNNs forget too fast
A vanilla updates its by completely overwriting it at each step:
This means the network replaces its entire memory every time step. Information from early in the sequence must survive through a chain of multiplications and squashing functions. After enough steps, gradients either vanish (the signal decays to nothing) or explode (it grows uncontrollably). In practice, vanilla RNNs struggle to connect events more than 10–20 steps apart.
By 2014, two gated alternatives had emerged: the well-established LSTM (1997) and the brand-new GRU (2014). Both add gates — learned switches that control information flow — but they differ in how many gates they use and what each gate controls.
LSTM: the four-gate heavyweight
The LSTM unit maintains a separate memory cell alongside the hidden state . Four components control it:
- Forget gate : decides how much of the old memory to keep. Think of a shredder — turn it up and old information is destroyed; turn it down and everything is preserved.
- Input gate : decides how much of the new candidate content to write into the cell. Like a mail slot — it controls what gets in.
- Cell update: . Old memory (partially forgotten) plus new content (partially admitted).
- Output gate : decides how much of the cell's content to reveal as the hidden state. A tinted window — the filing cabinet is there, but you only see what the gate lets through: .
This design is powerful: the cell can carry information across hundreds of steps unchanged. But it requires three gating vectors (, , ) plus a candidate (), meaning four matrix multiplications per step — a heavy computational footprint.
GRU: same idea, fewer parts
The GRU asks: can we get the benefits of gating without a separate memory cell and without an output gate? Its answer is yes, using just two gates:
Update gate — controls the blend between old and new. When , the unit adopts the new candidate entirely; when , it copies the old state unchanged. This single gate plays both the roles of the LSTM's forget and input gates — they are tied together: whatever fraction you forget is exactly the fraction you fill with new content.
Reset gate — controls how much of the previous state is visible when computing the candidate. When , the unit behaves as if it is reading the very first element of the sequence — it can make a fresh start. When , all previous context flows in.
Think of a highway with a single adjustable lane divider (update gate): one side carries old information forward, the other carries new information in. The reset gate is a filter on the rearview mirror — it controls how much of the road behind you influences your next turn.
LSTM vs GRU: a side-by-side view
The two units share a fundamental idea: additive updates create shortcut paths that let gradients flow across many time steps without vanishing. In both units, the new state is not computed by overwriting; it is blended with the old state via learned gates.
But they differ in architecture:
- LSTM has a separate memory cell that is never directly exposed; the output gate controls what the rest of the network sees. GRU has no separate cell — the hidden state is the memory.
- LSTM's forget and input gates are independent: you can forget 80% and write 50%. In GRU they are tied: forget 80% means write exactly 80%.
- LSTM uses 4 weight matrices per gate/candidate (≈ 4 matrix multiplies per step). GRU uses 3 (update gate, reset gate, candidate). Fewer parameters for the same hidden size.
The shared secret: additive shortcuts for gradients
The paper highlights that the most important feature shared by LSTM and GRU — and absent from vanilla RNNs — is their additive update path.
In a vanilla RNN, is computed by a nonlinear function of , creating a chain of bounded nonlinearities. Back-propagating through this chain multiplies many Jacobians together, and if each has spectral radius < 1, the gradient shrinks exponentially — the problem.
Both LSTM and GRU replace this chain with an additive shortcut: part of the old state is added directly to the new state. When the gate is saturated near 1, the flows through unchanged — like a highway that bypasses city traffic. This is the same principle behind residual connections in deep feedforward networks, applied here to the time axis.
The experiment: same parameters, different units
To make the comparison fair, the authors gave all three models approximately the same total number of parameters — about 20K for music tasks and 169K for speech tasks. Because GRU has fewer matrices per unit than LSTM, it can afford more hidden units at the same parameter count: 46 GRU units vs 36 LSTM units for music, 227 vs 195 for speech.
They tested on two domains:
- Polyphonic music modeling (4 datasets: Nottingham, JSB Chorales, MuseData, Piano-midi): each time step is a multi-hot binary representing which notes are being played simultaneously.
- Speech signal modeling (2 Ubisoft datasets): raw one-dimensional audio. The reads 20 consecutive samples and predicts the next 10. One dataset has sequences of length 500 (short), the other 8,000 (long) — a direct test of long-range dependency capture.
All models were trained with , weight noise (σ = 0.075), and (norm ≤ 1). Learning rates were tuned by random search.
Results: gated units win, GRU matches LSTM
The results tell a clear two-level story:
Level 1 — Gates matter. On the challenging speech datasets, both LSTM and GRU dramatically outperformed the vanilla tanh unit. On Ubisoft B (long sequences), tanh achieved a test negative log-likelihood of 7.62 while GRU reached 0.88 and LSTM 1.26 — an order of magnitude better. The gating mechanism is not optional; it is essential for non-trivial .
Level 2 — GRU ≈ LSTM. Across all six datasets, neither gated unit consistently dominated the other. GRU won on most music datasets and on the long speech dataset; LSTM won on the short speech dataset. On music, the differences were small. The authors' conclusion: the choice between them depends on the task and dataset, but GRU's simpler design carries no systematic penalty.
The learning curves revealed an additional advantage: GRU often converged faster in both wall-clock time and number of updates, likely because fewer parameters per unit means each gradient step is cheaper and the optimization landscape is smoother.
The GRU in code
Simplified to show the idea — not the real implementation.
import numpy as np
def sigmoid(x):
return 1 / (1 + np.exp(-x))
def gru_step(x_t, h_prev, W_z, U_z, W_r, U_r, W_h, U_h):
"""One GRU time step. x_t: input, h_prev: previous hidden state."""
# Update gate: how much of old state to keep
z_t = sigmoid(W_z @ x_t + U_z @ h_prev)
# Reset gate: how much past context to use for candidate
r_t = sigmoid(W_r @ x_t + U_r @ h_prev)
# Candidate activation: new content proposal
h_tilde = np.tanh(W_h @ x_t + U_h @ (r_t * h_prev))
# Final state: blend old and new
h_t = (1 - z_t) * h_prev + z_t * h_tilde
return h_t
# Compare with LSTM: no separate cell state, no output gate.
# Just two gates and one candidate — that's the whole unit.Why this paper mattered
Before this paper, GRU had only been tested in one context (machine translation by Cho et al., 2014). This systematic evaluation across music and speech gave the community the empirical confidence to adopt GRU broadly. The practical implications were significant:
- Faster training: fewer parameters per unit means faster gradient computation and often faster .
- Smaller models: for the same hidden dimension, GRU uses ~25% fewer parameters than LSTM — crucial for edge deployment.
- Design philosophy: the paper showed that architectural complexity does not automatically buy better performance. Sometimes simpler is just as good — a lesson that echoes through modern architecture design.
CitationChung, Gulcehre, Cho, Bengio. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. NIPS 2014 Workshop, 2014.
Terms in this paper
- GRUالوحدة العودية البوابية
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Update Gateبوابة التحديث
- Reset Gateبوابة إعادة التعيين
- Hidden Stateالحالة المخفية
- Vanishing Gradientاضمحلال متجهات الميل
- Sequence Modelingنمذجة التتابعات
- Recurrent Neural Network (RNN)الشبكة العصبية التكرارية