RNNs & Sequence Models2014intermediate11 min read

Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

تقييم تجريبي للشبكات العصبية التكرارية ذات البوابات في نمذجة التتابعات

Chung, J. · Gulcehre, C. · Cho, K. · Bengio, Y. — NIPS 2014 Workshop

The problem

By 2014, the had become the go-to solution for tasks — music, speech, translation — because vanilla RNNs suffered from vanishing gradients on long sequences. But the LSTM's four-gate architecture was complex, slow to train, and -heavy. Cho et al. had just proposed a simpler gated unit () for , but no one had systematically compared it against LSTM and vanilla tanh RNNs across multiple domains. The field needed an empirical answer: does the simpler design sacrifice quality?

The contribution

A controlled head-to-head comparison of three recurrent units — vanilla tanh, LSTM, and GRU — on polyphonic music modeling and raw speech signal modeling. All models were given the same number of parameters so the comparison is fair. Key finding: gated units (LSTM and GRU) clearly outperform tanh units. GRU is comparable to LSTM, often converging faster. The simpler two-gate design does not sacrifice quality.

The impact

This paper gave the community confidence that GRU is a legitimate alternative to LSTM. It accelerated adoption of GRU in production systems where faster and fewer parameters matter — edge devices, real-time applications, and rapid prototyping. The paper's central insight — that simpler gating can match complex gating — also influenced later work on minimal recurrent architectures and helped shape the design philosophy that led to streamlined sequence models.

Think of a as writing on a whiteboard and erasing everything before writing the next line — by page 100 you've lost page 1 entirely.

An LSTM fixes this by adding a locked filing cabinet (memory cell), a shredder (forget gate), a mail slot (input gate), and a tinted window (output gate). Nothing is lost accidentally, but the machinery is heavy.

A GRU asks: what if we skip the filing cabinet and the window, and instead give the writer just two dials — one that controls how much of the old text to keep () and one that controls how much old context to consider before drafting new text ()? The result: nearly the same recall, half the moving parts.

The problem: vanilla RNNs forget too fast

A vanilla updates its by completely overwriting it at each step:

ht=tanh⁡(Wxt+Uht−1)h_t = \tanh(W x_t + U h_{t-1})

This means the network replaces its entire memory every time step. Information from early in the sequence must survive through a chain of multiplications and squashing functions. After enough steps, gradients either vanish (the signal decays to nothing) or explode (it grows uncontrollably). In practice, vanilla RNNs struggle to connect events more than 10–20 steps apart.

By 2014, two gated alternatives had emerged: the well-established LSTM (1997) and the brand-new GRU (2014). Both add gates — learned switches that control information flow — but they differ in how many gates they use and what each gate controls.

Open in Lab
Watch how a vanilla RNN's memory of an early event fades compared to gated units.
The demo wakes as you arrive…

LSTM: the four-gate heavyweight

The LSTM unit maintains a separate memory cell ctc_t alongside the hidden state hth_t. Four components control it:

  • Forget gate ftf_t: decides how much of the old memory to keep. Think of a shredder — turn it up and old information is destroyed; turn it down and everything is preserved.
  • Input gate iti_t: decides how much of the new candidate content c~t\tilde{c}_t to write into the cell. Like a mail slot — it controls what gets in.
  • Cell update: ct=ft⊙ct−1+it⊙c~tc_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t. Old memory (partially forgotten) plus new content (partially admitted).
  • Output gate oto_t: decides how much of the cell's content to reveal as the hidden state. A tinted window — the filing cabinet is there, but you only see what the gate lets through: ht=ot⊙tanh⁡(ct)h_t = o_t \odot \tanh(c_t).

This design is powerful: the cell can carry information across hundreds of steps unchanged. But it requires three gating vectors (ftf_t, iti_t, oto_t) plus a candidate (c~t\tilde{c}_t), meaning four matrix multiplications per step — a heavy computational footprint.

ct=ft⊙ct−1+it⊙c~t,ht=ot⊙tanh⁡(ct)c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t, \qquad h_t = o_t \odot \tanh(c_t)
LSTM cell update and output — f = forget gate (what to erase) · i = input gate (what to write) · o = output gate (what to reveal) · c̃ = candidate content · ⊙ = element-wise multiplication

GRU: same idea, fewer parts

The GRU asks: can we get the benefits of gating without a separate memory cell and without an output gate? Its answer is yes, using just two gates:

Update gate ztz_t — controls the blend between old and new. When zt≈1z_t \approx 1, the unit adopts the new candidate entirely; when zt≈0z_t \approx 0, it copies the old state unchanged. This single gate plays both the roles of the LSTM's forget and input gates — they are tied together: whatever fraction you forget is exactly the fraction you fill with new content.

Reset gate rtr_t — controls how much of the previous state is visible when computing the candidate. When rt≈0r_t \approx 0, the unit behaves as if it is reading the very first element of the sequence — it can make a fresh start. When rt≈1r_t \approx 1, all previous context flows in.

Think of a highway with a single adjustable lane divider (update gate): one side carries old information forward, the other carries new information in. The reset gate is a filter on the rearview mirror — it controls how much of the road behind you influences your next turn.

zt=σ(Wzxt+Uzht−1)z_t = \sigma(W_z x_t + U_z h_{t-1})
Update gate — Decides the blend ratio: z ≈ 1 → adopt new content; z ≈ 0 → keep old state unchanged
rt=σ(Wrxt+Urht−1)r_t = \sigma(W_r x_t + U_r h_{t-1})
Reset gate — Controls how much past state the candidate sees: r ≈ 0 → ignore history (fresh start); r ≈ 1 → full context
h~t=tanh⁡(Wxt+U(rt⊙ht−1))\tilde{h}_t = \tanh(W x_t + U(r_t \odot h_{t-1}))
Candidate activation — New content proposal, computed from current input and (optionally reset) past state
ht=(1−zt)⊙ht−1+zt⊙h~th_t = (1 - z_t) \odot h_{t-1} + z_t \odot \tilde{h}_t
GRU hidden state update — the core equation — A linear interpolation: (1−z) × old state + z × new candidate. If z=0, the old state passes through untouched — a shortcut path for the gradient.
Open in Lab
Drag the sliders to see how the update and reset gates control information flow step by step.
The demo wakes as you arrive…

LSTM vs GRU: a side-by-side view

The two units share a fundamental idea: additive updates create shortcut paths that let gradients flow across many time steps without vanishing. In both units, the new state is not computed by overwriting; it is blended with the old state via learned gates.

But they differ in architecture:

  • LSTM has a separate memory cell ctc_t that is never directly exposed; the output gate controls what the rest of the network sees. GRU has no separate cell — the hidden state is the memory.
  • LSTM's forget and input gates are independent: you can forget 80% and write 50%. In GRU they are tied: forget 80% means write exactly 80%.
  • LSTM uses 4 weight matrices per gate/candidate (≈ 4 matrix multiplies per step). GRU uses 3 (update gate, reset gate, candidate). Fewer parameters for the same hidden size.
Open in Lab
Toggle between the two architectures. Notice how GRU merges the memory cell into the hidden state.
The demo wakes as you arrive…

The shared secret: additive shortcuts for gradients

The paper highlights that the most important feature shared by LSTM and GRU — and absent from vanilla RNNs — is their additive update path.

In a vanilla RNN, hth_t is computed by a nonlinear function of ht−1h_{t-1}, creating a chain of bounded nonlinearities. Back-propagating through this chain multiplies many Jacobians together, and if each has spectral radius < 1, the gradient shrinks exponentially — the problem.

Both LSTM and GRU replace this chain with an additive shortcut: part of the old state is added directly to the new state. When the gate is saturated near 1, the flows through unchanged — like a highway that bypasses city traffic. This is the same principle behind residual connections in deep feedforward networks, applied here to the time axis.

Open in Lab
Compare gradient magnitude across time steps for vanilla, LSTM, and GRU units.
The demo wakes as you arrive…

The experiment: same parameters, different units

To make the comparison fair, the authors gave all three models approximately the same total number of parameters — about 20K for music tasks and 169K for speech tasks. Because GRU has fewer matrices per unit than LSTM, it can afford more hidden units at the same parameter count: 46 GRU units vs 36 LSTM units for music, 227 vs 195 for speech.

They tested on two domains:

  • Polyphonic music modeling (4 datasets: Nottingham, JSB Chorales, MuseData, Piano-midi): each time step is a multi-hot binary representing which notes are being played simultaneously.
  • Speech signal modeling (2 Ubisoft datasets): raw one-dimensional audio. The reads 20 consecutive samples and predicts the next 10. One dataset has sequences of length 500 (short), the other 8,000 (long) — a direct test of long-range dependency capture.

All models were trained with , weight noise (σ = 0.075), and (norm ≤ 1). Learning rates were tuned by random search.

Open in Lab
Same parameter budget, different unit counts. GRU fits more units because each unit has fewer matrices.
The demo wakes as you arrive…

Results: gated units win, GRU matches LSTM

The results tell a clear two-level story:

Level 1 — Gates matter. On the challenging speech datasets, both LSTM and GRU dramatically outperformed the vanilla tanh unit. On Ubisoft B (long sequences), tanh achieved a test negative log-likelihood of 7.62 while GRU reached 0.88 and LSTM 1.26 — an order of magnitude better. The gating mechanism is not optional; it is essential for non-trivial .

Level 2 — GRU ≈ LSTM. Across all six datasets, neither gated unit consistently dominated the other. GRU won on most music datasets and on the long speech dataset; LSTM won on the short speech dataset. On music, the differences were small. The authors' conclusion: the choice between them depends on the task and dataset, but GRU's simpler design carries no systematic penalty.

The learning curves revealed an additional advantage: GRU often converged faster in both wall-clock time and number of updates, likely because fewer parameters per unit means each gradient step is cheaper and the optimization landscape is smoother.

Open in Lab
Click each dataset to see how the three units compared in test performance and convergence speed.
The demo wakes as you arrive…

The GRU in code

GRU forward step — from scratchpython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(x):
    return 1 / (1 + np.exp(-x))

def gru_step(x_t, h_prev, W_z, U_z, W_r, U_r, W_h, U_h):
    """One GRU time step. x_t: input, h_prev: previous hidden state."""
    # Update gate: how much of old state to keep
    z_t = sigmoid(W_z @ x_t + U_z @ h_prev)

    # Reset gate: how much past context to use for candidate
    r_t = sigmoid(W_r @ x_t + U_r @ h_prev)

    # Candidate activation: new content proposal
    h_tilde = np.tanh(W_h @ x_t + U_h @ (r_t * h_prev))

    # Final state: blend old and new
    h_t = (1 - z_t) * h_prev + z_t * h_tilde
    return h_t

# Compare with LSTM: no separate cell state, no output gate.
# Just two gates and one candidate — that's the whole unit.

Why this paper mattered

Before this paper, GRU had only been tested in one context (machine translation by Cho et al., 2014). This systematic evaluation across music and speech gave the community the empirical confidence to adopt GRU broadly. The practical implications were significant:

  • Faster training: fewer parameters per unit means faster gradient computation and often faster .
  • Smaller models: for the same hidden dimension, GRU uses ~25% fewer parameters than LSTM — crucial for edge deployment.
  • Design philosophy: the paper showed that architectural complexity does not automatically buy better performance. Sometimes simpler is just as good — a lesson that echoes through modern architecture design.

CitationChung, Gulcehre, Cho, Bengio. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. NIPS 2014 Workshop, 2014.

Terms in this paper