Core ML2013intermediate11 min read

On the Difficulty of Training Recurrent Neural Networks

صعوبة تدريب الشبكات العصبية التكرارية

Pascanu, R. · Mikolov, T. · Bengio, Y. — ICML

The problem

Recurrent neural networks unroll across many time steps during , creating long chains of multiplications. If the largest of the recurrent matrix exceeds 1, norms explode exponentially — a single bad step can catapult parameters far from any useful solution. If eigenvalues are below 1, gradients vanish and the network cannot learn long-range dependencies. By 2013, these twin pathologies were the primary obstacle blocking RNNs from scaling to real-world tasks.

The contribution

The paper analyzes exploding and vanishing gradients from three perspectives — analytical (eigenvalue conditions on the ), geometric (high-curvature walls in the surface), and dynamical systems (bifurcations and basin-of-attraction crossings). It then proposes gradient clipping: whenever the gradient norm exceeds a threshold, rescale the entire gradient to that threshold. For vanishing gradients, it introduces a soft term that encourages the Jacobian to preserve gradient norm across time steps. Together, + clipping + regularization (SGD-CR) solves pathological long-dependency benchmarks up to 200 steps and generalizes to sequences twice as long.

The impact

became one of the most universally adopted tricks in . Every major framework — PyTorch, TensorFlow, JAX — ships it as a built-in utility. Transformers, LSTMs, GANs, and virtually every trained with SGD-family optimizers uses clipping. The paper's analysis of error surface geometry also laid groundwork for understanding landscapes that informed Adam, warmup, and modern training stabilization techniques.

Imagine you're driving through a mountain valley on a narrow road. On one side is a sheer cliff — one wrong turn of the wheel and you plummet. On the other side the road gently slopes into fog, and if you drift that way you slowly lose speed until you stop, stranded.

Training an is driving that road. The cliff is the : the error surface suddenly becomes a near-vertical wall, and a normal-sized gradient step launches you off the edge. The foggy slope is the : the signal fades so gradually you can't steer at all.

This paper's solution? Speed-limit your steering. If the wheel jerks toward the cliff, cap the turn angle — you stay on the road. That's gradient clipping. For the fog, install headlights that amplify the road markings — a regularizer that keeps the signal visible across long stretches.

The problem: why RNNs break on long sequences

A recurrent neural network processes sequences by maintaining a that gets updated at every time step. To learn, we unroll the network across time and backpropagate — a process called (BPTT). The gradient of the at time tt with respect to parameters at an earlier step kk involves a product of Jacobian matrices, one for each step between kk and tt:

∂xt∂xk=∏t≥i>k∂xi∂xi−1=∏t≥i>kWrec⊤ diag ⁣(σ′(xi−1))\frac{\partial \mathbf{x}_t}{\partial \mathbf{x}_k} = \prod_{t \geq i > k} \frac{\partial \mathbf{x}_i}{\partial \mathbf{x}_{i-1}} = \prod_{t \geq i > k} \mathbf{W}_{\text{rec}}^\top \, \text{diag}\!\left(\sigma'(\mathbf{x}_{i-1})\right)
Chain of Jacobians — the engine of both exploding and vanishing gradients — During backpropagation through time, gradient information must pass through every recurrent step between an earlier state and a later one. Because each step scales the gradient, repeatedly applying many such transformations can cause the signal to shrink toward zero or grow uncontrollably. This phenomenon is the root cause of the vanishing- and exploding-gradient problems that make long-range learning difficult in traditional recurrent neural networks.

Think of it like compound interest. If each Jacobian factor amplifies the signal by a factor λ>1\lambda > 1, after t−kt - k steps the signal grows as λt−k\lambda^{t-k} — exponential explosion. If λ<1\lambda < 1, it shrinks as λt−k\lambda^{t-k} — exponential decay. A product of 100 matrices with λ=1.1\lambda = 1.1 amplifies by 1.1100≈13,7811.1^{100} \approx 13{,}781. With λ=0.9\lambda = 0.9 it shrinks to 0.9100≈0.0000270.9^{100} \approx 0.000027.

More precisely, if λ1\lambda_1 is the largest eigenvalue of Wrec\mathbf{W}_{\text{rec}} and γ\gamma is the upper bound on ∣σ′(x)∣|\sigma'(x)| (for tanh, γ=1\gamma = 1; for sigmoid, γ=14\gamma = \tfrac{1}{4}), then:

  • Sufficient condition for vanishing: λ1<1γ\lambda_1 < \frac{1}{\gamma} — long-term components decay exponentially.
  • Necessary condition for exploding: λ1>1γ\lambda_1 > \frac{1}{\gamma} — without this, gradients cannot explode.
Open in Lab
Drag the eigenvalue slider to see how gradient magnitude changes across 50 time steps. Above 1.0, it rockets upward; below 1.0, it collapses to zero.
The demo wakes as you arrive…

The geometric view: walls in the error surface

The paper reveals a striking geometric pattern. When gradients explode, they do so along a specific direction v\mathbf{v} — the of the largest eigenvalue. The curvature of the error surface explodes along the same direction, creating a steep wall in the loss landscape.

Imagine a wide, gentle valley with a single sheer cliff face cutting across it. Standard SGD computes a step proportional to the gradient. Near the cliff, that gradient is enormous, and the step launches you clear across the valley — or worse, out of it entirely. Training becomes unstable: the parameters bounce wildly instead of converging.

The key insight is that the valley is wide on both sides of the wall. If instead of taking the enormous step, you cap the step size and take a small step in the gradient direction, you land back in the smooth region beside the wall, where SGD can resume normal exploration.

Open in Lab
The 3D surface shows a valley with a steep wall. Toggle clipping to see how it keeps the optimization trajectory on the smooth side.
The demo wakes as you arrive…

The dynamical systems view: attractors and bifurcations

The authors borrow tools from dynamical systems theory to explain when gradient explosions happen. An RNN's hidden state behaves like a dynamical system: given a set of parameters, the state converges toward attractors — stable equilibria. The space is divided into basins of attraction, each pulling the state toward a different .

As parameters change during training, the landscape of attractors shifts smoothly — almost everywhere. At special points called boundaries, attractors appear, disappear, or change shape abruptly. When the training path crosses a boundary between two basins of attraction, a tiny parameter change causes a huge change in the hidden state, which means the gradient explodes.

Think of it like a marble rolling on a landscape with two bowls. Most of the time, tilting the landscape a little just shifts where the marble settles. But at the ridge between the bowls, an infinitesimal tilt sends the marble to an entirely different bowl — a discontinuous jump. That ridge is the bifurcation boundary, and the jump is the gradient explosion.

Open in Lab
Drag the bias slider to explore the bifurcation diagram. Watch how attractors appear and disappear, and how crossing a boundary causes a discontinuous jump in the state.
The demo wakes as you arrive…

The solution: gradient clipping

The geometric insight — a wide valley with a single steep wall — suggests an elegant fix. When the gradient norm spikes (indicating you've hit the wall), don't take the enormous step. Instead, rescale the gradient so its norm equals a fixed threshold, then step. You preserve the direction (still descending) but limit the magnitude (staying near the wall instead of flying across the valley).

The algorithm is strikingly simple:

g^←∂E∂θ,if ∥g^∥≥τ then g^←τ∥g^∥g^\hat{\mathbf{g}} \leftarrow \frac{\partial E}{\partial \boldsymbol{\theta}}, \quad \text{if } \|\hat{\mathbf{g}}\| \geq \tau \text{ then } \hat{\mathbf{g}} \leftarrow \frac{\tau}{\|\hat{\mathbf{g}}\|} \hat{\mathbf{g}}
Gradient norm clipping — the complete algorithm in one line — Compute the gradient. If its norm exceeds threshold τ, rescale it to have norm τ. Direction preserved, magnitude bounded. That's it.

Why does something so simple work? Because of the valley geometry. The wall is narrow — the explosion happens in a localized region of parameter space. On both sides of the wall, the error surface is smooth and well-behaved. A clipped step from the wall just pushes you back to the smooth side, where standard SGD works fine.

Compared to second-order methods like -Free optimization, clipping has three advantages: it is trivially cheap to compute, it handles sudden curvature changes (which smoothed Hessian estimates miss), and it works even when the gradient and curvature grow at different rates (which would cause a second-order correction to itself explode).

Open in Lab
Watch SGD with and without clipping navigate a 2D loss surface with a cliff. The clipped version stays stable; the unclipped one flies off.
The demo wakes as you arrive…
Gradient clipping in NumPy — the full implementationpython

Simplified to show the idea — not the real implementation.

import numpy as np

def clip_gradient_norm(gradient, max_norm=1.0):
    """Clip gradient vector when its norm exceeds max_norm.

    This is the exact algorithm from Pascanu et al. (2013):
    preserve direction, bound magnitude.
    """
    grad_norm = np.linalg.norm(gradient)
    if grad_norm > max_norm:
        gradient = (max_norm / grad_norm) * gradient
    return gradient

# Example: gradient with norm 50 gets rescaled to norm 1
g = np.random.randn(100) * 50  # a "gradient explosion"
g_clipped = clip_gradient_norm(g, max_norm=1.0)
print(f"Before: {np.linalg.norm(g):.1f}")   # ~500
print(f"After:  {np.linalg.norm(g_clipped):.1f}")  # 1.0
# Direction unchanged, magnitude bounded. Training continues.

Vanishing gradients: the regularization approach

Clipping tames explosions, but what about the opposite problem — vanishing gradients? When gradient components shrink exponentially, the model cannot learn dependencies between distant time steps.

The paper introduces a regularization term Ω\Omega that penalizes the Jacobian whenever it fails to preserve the gradient norm as the error signal travels backward through time. The idea: force the back-propagated error signal to neither grow nor shrink as it passes through each time step.

Ω=∑k(∥∂E∂xk+1∂xk+1∂xk∥/∥∂E∂xk+1∥−1)2\Omega = \sum_k \left( \left\| \frac{\partial E}{\partial \mathbf{x}_{k+1}} \frac{\partial \mathbf{x}_{k+1}}{\partial \mathbf{x}_k} \right\| \bigg/ \left\| \frac{\partial E}{\partial \mathbf{x}_{k+1}} \right\| - 1 \right)^2
Vanishing gradient regularizer — preserve the error signal's magnitude — For each time step, measure how much the Jacobian changes the norm of the back-propagated error. Penalize deviations from 1 (perfect preservation). The penalty is directional — it only cares about the relevant error direction, not all eigenvalues.

Two design choices make this regularizer practical. First, it only penalizes the Jacobian's behavior in the direction of the current error gradient, not across all directions. This avoids the expensive requirement of making all eigenvalues close to 1. Second, the paper uses only the "immediate" for efficiency — treating the hidden state and error as constants when differentiating Ω\Omega with respect to Wrec\mathbf{W}_{\text{rec}}.

There is a natural tension between the two solutions. Preventing vanishing gradients means keeping the model sensitive to distant inputs — but this also means operating near basin boundaries where explosions are more likely. The paper combines both: the regularizer pushes the model toward richer regimes where long-term memory is possible, and clipping acts as a safety net when the inevitable explosions occur.

Experimental results

The combined approach — SGD-CR (SGD + Clipping + Regularization) — was tested on both synthetic pathological benchmarks and real-world tasks.

On the temporal order problem, which requires remembering symbols across long sequences, SGD alone fails beyond length 20. SGD with clipping (SGD-C) improves but still struggles. SGD-CR achieves 100% success on sequences up to 200 steps — and generalizes to sequences twice as long (400 steps) without retraining.

On real-world tasks, clipping alone improved performance on polyphonic music prediction (Piano-midi.de, Nottingham, MuseData) and character-level language modeling (Penn Treebank). Adding the regularizer brought further gains, especially on the modified task of predicting the 5th-next character, where long-range dependencies are more important.

Open in Lab
Compare success rates of SGD, SGD-C, and SGD-CR on the temporal order problem across different sequence lengths.
The demo wakes as you arrive…

Context: what came before

The vanishing gradient problem was first identified by Bengio et al. (1994), which showed that gradient-based learning in RNNs becomes increasingly difficult as the time gap between relevant events grows. That paper provided the theoretical foundation — eigenvalue conditions for gradient decay — that Pascanu et al. extend to the exploding case.

(1997) solved the vanishing gradient architecturally by introducing gated memory cells with a constant error carousel. However, LSTM did not address exploding gradients explicitly. Hessian-Free optimization (Martens and Sutskever, 2011) handled both problems through second-order information but was far more expensive.

Gradient clipping bridges the gap: it is cheap enough for standard SGD, handles explosions that architectural changes like LSTM do not address, and complements rather than replaces these other approaches.

Why it still matters

Gradient clipping outlived the RNN era it was designed for. When Transformers replaced RNNs as the dominant architecture, clipping came along — not because Transformers have the same Jacobian chain problem, but because deep networks of any kind can encounter sudden gradient spikes during training due to unusual mini-batches, learning rate schedules, or loss landscape geometry.

Today, torch.nn.utils.clip_grad_norm_ is called in virtually every large-scale training loop. GPT, Claude, Gemini, LLaMA — all use gradient clipping. It is one of those rare ideas that is so simple and so effective that it became invisible infrastructure, the kind of thing every practitioner uses without thinking about the paper that introduced it.

  1. 1994

    Bengio et al. — Vanishing Gradient

    First formal analysis showing that learning long-term dependencies with gradient descent is fundamentally difficult. Eigenvalue conditions for gradient decay established.

  2. 1997

    LSTM — Architectural Fix

    Hochreiter and Schmidhuber introduced gated memory cells to solve vanishing gradients architecturally, but exploding gradients remained unaddressed.

  3. 2013

    This Paper — Gradient Clipping

    Pascanu, Mikolov, and Bengio unified the analysis of both problems and proposed gradient norm clipping plus vanishing gradient regularization. Simple, cheap, effective.

  4. 2014

    Seq2Seq with Clipping

    Sutskever et al. used gradient clipping as a key ingredient in the first successful neural machine translation system, demonstrating its importance beyond benchmarks.

  5. 2017

    Transformers Adopt Clipping

    The original Transformer paper used gradient clipping in training. It survived the architectural revolution from RNNs to attention — proving its universality.

CitationPascanu, Mikolov, Bengio. On the Difficulty of Training Recurrent Neural Networks. ICML, 2013.

Terms in this paper