Deep Learning2010foundational8 min read

Understanding the Difficulty of Training Deep Feedforward Neural Networks

لماذا يصعب تدريب الشبكات العصبية العميقة ذات التغذية الأمامية؟

Glorot, X. · Bengio, Y. — AISTATS

The problem

Before 2006, deep neural networks from random initialization was widely considered impractical. By 2010, -wise had shown that deep networks could succeed, but the underlying reason remained unclear: why did ordinary descent fail when starting from random weights? The culprit turned out to be unstable signal propagation. As data passed through many layers, activations and gradients tended to shrink or grow uncontrollably, pushing neurons into and making learning increasingly difficult. Every additional layer amplified the problem.

The contribution

Glorot and Bengio showed that successful deep learning requires preserving the scale of information as it moves both forward through the network and backward through gradient updates. Their solution was a carefully designed initialization scheme that balances the number of incoming and outgoing connections for each layer, keeping activations and gradients within a stable range. This method, later known as Xavier initialization, made it possible to train substantially deeper networks directly from random initialization. The paper also highlighted the importance of activation functions and showed that some saturate more easily than others, leading to slower and less stable learning.

The impact

Xavier initialization became the default initialization for nearly every deep network. It removed the need for greedy pretraining and enabled direct end-to-end training of deep architectures. Its -preservation principle directly inspired He initialization (for ) and informed — the two techniques that, combined with Xavier's insight, made modern deep learning practical.

Imagine a row of people passing buckets of water down a chain. If someone has arms that are too strong they fling the bucket too hard and splash everything; too weak and the bucket barely reaches the next person. After ten handoffs the chain either floods or runs dry.

Xavier initialization sizes everyone's arms so the amount of water stays the same from start to end — no matter how long the chain is.

That "water" is information flowing forward through activations and gradients flowing backward through . The "arm strength" is the variance of the weights.

The problem: activations die or explode as networks deepen

A with LL layers computes its output by passing input through a chain of linear transformations and nonlinearities. During backpropagation, gradients travel the reverse path. Two things can go wrong:

  • Vanishing activations. If weights are too small, each layer shrinks the signal. By layer 5 the activations are near zero, the nonlinearity is stuck in its linear regime, and the network learns nothing useful.

  • Exploding/saturating activations. If weights are too large, activations hit the flat tails of or where the derivative ≈ 0. Gradients vanish even though the activations are large — a subtle trap the paper calls saturation.

The standard initialization of the era was Wij∼U[−1/nin,  1/nin]W_{ij} \sim U[-1/\sqrt{n_{in}},\; 1/\sqrt{n_{in}}], which only considered the fan-in. Glorot and Bengio showed this is fundamentally lopsided: it ignores what happens during backpropagation.

Open in Lab
Watch how signal variance changes across 5 layers. Toggle between standard init and Xavier init to see the difference.
The demo wakes as you arrive…

Diagnosing the disease: activation histograms

Glorot and Bengio trained networks with 1–5 hidden layers (1000 units each) on Shapeset-3×2 and monitored activations layer by layer. With tanh and standard init, layer 1 quickly pushed activations toward ±1 — the saturated region where tanh's gradient is near zero. Deeper layers followed: by epoch 100, the top layer was almost entirely saturated while the bottom layer had barely begun learning.

Think of it as a highway with toll booths. The first toll booth (layer 1) sets its barrier so high that traffic (gradients) can barely trickle through. Every subsequent toll booth compounds the blockage. By the time you reach the entrance (layer 1 from the gradient's perspective), there is no signal left to direct learning.

The softsign activation f(x)=x/(1+∣x∣)f(x) = x/(1+|x|) avoids this hard saturation because its tails decay polynomially rather than exponentially. The paper showed softsign networks train more stably, though the community ultimately adopted ReLU — which eliminates saturation entirely on its positive side.

Open in Lab
Compare activation distributions across layers for sigmoid, tanh, and softsign. See how saturation spreads with depth under standard init.
The demo wakes as you arrive…

The derivation: keeping variance constant

Consider a single layer computing si=Wizi−1+bis^i = W^i z^{i-1} + b^i where zi=f(si)z^{i} = f(s^i). Assume the is symmetric with f′(0)=1f'(0) = 1 (true for tanh and softsign near zero), weights have zero mean and are independent of the input, and the input also has zero mean. Under these assumptions, the variance of layer ii's output is a product of all preceding layer variances:

Var[zi]=Var[x]∏i′=0i−1ni′ Var[Wi′]\text{Var}[z^i] = \text{Var}[x] \prod_{i'=0}^{i-1} n_{i'}\,\text{Var}[W^{i'}]
Forward propagation variance — If n·Var[W] > 1 for any layer, variance grows exponentially — activations explode. If n·Var[W] < 1, variance shrinks — activations vanish. We need n·Var[W] = 1 at every layer.

The same analysis applies in reverse. During backpropagation, the gradient variance across layers is:

Var ⁣[∂L∂si]=Var ⁣[∂L∂sd]∏i′=idni′+1 Var[Wi′]\text{Var}\!\left[\frac{\partial \mathcal{L}}{\partial s^i}\right] = \text{Var}\!\left[\frac{\partial \mathcal{L}}{\partial s^d}\right] \prod_{i'=i}^{d} n_{i'+1}\,\text{Var}[W^{i'}]
Backward propagation gradient variance — To keep gradients stable, we need n_{out}·Var[W] = 1 at every layer. But the forward condition requires n_{in}·Var[W] = 1. These two constraints collide unless the layer is square.
Var[Wi]=2ni+ni+1\text{Var}[W^i] = \frac{2}{n_{i} + n_{i+1}}
Xavier initialization — the variance rule — Average the fan-in and fan-out requirements. For a uniform distribution this becomes W ~ U[-√(6/(nᵢₙ+nₒᵤₜ)), √(6/(nᵢₙ+nₒᵤₜ))]. This single formula keeps both forward activations and backward gradients in a stable variance band.
Open in Lab
Adjust n_in and n_out to see how Xavier initialization adapts the weight range compared to standard 1/√n init.
The demo wakes as you arrive…

Gradient flow: what the paper revealed

The paper's most striking contribution beyond the formula was a set of diagnostic visualizations. By plotting activation histograms and gradient standard deviations layer by layer during training, Glorot and Bengio created an X-ray of what happens inside a deep network:

With standard init + tanh: layer 1 saturates first (activations cluster at ±1), then layers 2, 3, 4 follow like dominoes. Gradients in the bottom layers collapse to near zero — those layers effectively stop learning. The top layer learns alone, making the 5-layer network behave like a 1-layer network.

With Xavier init + tanh: activations stay centered near zero in the linear regime of tanh across all layers. Gradient standard deviations remain similar across layers throughout training — the signal reaches every layer equally.

With softsign: even under standard init, softsign's polynomial tails prevent hard saturation. Activations drift toward the tails but never lock there. Gradients are healthier, though not as uniform as Xavier + tanh.

Open in Lab
Compare gradient magnitudes across 5 layers. Standard init creates a gradient cliff; Xavier init keeps a flat gradient landscape.
The demo wakes as you arrive…

Activation functions: sigmoid, tanh, and softsign

The paper compared three activation functions and their behavior under deep training:

Sigmoid σ(x)=1/(1+e−x)\sigma(x) = 1/(1+e^{-x}): outputs in (0,1)(0,1) with mean ≈ 0.5, not zero. This non-zero mean causes each layer to systematically shift activations, compounding through depth. Glorot and Bengio showed that sigmoid networks showed the worst saturation patterns.

Tanh tanh⁡(x)\tanh(x): outputs in (−1,1)(-1, 1) with mean 0 — better centered. But its gradient vanishes exponentially at the tails: tanh⁡′(x)→0\tanh'(x) \to 0 as ∣x∣→∞|x| \to \infty. Under standard init, the first layer saturates within the first few epochs.

Softsign x/(1+∣x∣)x/(1+|x|): outputs in (−1,1)(-1, 1) with mean 0, like tanh, but the tails decay as 1/x21/x^2 rather than exponentially. This means the gradient shrinks slowly, not suddenly — a more graceful degradation that prevents the hard saturation trap.

Open in Lab
Drag the x slider and compare how sigmoid, tanh, and softsign respond — and how their derivatives differ at the extremes.
The demo wakes as you arrive…

Xavier initialization in code

Xavier initialization — NumPy and PyTorchpython

Simplified to show the idea — not the real implementation.

import numpy as np

def xavier_uniform(n_in, n_out):
    """Xavier/Glorot uniform initialization.
    Keeps Var[activation] ≈ Var[input] across layers."""
    limit = np.sqrt(6.0 / (n_in + n_out))
    return np.random.uniform(-limit, limit, size=(n_in, n_out))

def xavier_normal(n_in, n_out):
    """Xavier/Glorot normal initialization.
    Same variance target, Gaussian distribution."""
    std = np.sqrt(2.0 / (n_in + n_out))
    return np.random.randn(n_in, n_out) * std

# Example: 3-layer network (784 → 256 → 128 → 10)
W1 = xavier_uniform(784, 256)   # Var ≈ 2/(784+256) = 0.00192
W2 = xavier_uniform(256, 128)   # Var ≈ 2/(256+128) = 0.00521
W3 = xavier_uniform(128, 10)    # Var ≈ 2/(128+10)  = 0.01449

# In PyTorch it's built-in:
# torch.nn.init.xavier_uniform_(layer.weight)
# torch.nn.init.xavier_normal_(layer.weight)

Why it changed everything

The initialization lineage

  1. 2006

    Greedy layer-wise pretraining

    Hinton showed deep networks could be trained by first pretraining each layer as an RBM. This proved depth works but required an expensive two-phase pipeline.

  2. 2010

    Xavier initialization (this paper)

    Glorot & Bengio derived a principled initialization that made pretraining unnecessary for tanh/sigmoid networks. One formula, one line of code, enormous impact.

  3. 2011

    Deep Sparse Rectifier Networks (ReLU)

    Glorot, Bordes & Bengio showed that ReLU networks train well and produce sparse activations. ReLU eliminated saturation on the positive side but created the "dying ReLU" problem.

  4. 2015

    He initialization

    He et al. adapted Xavier's principle for ReLU by noting that ReLU zeroes half its input, doubling the needed variance to Var[W] = 2/n_in.

  5. 2015

    Batch Normalization

    Ioffe & Szegedy took the variance-preservation idea further: instead of only fixing initialization, normalize activations at every layer during every forward pass.

CitationGlorot, Bengio. Understanding the Difficulty of Training Deep Feedforward Neural Networks. AISTATS, 2010.

Terms in this paper