Deep Learning2010foundational8 min read
Understanding the Difficulty of Training Deep Feedforward Neural Networks
لماذا يصعب تدريب الشبكات العصبية العميقة ذات التغذية الأمامية؟
Glorot, X. · Bengio, Y. — AISTATS
The problem
Before 2006, deep neural networks from random initialization was widely considered impractical. By 2010, -wise had shown that deep networks could succeed, but the underlying reason remained unclear: why did ordinary descent fail when starting from random weights? The culprit turned out to be unstable signal propagation. As data passed through many layers, activations and gradients tended to shrink or grow uncontrollably, pushing neurons into and making learning increasingly difficult. Every additional layer amplified the problem.
The contribution
Glorot and Bengio showed that successful deep learning requires preserving the scale of information as it moves both forward through the network and backward through gradient updates. Their solution was a carefully designed initialization scheme that balances the number of incoming and outgoing connections for each layer, keeping activations and gradients within a stable range. This method, later known as Xavier initialization, made it possible to train substantially deeper networks directly from random initialization. The paper also highlighted the importance of activation functions and showed that some saturate more easily than others, leading to slower and less stable learning.
The impact
Xavier initialization became the default initialization for nearly every deep network. It removed the need for greedy pretraining and enabled direct end-to-end training of deep architectures. Its -preservation principle directly inspired He initialization (for ) and informed — the two techniques that, combined with Xavier's insight, made modern deep learning practical.
Imagine a row of people passing buckets of water down a chain. If someone has arms that are too strong they fling the bucket too hard and splash everything; too weak and the bucket barely reaches the next person. After ten handoffs the chain either floods or runs dry.
Xavier initialization sizes everyone's arms so the amount of water stays the same from start to end — no matter how long the chain is.
That "water" is information flowing forward through activations and gradients flowing backward through . The "arm strength" is the variance of the weights.
The problem: activations die or explode as networks deepen
A with layers computes its output by passing input through a chain of linear transformations and nonlinearities. During backpropagation, gradients travel the reverse path. Two things can go wrong:
-
Vanishing activations. If weights are too small, each layer shrinks the signal. By layer 5 the activations are near zero, the nonlinearity is stuck in its linear regime, and the network learns nothing useful.
-
Exploding/saturating activations. If weights are too large, activations hit the flat tails of or where the derivative ≈ 0. Gradients vanish even though the activations are large — a subtle trap the paper calls saturation.
The standard initialization of the era was , which only considered the fan-in. Glorot and Bengio showed this is fundamentally lopsided: it ignores what happens during backpropagation.
Diagnosing the disease: activation histograms
Glorot and Bengio trained networks with 1–5 hidden layers (1000 units each) on Shapeset-3×2 and monitored activations layer by layer. With tanh and standard init, layer 1 quickly pushed activations toward ±1 — the saturated region where tanh's gradient is near zero. Deeper layers followed: by epoch 100, the top layer was almost entirely saturated while the bottom layer had barely begun learning.
Think of it as a highway with toll booths. The first toll booth (layer 1) sets its barrier so high that traffic (gradients) can barely trickle through. Every subsequent toll booth compounds the blockage. By the time you reach the entrance (layer 1 from the gradient's perspective), there is no signal left to direct learning.
The softsign activation avoids this hard saturation because its tails decay polynomially rather than exponentially. The paper showed softsign networks train more stably, though the community ultimately adopted ReLU — which eliminates saturation entirely on its positive side.
The derivation: keeping variance constant
Consider a single layer computing where . Assume the is symmetric with (true for tanh and softsign near zero), weights have zero mean and are independent of the input, and the input also has zero mean. Under these assumptions, the variance of layer 's output is a product of all preceding layer variances:
The same analysis applies in reverse. During backpropagation, the gradient variance across layers is:
Gradient flow: what the paper revealed
The paper's most striking contribution beyond the formula was a set of diagnostic visualizations. By plotting activation histograms and gradient standard deviations layer by layer during training, Glorot and Bengio created an X-ray of what happens inside a deep network:
With standard init + tanh: layer 1 saturates first (activations cluster at ±1), then layers 2, 3, 4 follow like dominoes. Gradients in the bottom layers collapse to near zero — those layers effectively stop learning. The top layer learns alone, making the 5-layer network behave like a 1-layer network.
With Xavier init + tanh: activations stay centered near zero in the linear regime of tanh across all layers. Gradient standard deviations remain similar across layers throughout training — the signal reaches every layer equally.
With softsign: even under standard init, softsign's polynomial tails prevent hard saturation. Activations drift toward the tails but never lock there. Gradients are healthier, though not as uniform as Xavier + tanh.
Activation functions: sigmoid, tanh, and softsign
The paper compared three activation functions and their behavior under deep training:
Sigmoid : outputs in with mean ≈ 0.5, not zero. This non-zero mean causes each layer to systematically shift activations, compounding through depth. Glorot and Bengio showed that sigmoid networks showed the worst saturation patterns.
Tanh : outputs in with mean 0 — better centered. But its gradient vanishes exponentially at the tails: as . Under standard init, the first layer saturates within the first few epochs.
Softsign : outputs in with mean 0, like tanh, but the tails decay as rather than exponentially. This means the gradient shrinks slowly, not suddenly — a more graceful degradation that prevents the hard saturation trap.
Xavier initialization in code
Simplified to show the idea — not the real implementation.
import numpy as np
def xavier_uniform(n_in, n_out):
"""Xavier/Glorot uniform initialization.
Keeps Var[activation] ≈ Var[input] across layers."""
limit = np.sqrt(6.0 / (n_in + n_out))
return np.random.uniform(-limit, limit, size=(n_in, n_out))
def xavier_normal(n_in, n_out):
"""Xavier/Glorot normal initialization.
Same variance target, Gaussian distribution."""
std = np.sqrt(2.0 / (n_in + n_out))
return np.random.randn(n_in, n_out) * std
# Example: 3-layer network (784 → 256 → 128 → 10)
W1 = xavier_uniform(784, 256) # Var ≈ 2/(784+256) = 0.00192
W2 = xavier_uniform(256, 128) # Var ≈ 2/(256+128) = 0.00521
W3 = xavier_uniform(128, 10) # Var ≈ 2/(128+10) = 0.01449
# In PyTorch it's built-in:
# torch.nn.init.xavier_uniform_(layer.weight)
# torch.nn.init.xavier_normal_(layer.weight)Why it changed everything
The initialization lineage
2006
Greedy layer-wise pretraining
Hinton showed deep networks could be trained by first pretraining each layer as an RBM. This proved depth works but required an expensive two-phase pipeline.
2010
Xavier initialization (this paper)
Glorot & Bengio derived a principled initialization that made pretraining unnecessary for tanh/sigmoid networks. One formula, one line of code, enormous impact.
2011
Deep Sparse Rectifier Networks (ReLU)
Glorot, Bordes & Bengio showed that ReLU networks train well and produce sparse activations. ReLU eliminated saturation on the positive side but created the "dying ReLU" problem.
2015
He initialization
He et al. adapted Xavier's principle for ReLU by noting that ReLU zeroes half its input, doubling the needed variance to Var[W] = 2/n_in.
2015
Batch Normalization
Ioffe & Szegedy took the variance-preservation idea further: instead of only fixing initialization, normalize activations at every layer during every forward pass.
CitationGlorot, Bengio. Understanding the Difficulty of Training Deep Feedforward Neural Networks. AISTATS, 2010.
Terms in this paper
- Activation Functionدالة التنشيط
- Varianceالتباين
- Gradientالتدرج التفاضلي
- Backpropagationالتحديث التراجعي
- Sigmoidدالة سيجمويد
- Tanhدالة الظل الزائدي
- Softmaxسوفت ماكس
- Weightالوزن البنيوي
- Layerالطبقة الحسابية
- Hidden Layerالطبقة الخفية
- Normalizationالمعايرة القياسية للبيانات
- Batch Normalizationتسوية الدفعات الحسابية
- Saturationالإشباع
- Convergenceالتقارب الحسابي