Computer Vision2015intermediate10 min read
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
الغوص في عالم المُقوِّمات: كيف تجاوزت الشبكات أداء البشر في تصنيف ImageNet
He, K. · Zhang, X. · Ren, S. · Sun, J. — ICCV
The problem
By 2015, deep networks used activations everywhere, yet their still assumed linear activations. Xavier initialization, the standard at the time, derived its scaling from a symmetry assumption that breaks when half the signal is zeroed out by ReLU. The result: networks deeper than about 8 layers could stall completely — gradients vanished and the never learned. Meanwhile, the fixed zero slope of ReLU wasted information in the negative half, and there was no principled way to decide what that slope should be.
The contribution
Two contributions that work together. First, Parametric ReLU (PReLU): the slope of the negative side becomes a learnable — each discovers whether to pass, dampen, or block negative signals, at essentially zero extra cost. Second, He initialization: by accounting for the fact that ReLU halves the of each layer's output, the correct scale is √(2/n) instead of Xavier's √(1/n). This lets networks with 30+ layers train from scratch without stalling. Together, PReLU-nets achieved 4.94% top-5 error on — surpassing reported human performance (5.1%) for the first time.
The impact
He initialization became the default for all ReLU-family networks and remains so today. It was a prerequisite for ResNet — without it, 152-layer networks could not have been trained. PReLU showed the community that activation functions need not be fixed, paving the way for GELU, Swish, and other learned or smooth activations. The paper's human-level result on ImageNet shifted public perception of what could achieve.
Imagine you're filling a chain of buckets with water, each one pouring into the next. If every bucket is slightly too small, the stream dwindles to a trickle by bucket 30. If every bucket is slightly too large, water sloshes over the sides and floods the floor.
Xavier initialization sized the buckets for a world where water flows equally in both directions. But ReLU blocks the backflow — half the water is drained at every bucket. He initialization doubles each bucket's capacity to compensate, so the stream arrives at the last bucket at full strength.
PReLU goes one step further: instead of a hard drain plug, each bucket gets an adjustable valve that learns during how much backflow to allow.
The problem: Xavier assumes linearity, but ReLU is not linear
When you stack many layers, the signal that enters the network gets multiplied by weight matrices over and over. If the variance grows at each layer, the signal explodes exponentially. If it shrinks, the signal — and its gradients — vanish to zero.
Glorot and Bengio's Xavier initialization (2010) addressed this by setting weight variance to , where is the number of inputs to a . But their derivation assumed activations are linear around zero — symmetric, passing positive and negative values equally. That assumption was valid for tanh and , the standard activations of 2010.
ReLU breaks this assumption: zeros out the entire negative half. In expectation, this halves the variance of the signal at every layer. Over layers the signal shrinks by a factor of . For a 30-layer network that is a factor of about one billionth. The network is mathematically alive but practically dead — all gradients round to zero in floating point.
PReLU: let the network choose its own activation shape
Standard ReLU has a hard rule: positive inputs pass through unchanged, negative inputs become exactly zero. But is zero always the right slope for the negative side? tried a fixed small slope (0.01), but that value was chosen by hand.
PReLU makes the negative slope a learnable parameter for each channel :
Think of it this way: ReLU is a one-way door — positive signals walk through, negative signals are turned away. Leaky ReLU is a door that lets negative signals squeeze through a fixed crack. PReLU is a door whose opening adjusts itself during training.
The for updating is simply the sum of values wherever . This adds almost no computation. For a channel-shared variant, a single parameter is shared across all channels of a layer — even that one extra parameter per layer measurably improves accuracy.
An interesting finding from the paper: early layers learn large values (around 0.6), meaning they keep much of the negative signal — they preserve information. Deeper layers learn small values, becoming more like standard ReLU — they become more selective and discriminative.
He initialization: accounting for the ReLU factor of ½
The goal of any initialization is simple: keep the variance of the signal roughly constant as it flows through the network, so it neither explodes nor vanishes.
Consider one layer. The output before activation is , where is a vector of inputs ( is the size, is the number of input channels). If the weights are zero-mean and independent:
Xavier assumed , which holds only if has zero mean. But after ReLU, , which is never negative, so its mean is not zero. Specifically, for a symmetric zero-mean input before ReLU, exactly half the values are zeroed, so:
Substituting this back:
To keep variance constant across layers, we need:
This means weights should be drawn from a zero-mean Gaussian with standard deviation . Compare this with Xavier's — the only difference is that factor of 2, but it makes all the difference for deep networks.
The same reasoning applies to the (gradient propagation), where the corresponding condition is with (using output channels instead of input channels). Either condition alone is sufficient — satisfying one approximately satisfies the other.
Extending to PReLU initialization
When the activation is PReLU with initial slope , the variance factor changes from to , because the negative side now scales by instead of being zeroed:
Why Xavier fails on deep ReLU networks
The paper demonstrates this dramatically. Consider VGG model B with 10 convolutional layers, all using 3×3 filters. The correct standard deviations per He initialization are 0.059, 0.042, 0.029, and 0.021 for layers with 64, 128, 256, and 512 filters respectively. VGG used a constant std of 0.01 throughout. The actual gradient propagated from layer 10 to layer 2 is roughly of what it should be — practically zero.
The paper trains a 30-layer model with both initializations. He initialization converges; Xavier completely stalls — the error never decreases, and monitoring confirms all gradients are vanishing. This is not a matter of tuning or patience. The signal is mathematically dead.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def he_init(shape, mode='fan_in'):
"""He (Kaiming) initialization for ReLU networks.
shape: (fan_out, fan_in) for FC, or (out_ch, in_ch, kH, kW) for conv.
"""
if len(shape) == 2:
fan_in = shape[1]
else:
fan_in = shape[1] * shape[2] * shape[3] # in_ch * kH * kW
std = np.sqrt(2.0 / fan_in)
return np.random.randn(*shape) * std
def prelu(x, a):
"""PReLU activation: f(x) = max(0,x) + a * min(0,x)"""
return np.maximum(0, x) + a * np.minimum(0, x)
# --- Demonstrate variance propagation ---
n_layers, n_units = 30, 256
# He initialization: variance stays near 1
signal = np.random.randn(1, n_units)
for _ in range(n_layers):
W = he_init((n_units, n_units))
signal = np.maximum(0, signal @ W.T) # ReLU
print(f"He — output variance after {n_layers} layers: {signal.var():.4f}")
# Xavier initialization: variance collapses
signal = np.random.randn(1, n_units)
for _ in range(n_layers):
W = np.random.randn(n_units, n_units) * np.sqrt(1.0 / n_units)
signal = np.maximum(0, signal @ W.T)
print(f"Xavier — output variance after {n_layers} layers: {signal.var():.6f}")PReLU in practice: what the learned slopes reveal
The paper trained a 14-layer model and examined the learned coefficients at every layer. Two patterns emerged:
The first — which extracts low-level features like edges and textures — learned large slopes (around 0.6). This means it keeps much of the negative signal. Early layers act as information collectors: they want to preserve as much of the raw input as possible, including patterns that happen to have negative filter responses.
Deeper layers learned progressively smaller slopes, approaching standard ReLU behavior. These layers are more discriminative — they have learned which signals matter and actively suppress the rest. The network self-organizes: preserve early, select late.
This gradient from information preservation to information selection mirrors how humans understand visual hierarchies: first gather all the details, then focus on what matters.
Surpassing human-level performance on ImageNet
Combining PReLU activations with He initialization, the authors trained models of increasing width and depth. Their best single model achieved 5.71% top-5 error — already better than all prior multi-model ensembles. A six-model ensemble reached 4.94% top-5 error, surpassing the reported human performance of 5.1% on the same dataset.
This was the first time a machine exceeded human accuracy on a large-scale visual recognition challenge. The human baseline was established by a trained annotator using a specialized interface with example images — not a casual guess. The result demonstrated that with the right training fundamentals (proper initialization, proper activations), depth and width could push accuracy to previously impossible levels.
The authors noted that human errors tended to come from fine-grained recognition (e.g., distinguishing 120 dog breeds) and class unawareness, while machine errors came from context-dependent recognition and multi-object scenes.
Why it still matters
2010
Xavier Initialization
Glorot and Bengio derive proper weight scaling for sigmoid/tanh activations by analyzing variance flow. Works well for linear activations, but not for ReLU.
2011
ReLU Gains Traction
Glorot et al. show that ReLU trains faster than sigmoid. Deep sparse rectifier networks become practical.
2015
He Initialization + PReLU
He et al. fix initialization for ReLU by adding the factor of 2 and propose learnable activation slopes. PReLU-nets achieve 4.94% top-5 error, surpassing humans.
2015
ResNet (same authors)
He initialization enabled training 152-layer networks with skip connections. ResNet won ILSVRC 2015 with 3.57% top-5 error.
2017
GELU, Swish, and Learned Activations
PReLU's idea of learnable activations inspired smooth alternatives like GELU and Swish, now standard in Transformers and modern architectures.
CitationHe, Zhang, Ren, Sun. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. ICCV, 2015.
Terms in this paper
- Weight initializationتهيئة الأوزان
- ReLUدالة الوحدة الخطية المصححة
- Activation Functionدالة التنشيط
- Vanishing Gradientاضمحلال متجهات الميل
- Varianceالتباين
- Forward Passالتمرير الأمامي
- Backward Passالتمرير الخلفي
- Leaky ReLUدالة التسريب الخطي المصححة
- ImageNetImageNet
- Top-5 Error Rateمعدل خطأ أعلى 5 احتمالات
- Batch Normalizationتسوية الدفعات الحسابية
- Deep Learningالتعلم العميق
- Convolutional Neural Network (CNN)الشبكة العصبية الالتفافية