Computer Vision2015intermediate10 min read

Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification

الغوص في عالم المُقوِّمات: كيف تجاوزت الشبكات أداء البشر في تصنيف ImageNet

He, K. · Zhang, X. · Ren, S. · Sun, J. — ICCV

The problem

By 2015, deep networks used activations everywhere, yet their still assumed linear activations. Xavier initialization, the standard at the time, derived its scaling from a symmetry assumption that breaks when half the signal is zeroed out by ReLU. The result: networks deeper than about 8 layers could stall completely — gradients vanished and the never learned. Meanwhile, the fixed zero slope of ReLU wasted information in the negative half, and there was no principled way to decide what that slope should be.

The contribution

Two contributions that work together. First, Parametric ReLU (PReLU): the slope of the negative side becomes a learnable — each discovers whether to pass, dampen, or block negative signals, at essentially zero extra cost. Second, He initialization: by accounting for the fact that ReLU halves the of each layer's output, the correct scale is √(2/n) instead of Xavier's √(1/n). This lets networks with 30+ layers train from scratch without stalling. Together, PReLU-nets achieved 4.94% top-5 error on — surpassing reported human performance (5.1%) for the first time.

The impact

He initialization became the default for all ReLU-family networks and remains so today. It was a prerequisite for ResNet — without it, 152-layer networks could not have been trained. PReLU showed the community that activation functions need not be fixed, paving the way for GELU, Swish, and other learned or smooth activations. The paper's human-level result on ImageNet shifted public perception of what could achieve.

Imagine you're filling a chain of buckets with water, each one pouring into the next. If every bucket is slightly too small, the stream dwindles to a trickle by bucket 30. If every bucket is slightly too large, water sloshes over the sides and floods the floor.

Xavier initialization sized the buckets for a world where water flows equally in both directions. But ReLU blocks the backflow — half the water is drained at every bucket. He initialization doubles each bucket's capacity to compensate, so the stream arrives at the last bucket at full strength.

PReLU goes one step further: instead of a hard drain plug, each bucket gets an adjustable valve that learns during how much backflow to allow.

The problem: Xavier assumes linearity, but ReLU is not linear

When you stack many layers, the signal that enters the network gets multiplied by weight matrices over and over. If the variance grows at each layer, the signal explodes exponentially. If it shrinks, the signal — and its gradients — vanish to zero.

Glorot and Bengio's Xavier initialization (2010) addressed this by setting weight variance to 1/n1/n, where nn is the number of inputs to a . But their derivation assumed activations are linear around zero — symmetric, passing positive and negative values equally. That assumption was valid for tanh and , the standard activations of 2010.

ReLU breaks this assumption: ReLU(x)=max⁡(0,x)\text{ReLU}(x) = \max(0, x) zeros out the entire negative half. In expectation, this halves the variance of the signal at every layer. Over LL layers the signal shrinks by a factor of (1/2)L(1/2)^L. For a 30-layer network that is a factor of about one billionth. The network is mathematically alive but practically dead — all gradients round to zero in floating point.

Open in Lab
Watch the signal variance collapse layer by layer under Xavier, and stay stable under He initialization.
The demo wakes as you arrive…

PReLU: let the network choose its own activation shape

Standard ReLU has a hard rule: positive inputs pass through unchanged, negative inputs become exactly zero. But is zero always the right slope for the negative side? tried a fixed small slope (0.01), but that value was chosen by hand.

PReLU makes the negative slope a learnable parameter aia_i for each channel ii:

f(yi)={yiif yi>0ai yiif yi≤0f(y_i) = \begin{cases} y_i & \text{if } y_i > 0 \\ a_i \, y_i & \text{if } y_i \le 0 \end{cases}
PReLU — Parametric Rectified Linear Unit — When ai=0a_i = 0 this is standard ReLU. When aia_i is a fixed small constant it is Leaky ReLU. PReLU learns aia_i jointly with the rest of the network via backpropagation.

Think of it this way: ReLU is a one-way door — positive signals walk through, negative signals are turned away. Leaky ReLU is a door that lets negative signals squeeze through a fixed crack. PReLU is a door whose opening adjusts itself during training.

The for updating aia_i is simply the sum of yiy_i values wherever yi≤0y_i \le 0. This adds almost no computation. For a channel-shared variant, a single parameter aa is shared across all channels of a layer — even that one extra parameter per layer measurably improves accuracy.

An interesting finding from the paper: early layers learn large aa values (around 0.6), meaning they keep much of the negative signal — they preserve information. Deeper layers learn small aa values, becoming more like standard ReLU — they become more selective and discriminative.

Open in Lab
Drag the slider to change the negative slope a. At 0 it is ReLU, at 1 it is linear (identity).
The demo wakes as you arrive…

He initialization: accounting for the ReLU factor of ½

The goal of any initialization is simple: keep the variance of the signal roughly constant as it flows through the network, so it neither explodes nor vanishes.

Consider one layer. The output before activation is yl=Wlxl+bly_l = W_l x_l + b_l, where xlx_l is a vector of nl=k2cn_l = k^2 c inputs (kk is the size, cc is the number of input channels). If the weights are zero-mean and independent:

Var[yl]=nl Var[wl] E[xl2]\text{Var}[y_l] = n_l \, \text{Var}[w_l] \, E[x_l^2]

Xavier assumed E[xl2]=Var[xl]E[x_l^2] = \text{Var}[x_l], which holds only if xlx_l has zero mean. But after ReLU, xl=max⁡(0,yl−1)x_l = \max(0, y_{l-1}), which is never negative, so its mean is not zero. Specifically, for a symmetric zero-mean input before ReLU, exactly half the values are zeroed, so:

E[xl2]=12 Var[yl−1]E[x_l^2] = \tfrac{1}{2}\,\text{Var}[y_{l-1}]

Substituting this back:

Var[yl]=12 nl Var[wl] Var[yl−1]\text{Var}[y_l] = \tfrac{1}{2}\, n_l\, \text{Var}[w_l]\, \text{Var}[y_{l-1}]
Variance propagation through a ReLU layer — The factor ½ is ReLU's signature — it halves the effective signal. Xavier ignores it; He initialization compensates for it.

To keep variance constant across layers, we need:

12 nl Var[wl]=1⟹Var[wl]=2nl\tfrac{1}{2}\, n_l\, \text{Var}[w_l] = 1 \quad \Longrightarrow \quad \text{Var}[w_l] = \frac{2}{n_l}

This means weights should be drawn from a zero-mean Gaussian with standard deviation σ=2/nl\sigma = \sqrt{2 / n_l}. Compare this with Xavier's σ=1/nl\sigma = \sqrt{1 / n_l} — the only difference is that factor of 2, but it makes all the difference for deep networks.

The same reasoning applies to the (gradient propagation), where the corresponding condition is 12n^lVar[wl]=1\frac{1}{2}\hat{n}_l \text{Var}[w_l] = 1 with n^l=k2dl\hat{n}_l = k^2 d_l (using output channels instead of input channels). Either condition alone is sufficient — satisfying one approximately satisfies the other.

Open in Lab
Compare the convergence of a 30-layer network with He vs Xavier initialization. Xavier stalls completely.
The demo wakes as you arrive…

Extending to PReLU initialization

When the activation is PReLU with initial slope aa, the variance factor changes from 12\frac{1}{2} to 12(1+a2)\frac{1}{2}(1 + a^2), because the negative side now scales by aa instead of being zeroed:

Var[wl]=2(1+a2) nl\text{Var}[w_l] = \frac{2}{(1 + a^2)\, n_l}
PReLU-aware initialization — When a=0a = 0, this reduces to He initialization (ReLU). When a=1a = 1, the activation is linear and this reduces to Xavier initialization.

Why Xavier fails on deep ReLU networks

The paper demonstrates this dramatically. Consider VGG model B with 10 convolutional layers, all using 3×3 filters. The correct standard deviations per He initialization are 0.059, 0.042, 0.029, and 0.021 for layers with 64, 128, 256, and 512 filters respectively. VGG used a constant std of 0.01 throughout. The actual gradient propagated from layer 10 to layer 2 is roughly 1/(1.7×104)1/(1.7 \times 10^4) of what it should be — practically zero.

The paper trains a 30-layer model with both initializations. He initialization converges; Xavier completely stalls — the error never decreases, and monitoring confirms all gradients are vanishing. This is not a matter of tuning or patience. The signal is mathematically dead.

The same idea in code

He initialization and PReLU in NumPypython

Simplified to show the idea — not the real implementation.

import numpy as np

def he_init(shape, mode='fan_in'):
    """He (Kaiming) initialization for ReLU networks.
    shape: (fan_out, fan_in) for FC, or (out_ch, in_ch, kH, kW) for conv.
    """
    if len(shape) == 2:
        fan_in = shape[1]
    else:
        fan_in = shape[1] * shape[2] * shape[3]   # in_ch * kH * kW
    std = np.sqrt(2.0 / fan_in)
    return np.random.randn(*shape) * std

def prelu(x, a):
    """PReLU activation: f(x) = max(0,x) + a * min(0,x)"""
    return np.maximum(0, x) + a * np.minimum(0, x)

# --- Demonstrate variance propagation ---
n_layers, n_units = 30, 256

# He initialization: variance stays near 1
signal = np.random.randn(1, n_units)
for _ in range(n_layers):
    W = he_init((n_units, n_units))
    signal = np.maximum(0, signal @ W.T)   # ReLU
print(f"He  — output variance after {n_layers} layers: {signal.var():.4f}")

# Xavier initialization: variance collapses
signal = np.random.randn(1, n_units)
for _ in range(n_layers):
    W = np.random.randn(n_units, n_units) * np.sqrt(1.0 / n_units)
    signal = np.maximum(0, signal @ W.T)
print(f"Xavier — output variance after {n_layers} layers: {signal.var():.6f}")

PReLU in practice: what the learned slopes reveal

The paper trained a 14-layer model and examined the learned aia_i coefficients at every layer. Two patterns emerged:

The first — which extracts low-level features like edges and textures — learned large slopes (around 0.6). This means it keeps much of the negative signal. Early layers act as information collectors: they want to preserve as much of the raw input as possible, including patterns that happen to have negative filter responses.

Deeper layers learned progressively smaller slopes, approaching standard ReLU behavior. These layers are more discriminative — they have learned which signals matter and actively suppress the rest. The network self-organizes: preserve early, select late.

This gradient from information preservation to information selection mirrors how humans understand visual hierarchies: first gather all the details, then focus on what matters.

Open in Lab
Learned PReLU slopes by layer depth — early layers preserve (high a), deep layers select (low a).
The demo wakes as you arrive…

Surpassing human-level performance on ImageNet

Combining PReLU activations with He initialization, the authors trained models of increasing width and depth. Their best single model achieved 5.71% top-5 error — already better than all prior multi-model ensembles. A six-model ensemble reached 4.94% top-5 error, surpassing the reported human performance of 5.1% on the same dataset.

This was the first time a machine exceeded human accuracy on a large-scale visual recognition challenge. The human baseline was established by a trained annotator using a specialized interface with example images — not a casual guess. The result demonstrated that with the right training fundamentals (proper initialization, proper activations), depth and width could push accuracy to previously impossible levels.

The authors noted that human errors tended to come from fine-grained recognition (e.g., distinguishing 120 dog breeds) and class unawareness, while machine errors came from context-dependent recognition and multi-object scenes.

Open in Lab
ImageNet top-5 error rate milestones — PReLU-nets crossed the human line.
The demo wakes as you arrive…

Why it still matters

  1. 2010

    Xavier Initialization

    Glorot and Bengio derive proper weight scaling for sigmoid/tanh activations by analyzing variance flow. Works well for linear activations, but not for ReLU.

  2. 2011

    ReLU Gains Traction

    Glorot et al. show that ReLU trains faster than sigmoid. Deep sparse rectifier networks become practical.

  3. 2015

    He Initialization + PReLU

    He et al. fix initialization for ReLU by adding the factor of 2 and propose learnable activation slopes. PReLU-nets achieve 4.94% top-5 error, surpassing humans.

  4. 2015

    ResNet (same authors)

    He initialization enabled training 152-layer networks with skip connections. ResNet won ILSVRC 2015 with 3.57% top-5 error.

  5. 2017

    GELU, Swish, and Learned Activations

    PReLU's idea of learnable activations inspired smooth alternatives like GELU and Swish, now standard in Transformers and modern architectures.

CitationHe, Zhang, Ren, Sun. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. ICCV, 2015.

Terms in this paper