Computer Vision2015intermediate8 min read

Deep Residual Learning for Image Recognition

التعلُّم المتبقّي العميق للتعرّف على الصور

He, K. · Zhang, X. · Ren, S. · Sun, J. — CVPR

The problem

Deeper networks should be more powerful — they have more capacity and can represent more complex functions. Yet experiments showed a baffling "": adding more layers to a plain made both training and test error increase. This was not (training error went up too). Something about the optimization landscape of very deep plain networks prevented descent from finding good solutions, even solutions as good as a shallower network could find.

The contribution

: instead of hoping each stack of layers learns the desired mapping H(x) directly, let it learn the residual F(x) = H(x) − x, then add x back at the end via a . If the optimal transformation is close to identity, learning a near-zero residual is far easier than learning an from scratch. This single change enabled training of networks with 152 layers — 8× deeper than VGG — with lower error and lower computational complexity. Bottleneck blocks (1×1 → 3×3 → 1×1 convolutions) made extreme depth practical. ResNets won every track at ILSVRC and COCO 2015.

The impact

ResNet made depth a reliable scaling axis — add layers and performance improves instead of collapses. Nearly every modern vision backbone (and many non-vision models) uses residual connections: DenseNet, EfficientNet, Vision Transformers, and even the 's own encoder and decoder blocks. The skip connection is now as fundamental to as the convolution itself.

Imagine climbing a skyscraper by staircase. At 20 floors you're fine. At 56 floors the stairs get so twisted that you actually end up lower than if you'd stopped at 20 — adding more stairs made you lose altitude.

The fix: an elevator shaft beside the stairs. At every floor, the elevator copies you up for free (the identity), and the staircase only needs to adjust the small difference from one floor to the next. Now 152 floors is easy — each floor can always fall back on the elevator if the staircase has nothing useful to add.

That elevator is the skip connection. The staircase is the residual F(x). Together they give you H(x) = F(x) + x.

The problem: deeper is not always better

By 2015, the deep learning community had a clear recipe: add more layers, get better results. VGG-16 and VGG-19 beat AlexNet's 8 layers. GoogLeNet went to 22. The obvious next step was 50, 100, 150 layers.

But something strange happened. A 56-layer plain CNN trained on CIFAR-10 had higher training error than its 20-layer counterpart. Not just higher test error — higher training error. This meant the problem was not overfitting; the deeper network simply couldn't learn as effectively.

This is the degradation problem: beyond a certain depth, adding layers makes the optimization landscape so difficult that the network can't even learn to copy the shallower network's solution (which it should be able to, since extra layers could in theory just perform identity mappings).

Open in Lab
Watch training error of a 56-layer plain network exceed a 20-layer one — the degradation problem in action.
The demo wakes as you arrive…

The idea: learn the residual, not the full mapping

He et al. made a simple but profound observation: if the optimal function for a block is close to the identity (i.e. the block should mostly pass data through), then learning H(x) = x from scratch is hard for a nonlinear stack of layers. But learning F(x) = 0 is easy — just push all weights toward zero.

The reformulates the task. Instead of learning H(x) directly, the layers learn the residual F(x) = H(x) − x, and a adds x back:

output = F(x) + x

If the block has nothing to contribute, F(x) = 0 and data passes through unchanged. If the block should transform the data, it learns only the difference from identity. Either way, optimization is easier.

y=F(x,{Wi})+xy = \mathcal{F}(x, \{W_i\}) + x
The residual block equation — F(x, {Wᵢ}) is the residual function learned by the stacked layers (e.g. two 3×3 convolutions with ReLU and batch normalization). x is the identity shortcut — added element-wise with no extra parameters.
Open in Lab
Click each part of the residual block to understand its role. The skip connection is the key innovation.
The demo wakes as you arrive…

Why skip connections fix gradient flow

The gradient of the loss with respect to an early layer x_l is:

∂L∂xl=∂L∂xL(1+∂∂xl∑i=lL−1Fi)\frac{\partial L}{\partial x_l} = \frac{\partial L}{\partial x_L}\left(1 +\frac{\partial}{\partial x_l}\sum_{i=l}^{L-1}\mathcal{F}_i\right)
derivative of the local loss — The derivative of the local loss with respect to the underlying layer input x using the chain rule.

That "1 +" is the identity path. Even if the learned residual gradients are small, the gradient can always flow backward through the skip connections undiminished. In plain networks, gradients must traverse every layer's weights — each multiplication can shrink (vanishing) or grow (exploding) the signal. The skip connection gives gradients a highway that runs straight back to any layer.

Open in Lab
Compare gradient magnitude through a plain network vs a residual network as depth grows.
The demo wakes as you arrive…

Bottleneck design: going really deep

For networks of 50+ layers, the basic two-layer residual block (two 3×3 convolutions) becomes computationally expensive. ResNet introduces the : a 1×1 convolution that reduces dimensions (e.g. 256 → 64), a 3×3 convolution that does the heavy spatial work on the smaller representation, and a 1×1 convolution that expands back (64 → 256).

The 3×3 conv now operates on 64 channels instead of 256, cutting compute by roughly 4×. This makes 101- and 152-layer ResNets practical with fewer FLOPs than VGG-16.

Open in Lab
Compare the basic 2-layer block with the 3-layer bottleneck block — same representational power, far fewer FLOPs.
The demo wakes as you arrive…

The family: from ResNet-18 to ResNet-152

ResNet is not one model but a family of architectures sharing the same residual design. The shallow variants — ResNet-18 (11.7M parameters) and ResNet-34 (21.8M) — use basic two-layer residual blocks. The deeper variants — ResNet-50 (25.6M), ResNet-101 (44.5M), and ResNet-152 (60.2M) — use bottleneck blocks. The interactive chart below shows their ImageNet top-5 error rates: from 10.92% (ResNet-18) down to 5.71% (ResNet-152).

Every variant follows the same pattern: an initial 7×7 conv + max pool, then four stages of residual blocks with doubling channels (64 → 128 → 256 → 512), and a global average pool + at the end.

Open in Lab
Watch how error drops as depth increases — the opposite of what plain networks do.
The demo wakes as you arrive…

Batch normalization: the silent partner

Every convolutional layer in ResNet is followed by (BN), which normalizes each channel's activations to zero mean and unit variance within each mini-batch. BN serves three roles:

  • Stabilizes training by preventing internal covariate shift — each layer receives inputs with consistent statistics.
  • Enables higher learning rates because normalized activations stay within the active range of ReLU, preventing both saturation and dead neurons.
  • Acts as mild regularization because the per-batch statistics add noise, reducing overfitting.

Without BN, training 152-layer networks would be impractical even with skip connections. ResNet pairs residual learning (fixing the optimization problem) with batch normalization (stabilizing the activations) — both are necessary.

A residual block in ~20 linespython

Simplified to show the idea — not the real implementation.

import numpy as np
def batch_norm(x, eps=1e-5):
    """Normalize each channel to zero mean, unit variance."""
    mean = x.mean(axis=(0, 2, 3), keepdims=True)
    var  = x.var(axis=(0, 2, 3), keepdims=True)
    return (x - mean) / np.sqrt(var + eps)

def conv3x3(x, W):
    """Placeholder for 3×3 convolution."""
    # Real implementation slides W across spatial dims
    return x  # simplified

def residual_block(x, W1, W2):
    """Core of ResNet: F(x) + x."""
    identity = x                       # save the input

    out = conv3x3(x, W1)               # first 3×3 conv
    out = batch_norm(out)               # normalize
    out = np.maximum(0, out)            # ReLU

    out = conv3x3(out, W2)              # second 3×3 conv
    out = batch_norm(out)               # normalize

    out = out + identity                # THE SKIP CONNECTION
    out = np.maximum(0, out)            # ReLU after addition
    return out

# That's it. `out + identity` is the entire innovation. # Without it: 56 layers > 20 layers in error. # With it: 152 layers sets world records.

Why it mattered

  1. 2015

    ResNet — skip connections conquer depth

    He et al. trained 152-layer networks with lower error and compute than VGG-16/19 by adding identity shortcuts. Won every track at ILSVRC and COCO 2015.

  2. 2016

    Identity Mappings in Deep Residual Networks

    He & Sun showed that moving BN and ReLU before the convolution ("pre-activation") improves gradient flow further, enabling 1001-layer ResNets on CIFAR.

  3. 2016

    ResNeXt — wider residual paths

    Xie et al. replaced single bottleneck paths with parallel grouped convolutions ("cardinality"), showing that width within residual blocks is as important as depth.

  4. 2017

    DenseNet — every layer connects to every layer

    Huang et al. extended skip connections so each layer receives feature maps from all preceding layers, maximizing feature reuse and gradient flow.

  5. 2017

    Mask R-CNN — ResNet backbone for instance segmentation

    He et al. used ResNet as the backbone for object detection and pixel-level segmentation, setting the template for modern detection architectures.

  6. 2019

    EfficientNet — scaling depth, width, and resolution together

    Tan & Le showed that scaling depth (à la ResNet), width, and input resolution in a balanced ratio outperforms scaling any one dimension alone.

  7. 2020

    Vision Transformer — residual connections in every block

    ViT replaced convolutions with self-attention for images but kept ResNet's residual connections in every encoder block — the skip connection transcended its origin architecture.

The skip connection is now everywhere. Every Transformer block — in GPT, Claude, Gemini, BERT — is a residual block. The idea that you should always give gradients an unobstructed path backward is now taken as an axiom of neural architecture design.

CitationHe, Zhang, Ren, Sun. Deep Residual Learning for Image Recognition. CVPR, 2016.

Terms in this paper