Computer Vision2015intermediate8 min read
Deep Residual Learning for Image Recognition
التعلُّم المتبقّي العميق للتعرّف على الصور
He, K. · Zhang, X. · Ren, S. · Sun, J. — CVPR
The problem
Deeper networks should be more powerful — they have more capacity and can represent more complex functions. Yet experiments showed a baffling "": adding more layers to a plain made both training and test error increase. This was not (training error went up too). Something about the optimization landscape of very deep plain networks prevented descent from finding good solutions, even solutions as good as a shallower network could find.
The contribution
: instead of hoping each stack of layers learns the desired mapping H(x) directly, let it learn the residual F(x) = H(x) − x, then add x back at the end via a . If the optimal transformation is close to identity, learning a near-zero residual is far easier than learning an from scratch. This single change enabled training of networks with 152 layers — 8× deeper than VGG — with lower error and lower computational complexity. Bottleneck blocks (1×1 → 3×3 → 1×1 convolutions) made extreme depth practical. ResNets won every track at ILSVRC and COCO 2015.
The impact
ResNet made depth a reliable scaling axis — add layers and performance improves instead of collapses. Nearly every modern vision backbone (and many non-vision models) uses residual connections: DenseNet, EfficientNet, Vision Transformers, and even the 's own encoder and decoder blocks. The skip connection is now as fundamental to as the convolution itself.
Imagine climbing a skyscraper by staircase. At 20 floors you're fine. At 56 floors the stairs get so twisted that you actually end up lower than if you'd stopped at 20 — adding more stairs made you lose altitude.
The fix: an elevator shaft beside the stairs. At every floor, the elevator copies you up for free (the identity), and the staircase only needs to adjust the small difference from one floor to the next. Now 152 floors is easy — each floor can always fall back on the elevator if the staircase has nothing useful to add.
That elevator is the skip connection. The staircase is the residual F(x). Together they give you H(x) = F(x) + x.
The problem: deeper is not always better
By 2015, the deep learning community had a clear recipe: add more layers, get better results. VGG-16 and VGG-19 beat AlexNet's 8 layers. GoogLeNet went to 22. The obvious next step was 50, 100, 150 layers.
But something strange happened. A 56-layer plain CNN trained on CIFAR-10 had higher training error than its 20-layer counterpart. Not just higher test error — higher training error. This meant the problem was not overfitting; the deeper network simply couldn't learn as effectively.
This is the degradation problem: beyond a certain depth, adding layers makes the optimization landscape so difficult that the network can't even learn to copy the shallower network's solution (which it should be able to, since extra layers could in theory just perform identity mappings).
The idea: learn the residual, not the full mapping
He et al. made a simple but profound observation: if the optimal function for a block is close to the identity (i.e. the block should mostly pass data through), then learning H(x) = x from scratch is hard for a nonlinear stack of layers. But learning F(x) = 0 is easy — just push all weights toward zero.
The reformulates the task. Instead of learning H(x) directly, the layers learn the residual F(x) = H(x) − x, and a adds x back:
output = F(x) + x
If the block has nothing to contribute, F(x) = 0 and data passes through unchanged. If the block should transform the data, it learns only the difference from identity. Either way, optimization is easier.
Why skip connections fix gradient flow
The gradient of the loss with respect to an early layer x_l is:
That "1 +" is the identity path. Even if the learned residual gradients are small, the gradient can always flow backward through the skip connections undiminished. In plain networks, gradients must traverse every layer's weights — each multiplication can shrink (vanishing) or grow (exploding) the signal. The skip connection gives gradients a highway that runs straight back to any layer.
Bottleneck design: going really deep
For networks of 50+ layers, the basic two-layer residual block (two 3×3 convolutions) becomes computationally expensive. ResNet introduces the : a 1×1 convolution that reduces dimensions (e.g. 256 → 64), a 3×3 convolution that does the heavy spatial work on the smaller representation, and a 1×1 convolution that expands back (64 → 256).
The 3×3 conv now operates on 64 channels instead of 256, cutting compute by roughly 4×. This makes 101- and 152-layer ResNets practical with fewer FLOPs than VGG-16.
The family: from ResNet-18 to ResNet-152
ResNet is not one model but a family of architectures sharing the same residual design. The shallow variants — ResNet-18 (11.7M parameters) and ResNet-34 (21.8M) — use basic two-layer residual blocks. The deeper variants — ResNet-50 (25.6M), ResNet-101 (44.5M), and ResNet-152 (60.2M) — use bottleneck blocks. The interactive chart below shows their ImageNet top-5 error rates: from 10.92% (ResNet-18) down to 5.71% (ResNet-152).
Every variant follows the same pattern: an initial 7×7 conv + max pool, then four stages of residual blocks with doubling channels (64 → 128 → 256 → 512), and a global average pool + at the end.
Batch normalization: the silent partner
Every convolutional layer in ResNet is followed by (BN), which normalizes each channel's activations to zero mean and unit variance within each mini-batch. BN serves three roles:
- Stabilizes training by preventing internal covariate shift — each layer receives inputs with consistent statistics.
- Enables higher learning rates because normalized activations stay within the active range of ReLU, preventing both saturation and dead neurons.
- Acts as mild regularization because the per-batch statistics add noise, reducing overfitting.
Without BN, training 152-layer networks would be impractical even with skip connections. ResNet pairs residual learning (fixing the optimization problem) with batch normalization (stabilizing the activations) — both are necessary.
Simplified to show the idea — not the real implementation.
import numpy as np
def batch_norm(x, eps=1e-5):
"""Normalize each channel to zero mean, unit variance."""
mean = x.mean(axis=(0, 2, 3), keepdims=True)
var = x.var(axis=(0, 2, 3), keepdims=True)
return (x - mean) / np.sqrt(var + eps)
def conv3x3(x, W):
"""Placeholder for 3×3 convolution."""
# Real implementation slides W across spatial dims
return x # simplified
def residual_block(x, W1, W2):
"""Core of ResNet: F(x) + x."""
identity = x # save the input
out = conv3x3(x, W1) # first 3×3 conv
out = batch_norm(out) # normalize
out = np.maximum(0, out) # ReLU
out = conv3x3(out, W2) # second 3×3 conv
out = batch_norm(out) # normalize
out = out + identity # THE SKIP CONNECTION
out = np.maximum(0, out) # ReLU after addition
return out
# That's it. `out + identity` is the entire innovation. # Without it: 56 layers > 20 layers in error. # With it: 152 layers sets world records.Why it mattered
2015
ResNet — skip connections conquer depth
He et al. trained 152-layer networks with lower error and compute than VGG-16/19 by adding identity shortcuts. Won every track at ILSVRC and COCO 2015.
2016
Identity Mappings in Deep Residual Networks
He & Sun showed that moving BN and ReLU before the convolution ("pre-activation") improves gradient flow further, enabling 1001-layer ResNets on CIFAR.
2016
ResNeXt — wider residual paths
Xie et al. replaced single bottleneck paths with parallel grouped convolutions ("cardinality"), showing that width within residual blocks is as important as depth.
2017
DenseNet — every layer connects to every layer
Huang et al. extended skip connections so each layer receives feature maps from all preceding layers, maximizing feature reuse and gradient flow.
2017
Mask R-CNN — ResNet backbone for instance segmentation
He et al. used ResNet as the backbone for object detection and pixel-level segmentation, setting the template for modern detection architectures.
2019
EfficientNet — scaling depth, width, and resolution together
Tan & Le showed that scaling depth (à la ResNet), width, and input resolution in a balanced ratio outperforms scaling any one dimension alone.
2020
Vision Transformer — residual connections in every block
ViT replaced convolutions with self-attention for images but kept ResNet's residual connections in every encoder block — the skip connection transcended its origin architecture.
The skip connection is now everywhere. Every Transformer block — in GPT, Claude, Gemini, BERT — is a residual block. The idea that you should always give gradients an unobstructed path backward is now taken as an axiom of neural architecture design.
CitationHe, Zhang, Ren, Sun. Deep Residual Learning for Image Recognition. CVPR, 2016.
Terms in this paper
- Skip Connectionالاتصال التجاوزي
- Residual Blockالكتلة المتبقّية
- Identity Mappingتحويل الهوية
- Degradation Problemمشكلة التدهور
- Bottleneck Blockكتلة عنق الزجاجة
- Residual Learningالتعلم المتبقّي
- Global Average Poolingالتجميع المتوسط العام