Computer Vision2018intermediate9 min read
Squeeze-and-Excitation Networks
شبكات الضغط والتنشيط الانتقائي
Hu, J. · Shen, L. · Sun, G. — CVPR
The problem
Standard convolutions entangle spatial and information in each layer: the filters learn local spatial patterns and combine channels, but they have no explicit mechanism to model which channels are most informative for a given input. Every channel's output is treated with equal weight, even when some channels carry far more discriminative information than others.
The contribution
The -and- (SE) block: a lightweight module that explicitly models channel interdependencies. It "squeezes" each channel to a single number via , then "excites" the channels through a small network (two fully-connected layers with a ) that outputs per-channel weights between 0 and 1. These weights rescale the original feature maps, amplifying informative channels and suppressing less useful ones. SE blocks can be plugged into any existing CNN at negligible extra cost, and the resulting SENet won first place in ILSVRC 2017 with a top-5 error of 2.251%.
The impact
SE blocks introduced the concept of to mainstream computer vision. The idea — letting the network learn to weight its own features — influenced EfficientNet's compound scaling, ConvNeXt's modernized convolutions, and the broader attention-in-vision movement. Channel attention is now a standard building block in detection, segmentation, and generation architectures, and the SE block remains the canonical reference point for this design pattern.
Imagine a TV production studio with dozens of camera operators, each filming a different aspect of a scene — close-ups of faces, wide establishing shots, overhead angles. In a standard broadcast, every camera feed gets mixed in equally. But a good director watches a quick summary from every camera, decides which angles matter for this scene, and turns up the good feeds while dimming the rest.
That's what a Squeeze-and-Excitation block does for a neural network's channels. Each channel is a "camera" capturing a different feature. The SE block watches a summary of every channel (squeeze), decides which ones matter (excitation), and adjusts the volume on each (recalibration).
The problem: all channels weigh the same
A convolutional layer applies a set of filters to an input, producing one per . Each feature map — each channel — captures a different pattern: one might detect horizontal edges, another diagonal textures, another color blobs.
But after the , every channel's contribution flows forward with equal weight. The network has no built-in way to say "for this image, the edge channel matters much more than the color channel." Channel interdependencies are only implicitly learned through stacked layers, never explicitly modeled.
This is like a committee where every member speaks at the same volume, regardless of expertise. The signal from the most relevant expert gets drowned by voices that happen to be loud but not informative.
The SE block: squeeze, excite, scale
The Squeeze-and-Excitation block adds three lightweight steps after any convolution:
Step 1 — Squeeze (Global Average Pooling). Each channel's H×W feature map is compressed to a single number: its spatial average. A of shape C×H×W becomes a of length C. This gives the network a global summary of what each channel detected across the entire spatial extent — like reading a one-line report from each camera operator.
Step 2 — Excitation (Bottleneck FC network). The C-length vector passes through two fully-connected layers with a bottleneck: the first reduces the dimension by a ratio (typically 16) and applies , the second restores it to C and applies . The output is C weights between 0 and 1 — an importance score for each channel. The bottleneck limits model complexity and forces the network to capture channel interactions in a compressed representation — like the director discussing with a small advisory panel rather than interviewing every camera operator.
Step 3 — Scale (Channel-wise multiplication). Each channel's original feature map is multiplied by its learned weight. Informative channels are amplified (weight near 1), irrelevant ones are suppressed (weight near 0). The spatial structure within each channel is preserved — only the channel's overall volume changes.
The math behind squeeze and excitation
Now that you understand the intuition — summarize, weigh, rescale — let's see it formally. The squeeze step compresses spatial information into a channel descriptor:
Next, the excitation step learns non-linear channel interdependencies through a bottleneck:
Finally, the scale step applies these learned weights to the original feature maps:
The reduction ratio: balancing capacity and cost
The bottleneck in the excitation step has a — the reduction ratio — that controls how much the channel vector is compressed before being expanded back. Think of it as the size of the advisory panel: too large (small ) and it becomes expensive and may overfit; too small (large ) and the panel lacks the capacity to capture complex channel relationships.
The authors tested ∈ 32 and found that gives the best balance between accuracy and computational cost. At , an SE block adds only about ~1% extra parameters to a ResNet-50 while providing a meaningful accuracy boost.
Plug-and-play: SE blocks in ResNet and Inception
A key strength of the SE block is that it is architecture-agnostic: it wraps around any transformation (a convolution, a residual block, an Inception module) and recalibrates its output before the result continues downstream.
In a ResNet, the SE block sits inside each residual block: the convolution path produces feature maps, the SE block recalibrates them, and then the identity shortcut is added. This means the SE module only adjusts the non-identity branch, while the skip connection still provides its highway.
In GoogLeNet/Inception, the SE block wraps the entire Inception module. All the parallel branches (1×1, 3×3, 5×5, pool) are concatenated first, then the SE block recalibrates the concatenated channels.
In both cases, the integration requires zero changes to the original architecture's code structure — just one extra module inserted after the main transformation.
The SE block in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
class SEBlock(nn.Module):
"""Squeeze-and-Excitation block."""
def __init__(self, channels, reduction=16):
super().__init__()
mid = channels // reduction
self.squeeze = nn.AdaptiveAvgPool2d(1) # C×H×W → C×1×1
self.excitation = nn.Sequential(
nn.Linear(channels, mid, bias=False), # FC₁: C → C/r
nn.ReLU(inplace=True),
nn.Linear(mid, channels, bias=False), # FC₂: C/r → C
nn.Sigmoid(), # gate each channel to [0,1]
)
def forward(self, x):
b, c, _, _ = x.shape
z = self.squeeze(x).view(b, c) # squeeze: spatial → scalar
s = self.excitation(z).view(b, c, 1, 1) # excite: learn channel weights
return x * s # scale: reweight channels
# Usage: insert after any conv block
# out = self.conv_block(x)
# out = self.se(out) # ← one line is all it takes
# out = out + x # then add the residual shortcutResults: small cost, big gains
Adding SE blocks to existing architectures consistently improved performance on ImageNet with minimal overhead:
- SE-ResNet-50 reduced top-1 error by ~1% over ResNet-50, adding only 2.5M parameters (~10% increase) and negligible extra compute.
- SE-ResNeXt-50 beat ResNeXt-50 and even approached ResNeXt-101 performance — a model with twice the depth.
- SE-Inception improved Inception-ResNet-v2 on both top-1 and top-5 metrics.
- SENet-154, an SE-modified ResNeXt-based architecture, won ILSVRC 2017 with a top-5 error of just 2.251%, a ~25% relative improvement over the 2016 winner.
The consistent improvement across diverse architectures — from VGG to ResNet to Inception — confirmed that channel attention is a universally beneficial mechanism, not an architecture- specific trick.
What do SE blocks actually learn?
The authors analyzed the learned channel weights across different layers and classes. In early layers (close to the input), the excitation patterns are nearly identical across classes — the network applies similar channel weighting regardless of what the image contains. This makes sense: early features like edges and textures are universally useful.
In deeper layers, the excitation becomes highly class-specific. For example, channels that detect fur texture get high weights for animal classes, while channels sensitive to geometric patterns get high weights for building classes. This shows the SE block learning a form of conditional computation: different inputs activate different channel weightings.
Interestingly, in the very last layers before classification, the excitation patterns converge again — suggesting the network has already selected its representation and needs less recalibration at the final stage.
Legacy: channel attention everywhere
2017
SENet wins ILSVRC
Hu, Shen & Sun win the 2017 ImageNet challenge with a 2.251% top-5 error. First time channel attention is shown to improve results across all major architectures.
2018
CBAM: channel + spatial attention
Woo et al. extend SE by adding a spatial attention branch, jointly learning which channels and which spatial locations matter.
2019
EfficientNet adopts SE blocks
Tan & Le's EfficientNet, found by neural architecture search, places SE blocks inside MBConv modules. Scaling with compound coefficients made it the best accuracy-per-FLOP model for years.
2020
ECA-Net — efficient channel attention
Wang et al. replace the SE bottleneck with a 1D convolution across channels, removing the FC layers entirely while matching SE's accuracy at lower cost.
2022
ConvNeXt modernizes convolutions
Liu et al. bring Transformer-era design lessons back to ConvNets. ConvNeXt includes channel-wise modulation inspired by SE — the "let each channel decide its own importance" principle lives on.
The SE block's influence extends beyond classification. Channel attention modules now appear in (YOLO variants), medical imaging (attention U-Nets), (RCAN), and generative models. The simplicity of the idea — a global pool, a small bottleneck, a channel-wise scale — set a template that later attention mechanisms refine but rarely abandon.
CitationHu, Shen, Sun. Squeeze-and-Excitation Networks. CVPR, 2018.
Terms in this paper
- Channel Attentionانتباه القنوات
- Global Average Poolingالتجميع المتوسط العام
- Feature Recalibrationإعادة معايرة السمات
- Squeezeالضغط
- Excitationالتنبيه
- Reduction Ratioنسبة الاختزال
- Gating Mechanismآلية البوابات
- Residual Connectionالوصلة التجاوزية