Computer Vision2018intermediate9 min read

Squeeze-and-Excitation Networks

شبكات الضغط والتنشيط الانتقائي

Hu, J. · Shen, L. · Sun, G. — CVPR

The problem

Standard convolutions entangle spatial and information in each layer: the filters learn local spatial patterns and combine channels, but they have no explicit mechanism to model which channels are most informative for a given input. Every channel's output is treated with equal weight, even when some channels carry far more discriminative information than others.

The contribution

The -and- (SE) block: a lightweight module that explicitly models channel interdependencies. It "squeezes" each channel to a single number via , then "excites" the channels through a small network (two fully-connected layers with a ) that outputs per-channel weights between 0 and 1. These weights rescale the original feature maps, amplifying informative channels and suppressing less useful ones. SE blocks can be plugged into any existing CNN at negligible extra cost, and the resulting SENet won first place in ILSVRC 2017 with a top-5 error of 2.251%.

The impact

SE blocks introduced the concept of to mainstream computer vision. The idea — letting the network learn to weight its own features — influenced EfficientNet's compound scaling, ConvNeXt's modernized convolutions, and the broader attention-in-vision movement. Channel attention is now a standard building block in detection, segmentation, and generation architectures, and the SE block remains the canonical reference point for this design pattern.

Imagine a TV production studio with dozens of camera operators, each filming a different aspect of a scene — close-ups of faces, wide establishing shots, overhead angles. In a standard broadcast, every camera feed gets mixed in equally. But a good director watches a quick summary from every camera, decides which angles matter for this scene, and turns up the good feeds while dimming the rest.

That's what a Squeeze-and-Excitation block does for a neural network's channels. Each channel is a "camera" capturing a different feature. The SE block watches a summary of every channel (squeeze), decides which ones matter (excitation), and adjusts the volume on each (recalibration).

The problem: all channels weigh the same

A convolutional layer applies a set of filters to an input, producing one per . Each feature map — each channel — captures a different pattern: one might detect horizontal edges, another diagonal textures, another color blobs.

But after the , every channel's contribution flows forward with equal weight. The network has no built-in way to say "for this image, the edge channel matters much more than the color channel." Channel interdependencies are only implicitly learned through stacked layers, never explicitly modeled.

This is like a committee where every member speaks at the same volume, regardless of expertise. The signal from the most relevant expert gets drowned by voices that happen to be loud but not informative.

Open in Lab
Without SE: all 4 channel outputs flow forward equally. With SE: the network learns to amplify the edge channel and suppress the noise channel for this input.
The demo wakes as you arrive…

The SE block: squeeze, excite, scale

The Squeeze-and-Excitation block adds three lightweight steps after any convolution:

Step 1 — Squeeze (Global Average Pooling). Each channel's H×W feature map is compressed to a single number: its spatial average. A of shape C×H×W becomes a of length C. This gives the network a global summary of what each channel detected across the entire spatial extent — like reading a one-line report from each camera operator.

Step 2 — Excitation (Bottleneck FC network). The C-length vector passes through two fully-connected layers with a bottleneck: the first reduces the dimension by a ratio rr (typically 16) and applies , the second restores it to C and applies . The output is C weights between 0 and 1 — an importance score for each channel. The bottleneck limits model complexity and forces the network to capture channel interactions in a compressed representation — like the director discussing with a small advisory panel rather than interviewing every camera operator.

Step 3 — Scale (Channel-wise multiplication). Each channel's original feature map is multiplied by its learned weight. Informative channels are amplified (weight near 1), irrelevant ones are suppressed (weight near 0). The spatial structure within each channel is preserved — only the channel's overall volume changes.

Open in Lab
Click each step to see how a C×H×W tensor is squeezed to a C-vector, excited into per-channel weights, and used to rescale the original feature maps.
The demo wakes as you arrive…

The math behind squeeze and excitation

Now that you understand the intuition — summarize, weigh, rescale — let's see it formally. The squeeze step compresses spatial information into a channel descriptor:

zc=Fsq(uc)=1H×W∑i=1H∑j=1Wuc(i,j)z_c = F_{sq}(u_c) = \frac{1}{H \times W} \sum_{i=1}^{H} \sum_{j=1}^{W} u_c(i,j)
Squeeze — global average pooling per channel — For each channel c, average all H×W spatial positions into a single scalar z_c. The result is a C-dimensional vector z — a global spatial summary.

Next, the excitation step learns non-linear channel interdependencies through a bottleneck:

s=Fex(z,W)=σ(W2⋅δ(W1⋅z))s = F_{ex}(z, W) = \sigma(W_2 \cdot \delta(W_1 \cdot z))
Excitation — gated bottleneck — W₁ ∈ ℝ^(C/r × C) reduces dimensionality · δ = ReLU adds non-linearity · W₂ ∈ ℝ^(C × C/r) restores dimensionality · σ = sigmoid squashes each weight to [0, 1]

Finally, the scale step applies these learned weights to the original feature maps:

x~c=Fscale(uc,sc)=sc⋅uc\tilde{x}_c = F_{scale}(u_c, s_c) = s_c \cdot u_c
Scale — channel-wise multiplication — Each channel's feature map u_c is multiplied by its scalar gate s_c. Channels deemed important (s_c ≈ 1) pass through nearly unchanged; irrelevant channels (s_c ≈ 0) are effectively muted.

The reduction ratio: balancing capacity and cost

The bottleneck in the excitation step has a rr — the reduction ratio — that controls how much the channel vector is compressed before being expanded back. Think of it as the size of the advisory panel: too large (small rr) and it becomes expensive and may overfit; too small (large rr) and the panel lacks the capacity to capture complex channel relationships.

The authors tested rr ∈ 32 and found that r=16r = 16 gives the best balance between accuracy and computational cost. At r=16r = 16, an SE block adds only about ~1% extra parameters to a ResNet-50 while providing a meaningful accuracy boost.

Open in Lab
Drag the slider to see how different reduction ratios change the bottleneck size and parameter count for a 256-channel layer.
The demo wakes as you arrive…

Plug-and-play: SE blocks in ResNet and Inception

A key strength of the SE block is that it is architecture-agnostic: it wraps around any transformation FtrF_{tr} (a convolution, a residual block, an Inception module) and recalibrates its output before the result continues downstream.

In a ResNet, the SE block sits inside each residual block: the convolution path produces feature maps, the SE block recalibrates them, and then the identity shortcut is added. This means the SE module only adjusts the non-identity branch, while the skip connection still provides its highway.

In GoogLeNet/Inception, the SE block wraps the entire Inception module. All the parallel branches (1×1, 3×3, 5×5, pool) are concatenated first, then the SE block recalibrates the concatenated channels.

In both cases, the integration requires zero changes to the original architecture's code structure — just one extra module inserted after the main transformation.

Open in Lab
Toggle between SE-ResNet and SE-Inception to see where the SE block is inserted.
The demo wakes as you arrive…

The SE block in code

Squeeze-and-Excitation block in PyTorchpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

class SEBlock(nn.Module):
    """Squeeze-and-Excitation block."""
    def __init__(self, channels, reduction=16):
        super().__init__()
        mid = channels // reduction
        self.squeeze = nn.AdaptiveAvgPool2d(1)       # C×H×W → C×1×1
        self.excitation = nn.Sequential(
            nn.Linear(channels, mid, bias=False),     # FC₁: C → C/r
            nn.ReLU(inplace=True),
            nn.Linear(mid, channels, bias=False),     # FC₂: C/r → C
            nn.Sigmoid(),                             # gate each channel to [0,1]
        )

    def forward(self, x):
        b, c, _, _ = x.shape
        z = self.squeeze(x).view(b, c)               # squeeze: spatial → scalar
        s = self.excitation(z).view(b, c, 1, 1)      # excite: learn channel weights
        return x * s                                  # scale: reweight channels

# Usage: insert after any conv block
# out = self.conv_block(x)
# out = self.se(out)          # ← one line is all it takes
# out = out + x              # then add the residual shortcut

Results: small cost, big gains

Adding SE blocks to existing architectures consistently improved performance on ImageNet with minimal overhead:

  • SE-ResNet-50 reduced top-1 error by ~1% over ResNet-50, adding only 2.5M parameters (~10% increase) and negligible extra compute.
  • SE-ResNeXt-50 beat ResNeXt-50 and even approached ResNeXt-101 performance — a model with twice the depth.
  • SE-Inception improved Inception-ResNet-v2 on both top-1 and top-5 metrics.
  • SENet-154, an SE-modified ResNeXt-based architecture, won ILSVRC 2017 with a top-5 error of just 2.251%, a ~25% relative improvement over the 2016 winner.

The consistent improvement across diverse architectures — from VGG to ResNet to Inception — confirmed that channel attention is a universally beneficial mechanism, not an architecture- specific trick.

Open in Lab
Comparison of top-1 error on ImageNet. Blue = original architecture, orange = with SE blocks. Lower is better.
The demo wakes as you arrive…

What do SE blocks actually learn?

The authors analyzed the learned channel weights across different layers and classes. In early layers (close to the input), the excitation patterns are nearly identical across classes — the network applies similar channel weighting regardless of what the image contains. This makes sense: early features like edges and textures are universally useful.

In deeper layers, the excitation becomes highly class-specific. For example, channels that detect fur texture get high weights for animal classes, while channels sensitive to geometric patterns get high weights for building classes. This shows the SE block learning a form of conditional computation: different inputs activate different channel weightings.

Interestingly, in the very last layers before classification, the excitation patterns converge again — suggesting the network has already selected its representation and needs less recalibration at the final stage.

Legacy: channel attention everywhere

  1. 2017

    SENet wins ILSVRC

    Hu, Shen & Sun win the 2017 ImageNet challenge with a 2.251% top-5 error. First time channel attention is shown to improve results across all major architectures.

  2. 2018

    CBAM: channel + spatial attention

    Woo et al. extend SE by adding a spatial attention branch, jointly learning which channels and which spatial locations matter.

  3. 2019

    EfficientNet adopts SE blocks

    Tan & Le's EfficientNet, found by neural architecture search, places SE blocks inside MBConv modules. Scaling with compound coefficients made it the best accuracy-per-FLOP model for years.

  4. 2020

    ECA-Net — efficient channel attention

    Wang et al. replace the SE bottleneck with a 1D convolution across channels, removing the FC layers entirely while matching SE's accuracy at lower cost.

  5. 2022

    ConvNeXt modernizes convolutions

    Liu et al. bring Transformer-era design lessons back to ConvNets. ConvNeXt includes channel-wise modulation inspired by SE — the "let each channel decide its own importance" principle lives on.

The SE block's influence extends beyond classification. Channel attention modules now appear in (YOLO variants), medical imaging (attention U-Nets), (RCAN), and generative models. The simplicity of the idea — a global pool, a small bottleneck, a channel-wise scale — set a template that later attention mechanisms refine but rarely abandon.

CitationHu, Shen, Sun. Squeeze-and-Excitation Networks. CVPR, 2018.

Terms in this paper