Computer Vision2022intermediate13 min read
A ConvNet for the 2020s
شبكة التفافية لعصر العشرينيات
Liu, Z. · Mao, H. · Wu, C.-Y. · Feichtenhofer, C. · Darrell, T. · Xie, S. — CVPR
The problem
By 2021, Vision Transformers (ViTs) had overtaken convolutional neural networks as the state-of-the-art in image recognition. Swin introduced hierarchical structure and windowed , dominating classification, detection, and segmentation benchmarks. The narrative was clear: convolutions are the past, attention is the future. But was the performance gap due to the inherent superiority of attention — or simply because ConvNets had not adopted the modern recipes and design principles that Transformers benefited from?
The contribution
ConvNeXt: a family of pure convolutional models built by systematically modernizing a standard ResNet-50 toward the design of Swin Transformer. Through a roadmap of incremental changes — modern training recipes, macro design adjustments (stage ratios, patchified stem), depthwise separable convolutions, inverted bottlenecks, large 7×7 kernels, , , and fewer activation/ layers — the authors show that each step closes the gap. The final ConvNeXt-T achieves 82.1% top-1 accuracy, surpassing Swin-T (81.3%) with similar . Scaled-up variants reach 87.8% and outperform Swin Transformers on COCO detection and ADE20K segmentation while maintaining higher .
The impact
ConvNeXt proved that the perceived gap between ConvNets and Transformers was largely a gap in training and design modernization, not in the fundamental mechanism. It revived confidence in convolutional architectures and became a widely-used in downstream vision tasks. Its design principles influenced ConvNeXt V2 and models like Depth Anything that chose ConvNeXt encoders for their efficiency and simplicity. The paper's methodology — systematic, controlled ablation from one architecture toward another — became a template for design-space exploration.
Picture a heritage building in a modern city. Developers want to tear it down and build a glass tower (the Transformer). But an architect asks: what if we keep the original structure and renovate room by room? New wiring, modern insulation, bigger windows — same footprint.
That's ConvNeXt. The "building" is the ResNet convolutional network. The "renovation" is a sequence of surgical upgrades borrowed from Transformer design. By the end, the heritage building matches the glass tower in performance — proving the foundations were never the problem.
The starting line: why did ConvNets fall behind?
By 2020, ResNets had dominated computer vision for five years. They were simple, well-understood, and deployed everywhere — from medical imaging to self-driving cars. Then Vision Transformers arrived, splitting images into patches and processing them with . Suddenly, the Transformer — originally designed for text — was beating CNNs on image benchmarks.
Swin Transformer made things worse for the camp: by adding hierarchical stages and windowed attention, it matched the multi-scale design that CNNs naturally have, and outperformed them on detection and segmentation too.
But the comparison was unfair. Swin Transformer used modern training techniques — 300- schedules, optimizer, aggressive , — while ResNets were still trained with the decade-old 90-epoch recipe. Was the Transformer winning because of attention, or because of training?
ConvNeXt's authors asked: what if we give the old ConvNet every advantage the Transformer has, one upgrade at a time, and measure the effect of each?
The modernization roadmap: from ResNet to ConvNeXt
The genius of this paper is its method: instead of proposing one big new idea, it takes ResNet-50 and applies one change at a time, measuring the accuracy after each step. This gives us a clear understanding of which design choices actually matter. The roadmap has two phases: first, modernize the training recipe (no architectural changes); second, modify the architecture step by step.
Step 0: modern training recipe
Before touching the architecture, the authors upgraded ResNet-50's training recipe to match what Vision Transformers use. This includes extending training from 90 to 300 epochs, switching from SGD to AdamW optimizer, applying data augmentation techniques like Mixup, Cutmix, and RandAugment, and adding regularization through stochastic depth and .
The result: ResNet-50 jumps from 76.1% to 78.8% on ImageNet — a 2.7% improvement with zero architectural changes. This single step shows how much performance was left on the table simply because CNNs were trained with outdated recipes. The gap between ResNet and Swin-T (81.3%) shrinks from 5.2% to just 2.5%.
Macro design: stage ratios and patchified stem
The first architectural changes affect the big picture — how the network is shaped at the highest level.
. ResNet-50 has four stages with (3, 4, 6, 3) blocks — roughly a 1:1:2:1 ratio. Swin-T uses (1, 1, 3, 1). Adjusting ResNet's stage ratio to (3, 3, 9, 3) — which matches Swin's 1:1:3:1 pattern — improves accuracy from 78.8% to 79.4%. The key insight is that the third stage, which operates at the medium resolution, should do the bulk of the computation. Think of it as giving the network more room to think at the scale where objects are most recognizable.
Patchified stem. ResNet's stem is a 7×7 with 2 followed by — a gradual downsampling. Vision Transformers instead use aggressive non-overlapping patches (): a single 4×4 convolution with stride 4 that splits the image into a grid in one step. Adopting this patchify strategy slightly improves accuracy to 79.5% and aligns the network's initial processing with how Transformers see the image.
ResNeXt-ify: depthwise separable convolutions
Standard convolutions mix spatial information (where patterns are) and information (what patterns are) in a single operation. This is powerful but expensive. splits this into two cheaper operations: a that processes each channel independently (spatial mixing), followed by a 1×1 convolution that mixes information across channels (channel mixing).
This separation mirrors what happens inside a Transformer: the self-attention layer mixes information across spatial positions (like depthwise conv), while the mixes across channels (like 1×1 conv). By adopting depthwise convolutions and widening the channels from 64 to 96, ConvNeXt reaches 80.5% accuracy.
Inverted bottleneck: expand, then compress
In a standard ResNet , the channels go wide → narrow → wide: input at 256 channels, squeeze to 64, then expand back to 256. This saves compute by doing the heavy 3×3 convolution in the narrow middle.
Transformers do the opposite. In the MLP block, the hidden dimension is 4× wider than the input: input at dimension d, expand to 4d, then compress back to d. The reasoning is that a wider hidden layer gives the network more room to learn complex transformations before compressing back.
ConvNeXt adopts this : input at 96 channels, expand to 384 (4×) via a 1×1 conv, then compress back to 96. This slightly improves accuracy to 80.6%. More importantly, combined with moving the depthwise convolution layer to operate on the narrow input (before expansion), it reduces FLOPs while maintaining representational power.
Large kernels: seeing the bigger picture
Vision Transformers owe much of their power to self-attention, which gives every token a global view of the image. Swin Transformer uses a window size of at least 7×7, which is the minimum attention span per block. Meanwhile, standard ResNets use 3×3 kernels — a much narrower view of the local neighborhood.
To bridge this gap, ConvNeXt increases the depthwise convolution from 3×3 to 7×7. This gives each spatial position a wider in a single layer — similar in spirit to the attention window. The key design choice is to place the large depthwise convolution early in the block (before the channel expansion), so the expensive computation happens on fewer channels.
The accuracy stays at 80.6% despite the larger kernel — but the FLOPs decrease because moving the depthwise layer up means it operates on fewer channels. The real benefit appears later when all changes are combined.
Micro design: activation, normalization, and less is more
The final refinements are small but collectively powerful. They change the low-level building blocks of each .
GELU instead of . Transformers use GELU (Gaussian Error Linear Unit) rather than ReLU as their . GELU is smoother — it doesn't hard-clip negative values to zero but instead applies a smooth gating based on the input's magnitude. ConvNeXt adopts GELU and, crucially, uses only one activation function in the entire block (between the two 1×1 conv layers), mirroring the Transformer MLP which has just one nonlinearity.
LayerNorm instead of BatchNorm. computes statistics across the batch dimension, which couples samples together and causes issues with small batch sizes. Layer normalization computes per-sample statistics, making it batch-size independent. Transformers use LayerNorm everywhere. ConvNeXt replaces all BatchNorm layers with a single LayerNorm placed before the depthwise convolution.
Fewer normalization and activation layers. ResNet blocks pile up alternating conv-BN-ReLU triplets. ConvNeXt strips this down to one LayerNorm and one GELU per block. This "less is more" philosophy follows the Transformer pattern and actually improves accuracy.
Separate downsampling layers. Instead of using strided convolutions within residual blocks for downsampling between stages (as ResNet does), ConvNeXt inserts dedicated 2×2 stride-2 convolution layers between stages, similar to the in Swin Transformer.
These micro-design changes collectively push accuracy from 80.6% to 82.0%, surpassing Swin-T's 81.3%.
The ConvNeXt block: putting it all together
The final ConvNeXt block is elegant in its simplicity:
- 7×7 depthwise conv — spatial mixing with a wide receptive field, operating on the narrow input channels.
- LayerNorm — a single normalization step.
- 1×1 conv (expand) — channel expansion to 4× the width.
- GELU — the only activation function in the block.
- 1×1 conv (compress) — channel compression back to the original width.
- — the input is added to the output via a , scaled by a learnable layer scale parameter.
This structure is strikingly similar to a Transformer block: the depthwise conv plays the role of self-attention (spatial mixing), the two 1×1 convs with GELU play the role of the MLP (channel mixing), and the skip connection provides the residual path.
Simplified to show the idea — not the real implementation.
import torch import torch.nn as nn
class ConvNeXtBlock(nn.Module):
"""A single ConvNeXt block."""
def __init__(self, dim, layer_scale_init=1e-6):
super().__init__()
# 1. Depthwise conv (spatial mixing)
self.dwconv = nn.Conv2d(dim, dim, kernel_size=7,
padding=3, groups=dim)
# 2. LayerNorm (channels-last format)
self.norm = nn.LayerNorm(dim)
# 3. Expand channels 4×
self.pwconv1 = nn.Linear(dim, 4 * dim)
# 4. GELU activation
self.act = nn.GELU()
# 5. Compress channels back
self.pwconv2 = nn.Linear(4 * dim, dim)
# 6. Layer scale (learnable)
self.gamma = nn.Parameter(
layer_scale_init * torch.ones(dim))
def forward(self, x):
shortcut = x # skip connection
x = self.dwconv(x) # 7×7 depthwise
x = x.permute(0, 2, 3, 1) # NCHW → NHWC
x = self.norm(x) # LayerNorm
x = self.pwconv1(x) # expand 4×
x = self.act(x) # GELU
x = self.pwconv2(x) # compress
x = self.gamma * x # layer scale
x = x.permute(0, 3, 1, 2) # NHWC → NCHW
return shortcut + x # residual addScaling up: from tiny to large
ConvNeXt defines four model sizes that parallel the Swin Transformer variants:
ConvNeXt-T (Tiny): 29M parameters, 4.5G FLOPs → 82.1% ImageNet top-1 (vs Swin-T 81.3%). ConvNeXt-S (Small): 50M parameters, 8.7G FLOPs → 83.1% (vs Swin-S 83.0%). ConvNeXt-B (Base): 89M parameters, 15.4G FLOPs → 83.8% (vs Swin-B 83.5%). ConvNeXt-L (Large): 198M parameters, 34.4G FLOPs → 84.3%.
When pre-trained on ImageNet-22K and fine-tuned at 384×384 resolution, ConvNeXt-B reaches 85.1% (vs Swin-B 84.5%) with 12.5% higher throughput. ConvNeXt-L reaches 85.5%, and ConvNeXt-XL reaches 87.8%.
On downstream tasks, ConvNeXt outperforms Swin Transformers as a backbone for Cascade Mask R-CNN on COCO (52.7 vs 51.9 box AP for the Base model) and UPerNet on ADE20K (52.6 vs 51.5 mIoU for the Base model).
The key takeaway: ConvNeXt achieves this with pure convolutions — no attention mechanism, no shifted windows, no relative position bias. Just standard convolutional building blocks, thoughtfully arranged.
Impact: what ConvNeXt changed
ConvNeXt's contribution extends beyond accuracy numbers. It demonstrated a rigorous methodology for comparing architectures: control every variable except the one you are testing. By modernizing ResNet one step at a time and measuring each change, the paper disentangled the contributions of training recipe, macro design, and micro design — something rarely done in the "propose a new architecture" tradition.
It also revived interest in convolutional architectures for practical deployment. ConvNets are simpler to implement, faster to compile on hardware (especially edge devices), and better supported by existing frameworks. Models like Depth Anything chose ConvNeXt as their encoder backbone precisely because of this practicality.
Perhaps most importantly, ConvNeXt reframed the narrative. The question stopped being "convolutions vs attention" and became "which design principles actually matter, regardless of the compute primitive?" Depthwise separable convolutions, inverted bottlenecks, fewer activations, LayerNorm — these ideas transcend the conv/attention divide.
Timeline: from ResNet to ConvNeXt and beyond
2015
ResNet — skip connections unlock depth
He et al. introduced residual connections that let gradients flow through identity paths, enabling networks with 150+ layers. The residual block (conv-BN-ReLU + skip) became the standard building block for CNNs.
2017
SE-Net — channel attention for CNNs
Squeeze-and-Excitation Networks introduced channel-wise attention: squeeze global spatial information into a channel descriptor, then excite channels by reweighting them. This showed CNNs can benefit from attention-like mechanisms.
2020
Vision Transformer (ViT) — attention replaces convolution
Dosovitskiy et al. applied a pure Transformer to image patches, achieving strong results with large-scale pre-training. Sparked the debate: are convolutions still necessary?
2021
Swin Transformer — hierarchical attention dominates
Liu et al. introduced shifted window attention with hierarchical stages, matching CNNs' multi-scale design. Swin set new records on classification, detection, and segmentation, becoming the de facto backbone standard.
2022
ConvNeXt — pure ConvNet fights back
Liu et al. systematically modernized ResNet to match Swin's design, proving the performance gap was about training and design choices, not the attention mechanism itself. ConvNeXt-T surpassed Swin-T with pure convolutions.
2023
ConvNeXt V2 — self-supervised scaling
Woo et al. co-designed ConvNeXt with masked autoencoders (FCMAE), adding Global Response Normalization. Showed ConvNets scale well with self-supervised pre-training too.
2024
Depth Anything — ConvNeXt as backbone
Yang et al. chose ConvNeXt as the encoder backbone for monocular depth estimation, demonstrating its practical value as an efficient, deployment-friendly feature extractor in downstream vision tasks.
ConvNeXt's lasting message is that the boundary between ConvNets and Transformers is thinner than it appears. The design principles that make Transformers work — wider layers, fewer activations, large receptive fields, modern normalization — can be transplanted into convolutional architectures with full effect. The compute primitive (convolution vs attention) matters less than the engineering recipe built around it.
CitationLiu, Mao, Wu, Feichtenhofer, Darrell, Xie. A ConvNet for the 2020s. CVPR, 2022.
Terms in this paper
- Depthwise Convolutionالالتفاف بالعمق
- Inverted Bottleneckعنق الزجاجة المقلوب
- Layer Normalizationالتسوية الطبقية
- GELUوحدة الخطأ الخطية الغاوسية (GELU)
- Stochastic Depthالعمق العشوائي
- Patchifyالتقطيع إلى رُقَع
- Stage Compute Ratioنسبة حوسبة المراحل
- 1x1 Convolutionالتفاف 1×1
- Group Normalizationالتسوية بالمجموعات
- Residual Blockالكتلة المتبقّية
- Backboneالبنية الأساسية
- Downstream Taskالمهمة اللاحقة
- Object Detectionرصد وتحديد الكائنات
- Semantic Segmentationالتجزئة الدلالية للصورة
- Feature Pyramid Networkشبكة الهرم الاستخلاصي للسمات