Computer Vision2021intermediate13 min read

Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows

محوِّل Swin: محوِّل رؤية هرمي باستخدام النوافذ المُنزاحة

Liu, Z. · Lin, Y. · Cao, Y. · Hu, H. · Wei, Y. · Zhang, Z. · Lin, S. · Guo, B. — ICCV

The problem

The original treats an image as a flat sequence of fixed-size patches and applies global across all of them. This creates two problems. First, the computational cost is quadratic in the number of patches — a 1024×1024 image with 16×16 patches produces 4,096 tokens, and global attention over them costs about 16 million pair-wise comparisons per layer. This makes ViT impractical for high- tasks like and . Second, ViT produces a single-scale with no built-in hierarchy, so it cannot easily replace backbones in frameworks like Feature Pyramid Networks that rely on multi-scale features.

The contribution

Swin introduces two key innovations. First, window-based self-attention that restricts computation to non-overlapping local windows of M×M patches, reducing complexity from quadratic to linear in the image size. Second, a shifted window scheme that alternates between regular and shifted partitions across consecutive Transformer blocks, enabling cross-window connections without extra cost. Combined with hierarchical merging that doubles the dimension while halving spatial resolution at each stage, Swin produces multi-scale feature maps compatible with FPN-style dense prediction frameworks. The result is 87.3% top-1 on -1K, 58.7 box AP on COCO test-dev, and 53.5 mIoU on ADE20K — surpassing prior state-of-the-art by significant margins.

The impact

Swin Transformer proved that Transformers can serve as general-purpose vision backbones competitive with — and often superior to — CNNs across all major vision tasks. It won the ICCV 2021 Best Paper Award (Marr Prize) and became one of the most cited computer vision papers of the decade. It directly inspired ConvNeXt (which showed CNNs could match Swin by adopting its design principles), SegFormer, and many other architectures. Swin established shifted windows and hierarchical patch merging as standard tools in the vision Transformer toolkit.

Imagine trying to understand a giant mural. The original Vision Transformer (ViT) approach is like hovering in a helicopter and trying to compare every brushstroke with every other brushstroke at the same time — exhausting, expensive, and you lose the fine details.

Swin Transformer takes a smarter approach. You stand in front of the mural with a movable frame that shows you one section at a time. You study the details within that frame, then shift the frame slightly so it overlaps with the previous view, connecting what you learned across sections. Every few steps, you take a step back to see a larger portion at lower resolution — building up from fine-grained textures to the overall composition, like constructing a pyramid of understanding.

The problem: global attention does not scale to pixels

The original Vision Transformer (ViT) splits an image into 16×16 patches and treats each patch as a . Self-attention then lets every token attend to every other token — exactly like words in a sentence. This works beautifully for image where the image is typically 224×224 pixels (producing only 196 tokens).

But vision is much more than classification. Object detection needs to find small objects in large images. Semantic segmentation needs to label every single pixel. These dense prediction tasks require high-resolution inputs — 800×1200 or more. At that resolution with 16×16 patches, you get 3,750 tokens. Global self-attention over 3,750 tokens means over 14 million pair-wise comparisons per layer per head. The cost grows quadratically: double the resolution, quadruple the computation.

There is a second problem. CNNs naturally produce hierarchical feature maps — early layers capture fine-grained details at high resolution, while deeper layers capture abstract semantics at lower resolution. This multi-scale pyramid is essential for frameworks like Feature Pyramid Networks (FPN) used in detection and segmentation. ViT, by contrast, produces a single-scale feature map — all tokens remain at the same resolution from start to finish. This makes it architecturally incompatible with the dense prediction pipelines that power modern computer vision.

Open in Lab
Compare ViT's quadratic cost with Swin's linear cost as image resolution increases. Drag the resolution slider to see how the gap explodes.
The demo wakes as you arrive…

The solution: a hierarchical pyramid with local windows

Swin Transformer borrows the best idea from CNNs — the multi-scale feature pyramid — and implements it using Transformers. The architecture has four stages:

Stage 1: The image is split into non-overlapping 4×4 patches. Each patch is flattened into a 48-dimensional vector (4 × 4 × 3 channels) and linearly projected to dimension CC. For a 224×224 image, this produces 56×56 = 3,136 tokens at the highest resolution.

Stage 2: A patch merging layer groups each 2×2 neighborhood of tokens, concatenates their features (producing a 4CC-dimensional vector), and projects them down to 2CC dimensions. The spatial resolution halves to 28×28, while the channel depth doubles. This is analogous to a strided convolution or .

Stages 3 and 4: Patch merging repeats, producing 14×14 tokens with 4CC channels, then 7×7 tokens with 8CC channels. Each stage applies multiple Swin Transformer blocks before merging.

The result is four feature maps at resolutions H4×W4\frac{H}{4} \times \frac{W}{4}, H8×W8\frac{H}{8} \times \frac{W}{8}, H16×W16\frac{H}{16} \times \frac{W}{16}, and H32×W32\frac{H}{32} \times \frac{W}{32} — exactly what FPN-style frameworks expect. Swin can therefore serve as a drop-in replacement for ResNet and other CNN backbones in existing detection and segmentation pipelines.

Open in Lab
Explore Swin's four-stage pyramid. Click each stage to see how resolution shrinks and channels grow — producing the multi-scale features that dense prediction needs.
The demo wakes as you arrive…

Window attention: seeing clearly by looking locally

The core efficiency trick in Swin is simple: instead of computing self-attention across the entire feature map, partition it into non-overlapping windows of M×MM \times M patches (typically M=7M = 7) and compute attention independently within each window.

In global self-attention, the computational cost for a feature map of h×wh \times w patches with CC channels is:

Ω(MSA)=4hwC2+2(hw)2C\Omega(\text{MSA}) = 4hwC^2 + 2(hw)^2C
Global self-attention complexity — quadratic in spatial size — The (hw)2(hw)^2 term is the killer: it means that doubling the image resolution (doubling hh and ww) increases the attention cost by 16×. This makes global attention impractical for high-resolution dense prediction tasks.

With window-based attention (W-MSA), attention is computed within each M×MM \times M window independently. The cost becomes:

Ω(W-MSA)=4hwC2+2M2hwC\Omega(\text{W-MSA}) = 4hwC^2 + 2M^2hwC
Window self-attention complexity — linear in spatial size — The critical difference: hwhw appears linearly (not squared). With fixed window size M=7M = 7, the 2M2hwC=98⋅hwC2M^2hwC = 98 \cdot hwC term grows linearly with image size. Double the resolution, double (not quadruple) the cost. This makes Swin practical for arbitrarily large images.

Within each window, the attention mechanism is the standard with one important addition — a learnable BB:

Attention(Q,K,V)=SoftMax ⁣(QKTd+B)V\text{Attention}(Q, K, V) = \text{SoftMax}\!\left(\frac{QK^T}{\sqrt{d}} + B\right)V
Window attention with relative position bias — QQ, KK, VV are query, key, and value matrices projected from M2M^2 patches within the window. dd is the head dimension. BB is sampled from a learnable bias table B^∈R(2M−1)×(2M−1)\hat{B} \in \mathbb{R}^{(2M-1) \times (2M-1)} indexed by the relative position of each patch pair. This bias tells the model "patches that are two steps apart behave differently from adjacent patches" — encoding spatial awareness without absolute position embeddings.

The key innovation: shifted windows for cross-window connections

Window-based attention has a limitation: tokens in one window cannot see tokens in neighboring windows. The model is blind at window boundaries — like reading a book through a slot that cuts off every sentence at the same position.

Swin solves this with a beautifully simple idea: alternate between two window configurations. In layer ℓ\ell, use the regular window partition. In layer ℓ+1\ell+1, shift the partition by (⌊M/2⌋,⌊M/2⌋)(\lfloor M/2 \rfloor, \lfloor M/2 \rfloor) pixels. The shifted windows straddle the boundaries of the previous layer's windows, creating cross-window connections.

Think of it like laying tiles: in one layer you lay a 7×7 grid. In the next layer, you offset the grid by 3.5 tiles. Every tile in the shifted grid overlaps with parts of four tiles from the previous grid. Information flows across boundaries without any explicit cross- — just by alternating the grid.

The two consecutive Swin Transformer blocks compute:

z^ℓ=W-MSA(LN(zℓ−1))+zℓ−1,zℓ=MLP(LN(z^ℓ))+z^ℓ\hat{z}^{\ell} = \text{W-MSA}(\text{LN}(z^{\ell-1})) + z^{\ell-1}, \quad z^{\ell} = \text{MLP}(\text{LN}(\hat{z}^{\ell})) + \hat{z}^{\ell}
Regular window attention block (layer ℓ) — W-MSA is window-based multi-head self-attention. LN is layer normalization applied before each sub-layer. The residual connections (+z+ z) ensure gradient flow. The MLP has two linear layers with GELU activation.
z^ℓ+1=SW-MSA(LN(zℓ))+zℓ,zℓ+1=MLP(LN(z^ℓ+1))+z^ℓ+1\hat{z}^{\ell+1} = \text{SW-MSA}(\text{LN}(z^{\ell})) + z^{\ell}, \quad z^{\ell+1} = \text{MLP}(\text{LN}(\hat{z}^{\ell+1})) + \hat{z}^{\ell+1}
Shifted window attention block (layer ℓ+1) — SW-MSA is identical to W-MSA except the window partition is shifted by (⌊M/2⌋,⌊M/2⌋)(\lfloor M/2 \rfloor, \lfloor M/2 \rfloor). This pair of blocks — regular then shifted — forms one logical unit. Consecutive pairs build up cross-window communication.
Open in Lab
Watch how shifting the window partition creates cross-boundary connections. Toggle between regular and shifted grids to see which patches can communicate.
The demo wakes as you arrive…

Efficient shifted windows: the cyclic shift trick

A naive implementation of shifted windows creates a problem: the shift produces windows of varying sizes at the borders. Some windows are smaller than M×MM \times M, which breaks the uniform batch computation that GPUs rely on for speed.

The solution is elegant: instead of creating irregular border windows, cyclically shift the entire feature map toward the top-left by (⌊M/2⌋,⌊M/2⌋)(\lfloor M/2 \rfloor, \lfloor M/2 \rfloor) positions. This wraps the bottom and right borders around to the top and left, creating a toroidal layout where all windows are again M×MM \times M. After attention, the feature map is shifted back to its original position.

But there is a catch: after the cyclic shift, some windows contain patches that are not spatially adjacent in the original image. For example, a window might contain patches from both the top-right and bottom-left corners. To prevent these non-adjacent patches from attending to each other, a masking mechanism is applied: the attention scores between non-adjacent sub-windows are set to −∞-\infty before , ensuring they contribute zero weight.

The result is that the number of windows stays exactly the same as in regular partitioning, enabling efficient batched computation with no wasted operations.

Open in Lab
Watch the cyclic shift in action: the feature map wraps around, attention is computed in uniform windows with masking, then everything shifts back.
The demo wakes as you arrive…

Inside the Swin block: LN → Attention → LN → MLP

Each Swin Transformer block follows the standard Transformer architecture with one key difference: the attention operates on local windows instead of the full sequence. The block has four components:

(LN): Applied before both the attention and sub-layers (pre-norm style). This stabilizes and allows higher learning rates.

Window Attention (W-MSA or SW-MSA): Multi-head self-attention computed within each window of M×MM \times M patches, with relative position bias added to each head.

: The input to each sub-layer is added to its output, ensuring flow through deep networks. This is the same used in ResNets and the original Transformer.

MLP: A two-layer feed-forward network with activation, expanding the hidden dimension by a ratio of 4× before projecting back. This is where the model adds non-linear capacity — the attention is, at its core, a weighted average, and the MLP adds the representational power to transform those averaged features.

Swin Transformer block — pseudocodepython

Simplified to show the idea — not the real implementation.

# Consecutive Swin Transformer blocks (one regular + one shifted)
# Block ℓ: regular window attention x_norm = LayerNorm(x) x = x + W_MSA(x_norm)          # window attention + residual x = x + MLP(LayerNorm(x))      # MLP + residual
# Block ℓ+1: shifted window attention x_norm = LayerNorm(x) x = x + SW_MSA(x_norm)         # shifted window attention + residual x = x + MLP(LayerNorm(x))      # MLP + residual

Patch merging: building the pyramid stage by stage

Between stages, patch merging reduces spatial resolution and increases channel depth — exactly like max pooling or strided convolution in a CNN. The process is straightforward:

Take a feature map of size H×W×CH \times W \times C. Group every 2×2 neighborhood of tokens. Concatenate the four tokens along the channel dimension, producing H2×W2×4C\frac{H}{2} \times \frac{W}{2} \times 4C. Then apply a linear layer that projects from 4C4C to 2C2C dimensions.

The result: spatial resolution halves, channel depth doubles, and total tokens reduce by 4×. This is how Swin transitions between its four stages, producing the multi-scale feature pyramid that dense prediction frameworks need.

Think of it as zooming out while switching to a higher-capacity lens. You see a larger area, but you now have richer features to describe it. Each stage operates at a different resolution, and the feature maps from all four stages are available for downstream tasks — just like FPN taps into multiple layers of a ResNet.

Open in Lab
Watch how patch merging groups 2×2 tokens, concatenates, and projects — shrinking spatial dimensions while enriching features.
The demo wakes as you arrive…

Model variants: from Tiny to Large

Swin Transformer comes in four sizes, scaling the base channel dimension CC and the number of blocks per stage:

Swin-T (Tiny): C=96C = 96, layers = [2, 2, 6, 2], ~29M parameters, 4.5 GFLOPs. Comparable in cost to ResNet-50.

Swin-S (Small): C=96C = 96, layers = [2, 2, 18, 2], ~50M parameters, 8.7 GFLOPs. Comparable to ResNet-101.

Swin-B (Base): C=128C = 128, layers = [2, 2, 18, 2], ~88M parameters, 15.4 GFLOPs. Comparable to ViT-B/DeiT-B.

Swin-L (Large): C=192C = 192, layers = [2, 2, 18, 2], ~197M parameters, 34.5 GFLOPs. The largest variant, pre-trained on ImageNet-22K.

Notice the pattern: Stage 3 has the most blocks (6 or 18) because it operates at 14×14 resolution — small enough to be computationally reasonable but large enough to capture rich spatial patterns. This design mirrors how ResNets concentrate their depth in the third stage (res4 block).

Open in Lab
Compare Swin-T, S, B, and L: parameters, FLOPs, and accuracy across classification, detection, and segmentation benchmarks.
The demo wakes as you arrive…

Results: Transformers conquer all vision tasks

Swin Transformer achieved state-of-the-art results across three major vision benchmarks:

Image Classification (ImageNet-1K): Swin-B achieved 83.5% top-1 accuracy without extra data. With ImageNet-22K pre-training, Swin-L reached 87.3% — surpassing all prior ViT and CNN models at comparable .

Object Detection (COCO): Using Cascade Mask R-CNN with Swin-L , the model achieved 58.7 box AP and 51.1 mask AP on COCO test-dev, surpassing the previous best by +2.7 box AP and +2.6 mask AP. This was the first time a pure Transformer backbone dominated COCO detection.

Semantic Segmentation (ADE20K): Swin-L achieved 53.5 mIoU, a +3.2 improvement over the previous state-of-the-art. The hierarchical multi-scale features proved essential for per-pixel prediction.

The key takeaway: Swin proved that a single Transformer architecture, with the right inductive biases (locality through windows, hierarchy through merging), can match and surpass the best CNNs on tasks they were specifically designed for.

Legacy: the design principles that outlived the model

  1. 2020

    Vision Transformer (ViT) — Google

    Showed that pure Transformers can match CNNs on image classification, but required massive pre-training data and couldn't handle dense prediction tasks.

  2. 2021

    Swin Transformer — Microsoft (this paper)

    Introduced shifted windows and hierarchical patch merging, making Transformers a general-purpose vision backbone. Won ICCV 2021 Best Paper (Marr Prize).

  3. 2021

    SegFormer — NVIDIA

    Adopted the hierarchical Transformer design for semantic segmentation, combining multi-scale features with a simple MLP decoder.

  4. 2022

    Swin Transformer V2 — Microsoft

    Extended Swin to 3 billion parameters with techniques like log-spaced continuous position bias and residual-post-norm, enabling training with higher-resolution images.

  5. 2022

    ConvNeXt — Meta AI

    Showed that pure CNNs can match Swin when modernized with its design principles: larger kernels (7×7), fewer activation functions, layer norm instead of batch norm. Proving Swin's design choices were more important than the attention mechanism itself.

Swin Transformer's deepest contribution is not the specific architecture but the design recipe it established: local attention with global communication through shifting, hierarchical feature maps through patch merging, and relative position bias instead of absolute embeddings. These principles now appear across the vision landscape — in video understanding (Video Swin), medical imaging, point cloud processing, and beyond. Even ConvNeXt, which deliberately avoided the Transformer attention mechanism, adopted Swin's overall blueprint and validated it as the right way to structure a vision backbone.

CitationLiu, Lin, Cao, Hu, Wei, Zhang, Lin, Guo. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. ICCV, 2021.

Terms in this paper