Computer Vision2020intermediate9 min read
An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale
صورة تساوي 16×16 كلمة: المحوِّلات للتعرف على الصور على نطاق واسع
Dosovitskiy, A. · Beyer, L. · Kolesnikov, A. · Weissenborn, D. · Zhai, X. · Unterthiner, T. · Dehghani, M. · Minderer, M. · Heigold, G. · Gelly, S. · Uszkoreit, J. · Houlsby, N. — ICLR
The problem
By 2020 CNNs dominated computer vision. The had revolutionized NLP, but attempts to apply to images either mixed it with convolutions or used it only on individual pixels — too expensive for real-world resolutions. The question was: can a pure Transformer, with no convolutions at all, match or beat the best CNNs on image ?
The contribution
: split the image into fixed-size 16×16 patches, flatten each patch into a , project it linearly into an , prepend a learnable [CLS] , add positional embeddings, and feed the sequence into a standard Transformer — exactly the one from "Attention Is All You Need." Pre-trained on large datasets (), ViT matched or exceeded the best CNNs (BiT, Noisy Student) on ImageNet, CIFAR-100, and VTAB while requiring substantially fewer compute resources to train.
The impact
ViT proved the Transformer is not language-specific — it is a general-purpose sequence processor. This unlocked a wave of vision Transformers (DeiT, Swin, BEiT, DINO, SAM), multimodal models (CLIP, DALL·E, Stable Diffusion), and ultimately the unified architectures behind today's multimodal AI. Two decades of -only vision ended with one paper.
A CNN reads an image like a person pressing a magnifying glass across a page: small area, high detail, one spot at a time. Useful features trickle upward through layers, and a top-left object can only influence a bottom-right decision after many stacked convolutions.
ViT replaces the magnifying glass with a conference table: the image is cut into 16×16 patches — like snapshots pinned to a board — and every patch sits at the table, free to address any other patch directly. The cat's ear can ask the grass in the background, "are you outdoor context?" in a single step.
The ceiling: CNNs see locally, not globally
CNNs process images with small filters (typically 3×3) that scan the image locally. To connect distant regions, you need many stacked layers — each layer's grows a little, but global context emerges only late in the network. This is a strong : locality and translation invariance are baked in.
These biases help when data is limited — they tell the how to look — but they also limit what the model can learn. In NLP, Transformers had already shown that with enough data, learned attention patterns outperform hand-coded structural assumptions. Could the same happen in vision?
The idea: treat image patches as tokens
ViT's insight is elegantly minimal: an image can be turned into a sequence the Transformer already knows how to process. The recipe:
- Split a 224×224 image into a grid of non-overlapping 16×16 patches → 196 patches.
- Flatten each patch (16 × 16 × 3 = 768 values) and project it linearly into a D-dimensional embedding.
- Prepend a learnable
[CLS]token — its final representation will carry the image's class. - Add learnable positional embeddings so the model knows which patch came from where.
- Feed the sequence of 197 tokens (196 patches + 1 CLS) into a standard Transformer encoder.
That's it. No convolutions, no pooling, no hand-designed feature extractors. The same architecture that reads sentences now reads images.
The architecture: a Transformer encoder, unchanged
ViT uses the standard Transformer encoder exactly as described in the original paper. Each layer applies:
- Multi-Head (MSA) — every patch attends to every other patch. A patch showing a wheel can attend to a patch showing a road, building "car on a road" context in one step.
- applied before each sub-layer (pre-norm), which stabilizes for deep models.
- block — two linear layers with GELU activation, applied independently to each patch position.
- Residual connections around both MSA and MLP.
After L layers, the [CLS] token's representation is passed through a small MLP head for classification. The model comes in three sizes: ViT-Base (86M params, 12 layers), ViT-Large (307M, 24 layers), and ViT-Huge (632M, 32 layers).
Why 16×16? The patch-size tradeoff
Patch size controls how many tokens the Transformer sees. For a 224×224 image:
- 32×32 patches → only 49 tokens — fast but coarse. Fine details are lost.
- 16×16 patches → 196 tokens — the sweet spot. Enough detail, tractable compute.
- 8×8 patches → 784 tokens — rich detail but self-attention costs O(n²), so 4× the patches means 16× the attention cost.
The naming convention encodes this: ViT-L/16 means the Large model with 16×16 patches. ViT-H/14 (Huge, 14×14 patches) was the largest variant tested.
Positional embeddings: how ViT knows where patches are
Unlike the original Transformer's fixed sinusoidal encodings, ViT uses learnable positional embeddings — one vector per patch position, trained from scratch. The paper explored 2D-aware alternatives but found simple 1D learned embeddings work just as well.
A striking finding: the learned positional embeddings, when visualized, show clear 2D structure. Nearby positions have similar embeddings, and the grid pattern emerges purely from learning — the model discovers spatial layout without being told about it.
The catch: ViT needs big data
The paper's central finding is a scaling story with two acts:
Act 1 — Small data: Trained on ImageNet alone (~1.3M images), ViT-Large underperforms a comparable ResNet (BiT). Without the CNN's inductive biases (locality, translation invariance), the Transformer must learn spatial structure from data — and 1.3M images are not enough.
Act 2 — Big data: Pre-trained on JFT-300M (~300M images), ViT-Huge surpasses the best CNN. Given sufficient data, learned attention patterns are more flexible and powerful than hard-coded convolutional structure. The Transformer's lack of inductive bias becomes an advantage — it doesn't assume, it discovers.
The crossover point is somewhere around 10–100M images. Below it, CNNs win on efficiency. Above it, ViT wins on capacity.
What does ViT actually learn?
Inspecting trained ViT models reveals a clear hierarchy across layers:
- Early layers — attention is mostly local. Attention heads attend to nearby patches, mimicking what convolutions do naturally. The model re-invents locality.
- Middle layers — a mix of local and global patterns. Some heads specialize in attending to patches of similar color or texture.
- Late layers — attention becomes global. Heads connect semantically related regions regardless of position: all parts of a dog, background vs. foreground. This is what CNNs struggle to achieve without very deep stacks.
The model also learns to use the [CLS] token as an information aggregator: by the final layer, it attends broadly to the most informative patches and ignores background clutter.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def create_patch_embeddings(image, patch_size=16, d_model=768):
"""Split image into patches and project to d_model dimensions."""
H, W, C = image.shape # e.g. (224, 224, 3)
n_patches = (H // patch_size) * (W // patch_size) # 196 for 224/16
# Reshape into (n_patches, patch_size*patch_size*C) = (196, 768)
patches = image.reshape(
H // patch_size, patch_size,
W // patch_size, patch_size, C
).transpose(0, 2, 1, 3, 4).reshape(n_patches, -1)
# Linear projection (in practice, a learned weight matrix)
W_e = np.random.randn(patches.shape[1], d_model) * 0.02
return patches @ W_e # (196, 768)
def vit_forward(image, n_layers=12, n_heads=12, d_model=768):
"""Simplified ViT forward pass."""
# 1. Patch embeddings
patch_emb = create_patch_embeddings(image, d_model=d_model)
# 2. Prepend [CLS] token
cls_token = np.zeros((1, d_model))
tokens = np.vstack([cls_token, patch_emb]) # (197, 768)
# 3. Add positional embeddings
pos_emb = np.random.randn(tokens.shape[0], d_model) * 0.02
z = tokens + pos_emb
# 4. Transformer encoder layers (simplified)
for _ in range(n_layers):
# Self-attention: every patch attends to every patch
z = layer_norm(z)
z = z + multi_head_attention(z, z, z, n_heads)
# MLP: per-patch "thinking"
z = layer_norm(z)
z = z + mlp(z)
# 5. Classification: take [CLS] token → MLP head
cls_output = z[0] # the [CLS] representation
return classification_head(cls_output)
# That's it. The SAME Transformer that reads text now reads images.
# The only vision-specific part: splitting the image into patches.Results: ViT vs. the best CNNs
When pre-trained on JFT-300M:
- ViT-H/14 achieved 88.55% top-1 accuracy on ImageNet — a new state of the art, beating the previous best CNN (Noisy Student EfficientNet at 88.4%) while using 4× less compute to pre-train.
- On CIFAR-100, ViT reached 94.55%.
- On the 19-task VTAB benchmark, ViT excelled on Natural and Specialized tasks but showed that Structured tasks (requiring geometric reasoning) remain harder.
The efficiency story is crucial: ViT-H/14 used ~2,500 TPUv3-core-days to pre-train, vs. ~9,900 for BiT-L (ResNet-152×4) — nearly 4× cheaper for better results.
What ViT unlocked
2020
Vision Transformer (ViT)
Proved a pure Transformer can match CNNs on image classification when pre-trained at scale. No convolutions needed.
2021
DeiT — Data-efficient Image Transformers
Showed ViT can be competitive with CNNs using ImageNet alone (no JFT) via knowledge distillation and strong augmentation. Made ViT practical for everyone.
2021
Swin Transformer — Shifted Windows
Introduced hierarchical feature maps and window-based attention, making Transformers competitive for detection and segmentation — tasks needing multi-scale features.
2021
CLIP — Contrastive Language–Image Pretraining
Paired a ViT image encoder with a text encoder, trained on 400M image-text pairs. Learned visual concepts from natural language, enabling zero-shot classification.
2021
DINO — Self-supervised ViT
Trained ViT without labels. The emerging attention maps spontaneously segment objects — proving the architecture discovers semantics on its own.
2022
BEiT & MAE — Masked Image Modeling
Applied BERT's masking idea to images — mask patches, predict them. Self-supervised pre-training that works with ViT out of the box.
2023
SAM — Segment Anything Model
A ViT-based foundation model for segmentation, trained on 1B masks. Prompt it with a click, box, or text — it segments any object in any image.
ViT's most lasting contribution is not a number on a leaderboard — it is the proof that the Transformer is domain-agnostic. Language, images, proteins, audio, video: the same architecture, the same attention mechanism, the same scaling laws. That unification is why today's frontier models handle text and images in a single forward pass.
CitationDosovitskiy, Beyer, Kolesnikov, Weissenborn, Zhai, Unterthiner, Dehghani, Minderer, Heigold, Gelly, Uszkoreit, Houlsby. An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale. ICLR, 2021.
Terms in this paper
- Vision Transformer (ViT)محوِّل الرؤية (ViT)
- Patch Embeddingتضمين الرُّقع
- Mean Attention Distanceمتوسط مسافة الانتباه
- JFT-300MJFT-300M