Computer Vision2020intermediate9 min read

An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale

صورة تساوي 16×16 كلمة: المحوِّلات للتعرف على الصور على نطاق واسع

Dosovitskiy, A. · Beyer, L. · Kolesnikov, A. · Weissenborn, D. · Zhai, X. · Unterthiner, T. · Dehghani, M. · Minderer, M. · Heigold, G. · Gelly, S. · Uszkoreit, J. · Houlsby, N. — ICLR

The problem

By 2020 CNNs dominated computer vision. The had revolutionized NLP, but attempts to apply to images either mixed it with convolutions or used it only on individual pixels — too expensive for real-world resolutions. The question was: can a pure Transformer, with no convolutions at all, match or beat the best CNNs on image ?

The contribution

: split the image into fixed-size 16×16 patches, flatten each patch into a , project it linearly into an , prepend a learnable [CLS] , add positional embeddings, and feed the sequence into a standard Transformer — exactly the one from "Attention Is All You Need." Pre-trained on large datasets (), ViT matched or exceeded the best CNNs (BiT, Noisy Student) on ImageNet, CIFAR-100, and VTAB while requiring substantially fewer compute resources to train.

The impact

ViT proved the Transformer is not language-specific — it is a general-purpose sequence processor. This unlocked a wave of vision Transformers (DeiT, Swin, BEiT, DINO, SAM), multimodal models (CLIP, DALL·E, Stable Diffusion), and ultimately the unified architectures behind today's multimodal AI. Two decades of -only vision ended with one paper.

A CNN reads an image like a person pressing a magnifying glass across a page: small area, high detail, one spot at a time. Useful features trickle upward through layers, and a top-left object can only influence a bottom-right decision after many stacked convolutions.

ViT replaces the magnifying glass with a conference table: the image is cut into 16×16 patches — like snapshots pinned to a board — and every patch sits at the table, free to address any other patch directly. The cat's ear can ask the grass in the background, "are you outdoor context?" in a single step.

The ceiling: CNNs see locally, not globally

CNNs process images with small filters (typically 3×3) that scan the image locally. To connect distant regions, you need many stacked layers — each layer's grows a little, but global context emerges only late in the network. This is a strong : locality and translation invariance are baked in.

These biases help when data is limited — they tell the how to look — but they also limit what the model can learn. In NLP, Transformers had already shown that with enough data, learned attention patterns outperform hand-coded structural assumptions. Could the same happen in vision?

Open in Lab
Compare how a CNN builds context gradually through stacked layers vs. ViT seeing everything in one step.
The demo wakes as you arrive…

The idea: treat image patches as tokens

ViT's insight is elegantly minimal: an image can be turned into a sequence the Transformer already knows how to process. The recipe:

  1. Split a 224×224 image into a grid of non-overlapping 16×16 patches → 196 patches.
  2. Flatten each patch (16 × 16 × 3 = 768 values) and project it linearly into a D-dimensional embedding.
  3. Prepend a learnable [CLS] token — its final representation will carry the image's class.
  4. Add learnable positional embeddings so the model knows which patch came from where.
  5. Feed the sequence of 197 tokens (196 patches + 1 CLS) into a standard Transformer encoder.

That's it. No convolutions, no pooling, no hand-designed feature extractors. The same architecture that reads sentences now reads images.

Open in Lab
Watch a 224×224 image get split into 196 patches, each flattened and projected into an embedding.
The demo wakes as you arrive…
z0=[xclass;  xp1E;  xp2E;  ⋯ ;  xpNE]+Epos\mathbf{z}_0 = [\mathbf{x}_{\text{class}}; \; \mathbf{x}_p^1 E; \; \mathbf{x}_p^2 E; \; \cdots; \; \mathbf{x}_p^N E] + \mathbf{E}_{pos}
ViT input assembly — patches become tokens — xₚⁱ = flattened patch i · E = linear projection to D dimensions · x_class = learnable [CLS] token · E_pos = positional embeddings · z₀ = the input sequence the Transformer sees

The architecture: a Transformer encoder, unchanged

ViT uses the standard Transformer encoder exactly as described in the original paper. Each layer applies:

  • Multi-Head (MSA) — every patch attends to every other patch. A patch showing a wheel can attend to a patch showing a road, building "car on a road" context in one step.
  • applied before each sub-layer (pre-norm), which stabilizes for deep models.
  • block — two linear layers with GELU activation, applied independently to each patch position.
  • Residual connections around both MSA and MLP.

After L layers, the [CLS] token's representation is passed through a small MLP head for classification. The model comes in three sizes: ViT-Base (86M params, 12 layers), ViT-Large (307M, 24 layers), and ViT-Huge (632M, 32 layers).

Open in Lab
Click any layer to see what it does inside the ViT encoder.
The demo wakes as you arrive…

Why 16×16? The patch-size tradeoff

Patch size controls how many tokens the Transformer sees. For a 224×224 image:

  • 32×32 patches → only 49 tokens — fast but coarse. Fine details are lost.
  • 16×16 patches → 196 tokens — the sweet spot. Enough detail, tractable compute.
  • 8×8 patches → 784 tokens — rich detail but self-attention costs O(n²), so 4× the patches means 16× the attention cost.

The naming convention encodes this: ViT-L/16 means the Large model with 16×16 patches. ViT-H/14 (Huge, 14×14 patches) was the largest variant tested.

Open in Lab
Drag the patch size and watch the token count and attention cost change.
The demo wakes as you arrive…

Positional embeddings: how ViT knows where patches are

Unlike the original Transformer's fixed sinusoidal encodings, ViT uses learnable positional embeddings — one vector per patch position, trained from scratch. The paper explored 2D-aware alternatives but found simple 1D learned embeddings work just as well.

A striking finding: the learned positional embeddings, when visualized, show clear 2D structure. Nearby positions have similar embeddings, and the grid pattern emerges purely from learning — the model discovers spatial layout without being told about it.

Open in Lab
Click a patch position and see which other positions have the most similar learned embeddings.
The demo wakes as you arrive…

The catch: ViT needs big data

The paper's central finding is a scaling story with two acts:

Act 1 — Small data: Trained on ImageNet alone (~1.3M images), ViT-Large underperforms a comparable ResNet (BiT). Without the CNN's inductive biases (locality, translation invariance), the Transformer must learn spatial structure from data — and 1.3M images are not enough.

Act 2 — Big data: Pre-trained on JFT-300M (~300M images), ViT-Huge surpasses the best CNN. Given sufficient data, learned attention patterns are more flexible and powerful than hard-coded convolutional structure. The Transformer's lack of inductive bias becomes an advantage — it doesn't assume, it discovers.

The crossover point is somewhere around 10–100M images. Below it, CNNs win on efficiency. Above it, ViT wins on capacity.

Open in Lab
Drag the dataset size to see where ViT overtakes the CNN.
The demo wakes as you arrive…

What does ViT actually learn?

Inspecting trained ViT models reveals a clear hierarchy across layers:

  • Early layers — attention is mostly local. Attention heads attend to nearby patches, mimicking what convolutions do naturally. The model re-invents locality.
  • Middle layers — a mix of local and global patterns. Some heads specialize in attending to patches of similar color or texture.
  • Late layers — attention becomes global. Heads connect semantically related regions regardless of position: all parts of a dog, background vs. foreground. This is what CNNs struggle to achieve without very deep stacks.

The model also learns to use the [CLS] token as an information aggregator: by the final layer, it attends broadly to the most informative patches and ignores background clutter.

The idea in code

ViT patch embedding + Transformer forward passpython

Simplified to show the idea — not the real implementation.

import numpy as np

def create_patch_embeddings(image, patch_size=16, d_model=768):
    """Split image into patches and project to d_model dimensions."""
    H, W, C = image.shape                  # e.g. (224, 224, 3)
    n_patches = (H // patch_size) * (W // patch_size)  # 196 for 224/16

    # Reshape into (n_patches, patch_size*patch_size*C) = (196, 768)
    patches = image.reshape(
        H // patch_size, patch_size,
        W // patch_size, patch_size, C
    ).transpose(0, 2, 1, 3, 4).reshape(n_patches, -1)

    # Linear projection (in practice, a learned weight matrix)
    W_e = np.random.randn(patches.shape[1], d_model) * 0.02
    return patches @ W_e                    # (196, 768)

def vit_forward(image, n_layers=12, n_heads=12, d_model=768):
    """Simplified ViT forward pass."""
    # 1. Patch embeddings
    patch_emb = create_patch_embeddings(image, d_model=d_model)

    # 2. Prepend [CLS] token
    cls_token = np.zeros((1, d_model))
    tokens = np.vstack([cls_token, patch_emb])   # (197, 768)

    # 3. Add positional embeddings
    pos_emb = np.random.randn(tokens.shape[0], d_model) * 0.02
    z = tokens + pos_emb

    # 4. Transformer encoder layers (simplified)
    for _ in range(n_layers):
        # Self-attention: every patch attends to every patch
        z = layer_norm(z)
        z = z + multi_head_attention(z, z, z, n_heads)
        # MLP: per-patch "thinking"
        z = layer_norm(z)
        z = z + mlp(z)

    # 5. Classification: take [CLS] token → MLP head
    cls_output = z[0]           # the [CLS] representation
    return classification_head(cls_output)

# That's it. The SAME Transformer that reads text now reads images.
# The only vision-specific part: splitting the image into patches.

Results: ViT vs. the best CNNs

When pre-trained on JFT-300M:

  • ViT-H/14 achieved 88.55% top-1 accuracy on ImageNet — a new state of the art, beating the previous best CNN (Noisy Student EfficientNet at 88.4%) while using 4× less compute to pre-train.
  • On CIFAR-100, ViT reached 94.55%.
  • On the 19-task VTAB benchmark, ViT excelled on Natural and Specialized tasks but showed that Structured tasks (requiring geometric reasoning) remain harder.

The efficiency story is crucial: ViT-H/14 used ~2,500 TPUv3-core-days to pre-train, vs. ~9,900 for BiT-L (ResNet-152×4) — nearly 4× cheaper for better results.

What ViT unlocked

  1. 2020

    Vision Transformer (ViT)

    Proved a pure Transformer can match CNNs on image classification when pre-trained at scale. No convolutions needed.

  2. 2021

    DeiT — Data-efficient Image Transformers

    Showed ViT can be competitive with CNNs using ImageNet alone (no JFT) via knowledge distillation and strong augmentation. Made ViT practical for everyone.

  3. 2021

    Swin Transformer — Shifted Windows

    Introduced hierarchical feature maps and window-based attention, making Transformers competitive for detection and segmentation — tasks needing multi-scale features.

  4. 2021

    CLIP — Contrastive Language–Image Pretraining

    Paired a ViT image encoder with a text encoder, trained on 400M image-text pairs. Learned visual concepts from natural language, enabling zero-shot classification.

  5. 2021

    DINO — Self-supervised ViT

    Trained ViT without labels. The emerging attention maps spontaneously segment objects — proving the architecture discovers semantics on its own.

  6. 2022

    BEiT & MAE — Masked Image Modeling

    Applied BERT's masking idea to images — mask patches, predict them. Self-supervised pre-training that works with ViT out of the box.

  7. 2023

    SAM — Segment Anything Model

    A ViT-based foundation model for segmentation, trained on 1B masks. Prompt it with a click, box, or text — it segments any object in any image.

ViT's most lasting contribution is not a number on a leaderboard — it is the proof that the Transformer is domain-agnostic. Language, images, proteins, audio, video: the same architecture, the same attention mechanism, the same scaling laws. That unification is why today's frontier models handle text and images in a single forward pass.

CitationDosovitskiy, Beyer, Kolesnikov, Weissenborn, Zhai, Unterthiner, Dehghani, Minderer, Heigold, Gelly, Uszkoreit, Houlsby. An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale. ICLR, 2021.

Terms in this paper