Computer Vision2021intermediate12 min read

BEiT: BERT Pre-Training of Image Transformers

BEiT: التدريب المسبق بأسلوب BERT لمحوِّلات الصور

Bao, H. · Dong, L. · Piao, S. · Wei, F. — ICLR

The problem

BERT revolutionized NLP by masking words and predicting them from context. But applying the same idea to images is not straightforward. In text, words come from a fixed vocabulary — you can predict "cat" from a list of 30,000 tokens. Image patches have no such vocabulary: each 16×16 patch is a continuous array of pixel values. A naive approach would predict raw pixels, but pixel-level regression wastes model capacity on low-level texture details instead of learning high-level visual semantics. Before BEiT, no method had successfully adapted BERT-style masked prediction to train Vision Transformers in a self-supervised way that outperformed supervised pre-.

The contribution

BEiT (Bidirectional representation from Image Transformers) introduces (MIM) — the visual analogue of BERT's masked language modeling. The key insight is a two-view framework: each image is represented both as raw patches (input to the ) and as discrete (prediction targets from a pre-trained dVAE tokenizer with an 8,192-entry ). During pre-training, ~40% of patches are masked with , and the model predicts the corresponding visual tokens. BEiT-Base achieves 83.2% top-1 on ImageNet-1K (vs. 81.8% for DeiT trained from scratch), and BEiT-Large reaches 86.3% — surpassing -L supervised on ImageNet-22K (85.2%). On ADE20K , BEiT-Large achieves 47.7 mIoU with intermediate .

The impact

BEiT proved that BERT-style self-supervised pre-training works for vision, launching the masked image modeling revolution. It directly inspired MAE (which dropped the tokenizer and predicted raw pixels with a much higher ), SimMIM, iBOT, and BEiT v2/v3. The idea that images can be "tokenized" into a discrete vocabulary — bridging vision and language — became foundational for multimodal models. BEiT showed the computer vision community that self-supervised Transformers can match or beat supervised training, ending the assumption that labeled data is always necessary.

Imagine you're learning a foreign language by doing fill-in-the-blank exercises. BERT does this with text: cover a word, guess it from context. But what if you tried this with photographs? You can't "guess a pixel" the way you guess a word — there's no dictionary of pixels.

BEiT's trick: first, create a visual dictionary. A separate model (a ) learns to compress each image patch into one of 8,192 "visual words." Now you have a vocabulary for images! Hide some patches, and instead of guessing raw pixels, guess which visual word each hidden patch represents. It's like learning to describe a covered puzzle piece using a fixed set of stickers — much more meaningful than guessing individual dots.

The problem: images have no vocabulary

BERT's masked language modeling works beautifully because language has a natural vocabulary. When you mask the word "cat" in a sentence, the model predicts it from a softmax over ~30,000 word tokens. The prediction target is a clean categorical distribution.

Images don't have this luxury. A Vision Transformer splits an image into patches (e.g., 16×16 pixels), but each patch is a continuous vector of 768 values (16 × 16 × 3 RGB channels). If you mask patches and try to predict raw pixel values, you're doing regression — and the model spends most of its capacity learning low-level textures, edge patterns, and color gradients rather than understanding what objects are in the image.

This is the fundamental challenge: how do you create a "vocabulary" for image patches so that masked prediction teaches high-level visual understanding, not pixel memorization?

Open in Lab
The two-view framework: image patches are the input, visual tokens are the prediction target. Toggle masking to see how the pipeline works.
The demo wakes as you arrive…

Step 1: Build a visual vocabulary with a discrete VAE

Before pre-training the Vision Transformer, BEiT first trains a separate model — a discrete variational (dVAE), borrowed from DALL-E's tokenizer — to create a "visual dictionary."

The dVAE works like a compression codebook. It takes a 224×224 image, encodes it to a 14×14 grid of continuous features, then quantizes each grid position to the nearest entry in a learnable codebook of 8,192 visual tokens. The reconstructs the image from these tokens. After training, the codebook becomes the image vocabulary — each of the 8,192 entries represents a recurring visual pattern (a type of texture, edge, color region, or object part).

Crucially, the dVAE is trained once and then frozen. During BEiT pre-training, only the Vision Transformer learns — the tokenizer just provides the prediction targets.

Open in Lab
See how the dVAE tokenizer compresses image patches into discrete visual tokens from a codebook of 8,192 entries.
The demo wakes as you arrive…

Step 2: Masked image modeling — the core pre-training task

With the visual vocabulary ready, BEiT pre-trains a standard Vision Transformer using masked image modeling (MIM). The process has three stages:

1. Patch the image: Split a 224×224 image into a 14×14 grid of 16×16 patches (196 patches total). Each patch is linearly projected into a d-dimensional .

2. Mask patches: Randomly select ~40% of the patches and replace them with a learnable [MASK] embedding. BEiT uses blockwise masking — instead of masking random scattered patches, it masks contiguous rectangular blocks. This prevents the model from trivially interpolating from immediate neighbors.

3. Predict visual tokens: Feed the corrupted patch sequence through the Transformer encoder. At each masked position, a softmax classifier over the 8,192-entry codebook predicts which visual token the original patch corresponds to. The loss is cross-entropy, computed only at masked positions — exactly like BERT's MLM.

Open in Lab
Compare blockwise masking (BEiT) vs. random masking. Notice how blocks remove entire spatial neighborhoods.
The demo wakes as you arrive…
LMIM=−∑i∈Mlog⁡p(zi∣x∖M)\mathcal{L}_{\text{MIM}} = -\sum_{i \in \mathcal{M}} \log p(z_i \mid \mathbf{x}_{\setminus\mathcal{M}})
Masked Image Modeling loss — For each masked position i in the set M, the model predicts the visual token z_i given all unmasked patches. The loss is the negative log-likelihood summed over masked positions — identical in form to BERT's MLM loss, but with visual tokens replacing word tokens.

The architecture: standard ViT, nothing new

BEiT's architecture is the standard Vision Transformer (ViT). No architectural changes, no new layers, no custom patterns. The innovation is entirely in the pre-training task.

Two model sizes:

  • BEiT-Base: 12 layers, 768 hidden dim, 12 attention heads — 86M parameters
  • BEiT-Large: 24 layers, 1024 hidden dim, 16 attention heads — 304M parameters

The input is a sequence of 196 patch embeddings (14×14 grid from a 224×224 image with 16×16 patches) plus a learnable [CLS] token. Position embeddings are learned. During pre-training, masked patches are replaced with a shared [MASK] embedding. A linear softmax head over the 8,192-token codebook is applied at each masked position.

h0=[e[CLS];  x1E;  …;  xNE]+Epos\mathbf{h}_0 = [\mathbf{e}_{\text{[CLS]}};\; \mathbf{x}_1\mathbf{E};\; \ldots;\; \mathbf{x}_N\mathbf{E}] + \mathbf{E}_{\text{pos}}
BEiT input representation — Each image patch x_i is linearly projected by matrix E into a d-dimensional embedding. Position embeddings E_pos are added element-wise. The [CLS] token is prepended. During pre-training, masked positions use a shared [MASK] embedding instead of the patch projection.

The BERT–BEiT parallel: a side-by-side comparison

BEiT deliberately mirrors BERT's design philosophy. Understanding the parallel makes the contribution crystal clear:

In BERT: the input units are word tokens from a fixed vocabulary (WordPiece, ~30K entries). The pre-training task is masked language modeling — mask 15% of tokens, predict them from bidirectional context. The prediction target is a softmax over the word vocabulary.

In BEiT: the input units are image patches (16×16 pixels), which have no natural vocabulary. The pre-training task is masked image modeling — mask ~40% of patches, predict visual tokens from the remaining patches. The prediction target is a softmax over the visual vocabulary (8,192 entries from the dVAE codebook).

The key difference: BERT's vocabulary is pre-existing (words are already discrete). BEiT must create its vocabulary first by training the dVAE tokenizer. This extra step is what makes the bridge from language to vision possible.

Open in Lab
Click each stage to see how BERT and BEiT solve the same problem in text vs. images.
The demo wakes as you arrive…

The idea in code

BEiT masked image modeling — the core pre-training looppython

Simplified to show the idea — not the real implementation.

import numpy as np

def tokenize_image(image, dvae_encoder, codebook):
    """Use the frozen dVAE to convert image patches to visual tokens."""
    features = dvae_encoder(image)          # (14, 14, d)
    # Quantize: find nearest codebook entry for each position
    tokens = np.argmin(
        np.linalg.norm(features[:,:,None] - codebook[None,None,:], axis=-1),
        axis=-1
    )                                        # (14, 14) of ints in [0, 8191]
    return tokens

def blockwise_mask(grid_h, grid_w, mask_ratio=0.4):
    """Generate a blockwise mask — contiguous rectangles, not scattered dots."""
    num_patches = grid_h * grid_w
    num_masked = int(num_patches * mask_ratio)
    mask = np.zeros(num_patches, dtype=bool)

    while mask.sum() < num_masked:
        # Random block dimensions
        bh = np.random.randint(2, grid_h // 2)
        bw = np.random.randint(2, grid_w // 2)
        top = np.random.randint(0, grid_h - bh + 1)
        left = np.random.randint(0, grid_w - bw + 1)
        for r in range(top, top + bh):
            for c in range(left, left + bw):
                mask[r * grid_w + c] = True
    return mask

def beit_mim_loss(vit, image, dvae_encoder, codebook):
    """One training step for BEiT's masked image modeling."""
    # Step 1: Get visual tokens (frozen dVAE)
    visual_tokens = tokenize_image(image, dvae_encoder, codebook)  # (14,14)

    # Step 2: Create blockwise mask
    mask = blockwise_mask(14, 14, mask_ratio=0.4)  # ~40% masked

    # Step 3: Forward pass — masked patches use [MASK] embedding
    hidden = vit.encode(image, mask)  # (196, d_model)

    # Step 4: Predict visual tokens at masked positions only
    loss = 0
    count = 0
    for i in range(196):
        if mask[i]:
            logits = hidden[i] @ vit.token_head.T   # (8192,)
            target = visual_tokens.flatten()[i]
            loss += cross_entropy(logits, target)
            count += 1
    return loss / count

# After pre-training: fine-tune with a task-specific head,
# exactly like BERT — same model, different output layer.

Fine-tuning: one pre-trained model, many vision tasks

After pre-training, BEiT is fine-tuned exactly like BERT: keep the pre-trained encoder and add a task-specific head.

For on ImageNet: the [CLS] token's final representation is passed through an over all patch representations, then a linear classifier. The entire model is fine-tuned end-to-end.

For semantic segmentation on ADE20K: BEiT serves as the encoder, and a task-specific decoder (such as UperNet) is added on top. An intermediate fine-tuning step — first fine-tuning on ImageNet classification, then transferring to segmentation — further boosts performance.

No part of the architecture changes between tasks. The same pre-trained weights provide the foundation; only the output head differs. This is the pre-train → fine-tune paradigm that BERT established for language, now proven to work for vision.

Results: self-supervised beats supervised

BEiT demonstrated that self-supervised pre-training can outperform supervised pre-training for Vision Transformers — a landmark result:

  • BEiT-Base on ImageNet-1K: 83.2% top-1 accuracy, vs. 81.8% for DeiT trained from scratch with the same architecture. A 1.4-point gain from pre-training alone.
  • BEiT-Large on ImageNet-1K: 86.3% top-1, surpassing ViT-L supervised on ImageNet-22K (85.2%) — even though BEiT-Large was pre-trained on ImageNet-1K without labels.
  • ADE20K semantic segmentation: BEiT-Large achieved 47.7 mIoU with intermediate fine-tuning (first fine-tune on ImageNet, then on ADE20K), improving on previous self-supervised methods.

The ablation studies were equally revealing: predicting visual tokens outperformed pixel regression, and blockwise masking outperformed random masking — especially on semantic segmentation where spatial reasoning matters most.

Open in Lab
Compare BEiT against supervised baselines and other self-supervised methods.
The demo wakes as you arrive…

What BEiT unlocked

  1. 2021

    BEiT — masked image modeling with visual tokens

    First BERT-style self-supervised method to outperform supervised ViT pre-training. Used dVAE tokenizer and blockwise masking at 40%.

  2. 2021

    MAE — drop the tokenizer, predict pixels at 75% masking

    Masked Autoencoders removed the dVAE entirely and showed that predicting raw pixels works if you mask 75% of patches and use an asymmetric encoder-decoder. Simpler pipeline, no separate tokenizer training.

  3. 2022

    SimMIM — simplify masking, predict pixels with a linear head

    Showed that even simple random masking with direct pixel prediction and a lightweight head can match BEiT, questioning the necessity of discrete tokenizers.

  4. 2022

    BEiT v2 — semantic tokenizer with VQ-KD

    Replaced dVAE with a CLIP-aligned tokenizer trained via vector-quantized knowledge distillation. BEiT v2-Base achieved 85.5% on ImageNet — semantic-level tokens proved even more powerful.

BEiT's deepest impact is methodological. Before BEiT, the default approach for training Vision Transformers was supervised pre-training on large labeled datasets like ImageNet-22K or JFT-300M. BEiT showed that a self-supervised objective — one that needs no labels — can produce better representations. This opened the door to training on unlimited unlabeled image data, the same shift that BERT brought to NLP.

The idea that images can be "tokenized" into discrete codes also laid the groundwork for unified vision-language models. If images and text both live in discrete token spaces, they can share the same Transformer architecture and pre-training objectives — exactly the direction taken by BEiT-3 and other multimodal foundation models.

CitationBao, Dong, Piao, Wei. BEiT: BERT Pre-Training of Image Transformers. ICLR, 2022.

Terms in this paper