Computer Vision2021intermediate12 min read
BEiT: BERT Pre-Training of Image Transformers
BEiT: التدريب المسبق بأسلوب BERT لمحوِّلات الصور
Bao, H. · Dong, L. · Piao, S. · Wei, F. — ICLR
The problem
BERT revolutionized NLP by masking words and predicting them from context. But applying the same idea to images is not straightforward. In text, words come from a fixed vocabulary — you can predict "cat" from a list of 30,000 tokens. Image patches have no such vocabulary: each 16×16 patch is a continuous array of pixel values. A naive approach would predict raw pixels, but pixel-level regression wastes model capacity on low-level texture details instead of learning high-level visual semantics. Before BEiT, no method had successfully adapted BERT-style masked prediction to train Vision Transformers in a self-supervised way that outperformed supervised pre-.
The contribution
BEiT (Bidirectional representation from Image Transformers) introduces (MIM) — the visual analogue of BERT's masked language modeling. The key insight is a two-view framework: each image is represented both as raw patches (input to the ) and as discrete (prediction targets from a pre-trained dVAE tokenizer with an 8,192-entry ). During pre-training, ~40% of patches are masked with , and the model predicts the corresponding visual tokens. BEiT-Base achieves 83.2% top-1 on ImageNet-1K (vs. 81.8% for DeiT trained from scratch), and BEiT-Large reaches 86.3% — surpassing -L supervised on ImageNet-22K (85.2%). On ADE20K , BEiT-Large achieves 47.7 mIoU with intermediate .
The impact
BEiT proved that BERT-style self-supervised pre-training works for vision, launching the masked image modeling revolution. It directly inspired MAE (which dropped the tokenizer and predicted raw pixels with a much higher ), SimMIM, iBOT, and BEiT v2/v3. The idea that images can be "tokenized" into a discrete vocabulary — bridging vision and language — became foundational for multimodal models. BEiT showed the computer vision community that self-supervised Transformers can match or beat supervised training, ending the assumption that labeled data is always necessary.
Imagine you're learning a foreign language by doing fill-in-the-blank exercises. BERT does this with text: cover a word, guess it from context. But what if you tried this with photographs? You can't "guess a pixel" the way you guess a word — there's no dictionary of pixels.
BEiT's trick: first, create a visual dictionary. A separate model (a ) learns to compress each image patch into one of 8,192 "visual words." Now you have a vocabulary for images! Hide some patches, and instead of guessing raw pixels, guess which visual word each hidden patch represents. It's like learning to describe a covered puzzle piece using a fixed set of stickers — much more meaningful than guessing individual dots.
The problem: images have no vocabulary
BERT's masked language modeling works beautifully because language has a natural vocabulary. When you mask the word "cat" in a sentence, the model predicts it from a softmax over ~30,000 word tokens. The prediction target is a clean categorical distribution.
Images don't have this luxury. A Vision Transformer splits an image into patches (e.g., 16×16 pixels), but each patch is a continuous vector of 768 values (16 × 16 × 3 RGB channels). If you mask patches and try to predict raw pixel values, you're doing regression — and the model spends most of its capacity learning low-level textures, edge patterns, and color gradients rather than understanding what objects are in the image.
This is the fundamental challenge: how do you create a "vocabulary" for image patches so that masked prediction teaches high-level visual understanding, not pixel memorization?
Step 1: Build a visual vocabulary with a discrete VAE
Before pre-training the Vision Transformer, BEiT first trains a separate model — a discrete variational (dVAE), borrowed from DALL-E's tokenizer — to create a "visual dictionary."
The dVAE works like a compression codebook. It takes a 224×224 image, encodes it to a 14×14 grid of continuous features, then quantizes each grid position to the nearest entry in a learnable codebook of 8,192 visual tokens. The reconstructs the image from these tokens. After training, the codebook becomes the image vocabulary — each of the 8,192 entries represents a recurring visual pattern (a type of texture, edge, color region, or object part).
Crucially, the dVAE is trained once and then frozen. During BEiT pre-training, only the Vision Transformer learns — the tokenizer just provides the prediction targets.
Step 2: Masked image modeling — the core pre-training task
With the visual vocabulary ready, BEiT pre-trains a standard Vision Transformer using masked image modeling (MIM). The process has three stages:
1. Patch the image: Split a 224×224 image into a 14×14 grid of 16×16 patches (196 patches total). Each patch is linearly projected into a d-dimensional .
2. Mask patches: Randomly select ~40% of the patches and replace them with a learnable [MASK] embedding. BEiT uses blockwise masking — instead of masking random scattered patches, it masks contiguous rectangular blocks. This prevents the model from trivially interpolating from immediate neighbors.
3. Predict visual tokens: Feed the corrupted patch sequence through the Transformer encoder. At each masked position, a softmax classifier over the 8,192-entry codebook predicts which visual token the original patch corresponds to. The loss is cross-entropy, computed only at masked positions — exactly like BERT's MLM.
The architecture: standard ViT, nothing new
BEiT's architecture is the standard Vision Transformer (ViT). No architectural changes, no new layers, no custom patterns. The innovation is entirely in the pre-training task.
Two model sizes:
- BEiT-Base: 12 layers, 768 hidden dim, 12 attention heads — 86M parameters
- BEiT-Large: 24 layers, 1024 hidden dim, 16 attention heads — 304M parameters
The input is a sequence of 196 patch embeddings (14×14 grid from a 224×224 image with 16×16 patches) plus a learnable [CLS] token. Position embeddings are learned. During pre-training, masked patches are replaced with a shared [MASK] embedding. A linear softmax head over the 8,192-token codebook is applied at each masked position.
The BERT–BEiT parallel: a side-by-side comparison
BEiT deliberately mirrors BERT's design philosophy. Understanding the parallel makes the contribution crystal clear:
In BERT: the input units are word tokens from a fixed vocabulary (WordPiece, ~30K entries). The pre-training task is masked language modeling — mask 15% of tokens, predict them from bidirectional context. The prediction target is a softmax over the word vocabulary.
In BEiT: the input units are image patches (16×16 pixels), which have no natural vocabulary. The pre-training task is masked image modeling — mask ~40% of patches, predict visual tokens from the remaining patches. The prediction target is a softmax over the visual vocabulary (8,192 entries from the dVAE codebook).
The key difference: BERT's vocabulary is pre-existing (words are already discrete). BEiT must create its vocabulary first by training the dVAE tokenizer. This extra step is what makes the bridge from language to vision possible.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def tokenize_image(image, dvae_encoder, codebook):
"""Use the frozen dVAE to convert image patches to visual tokens."""
features = dvae_encoder(image) # (14, 14, d)
# Quantize: find nearest codebook entry for each position
tokens = np.argmin(
np.linalg.norm(features[:,:,None] - codebook[None,None,:], axis=-1),
axis=-1
) # (14, 14) of ints in [0, 8191]
return tokens
def blockwise_mask(grid_h, grid_w, mask_ratio=0.4):
"""Generate a blockwise mask — contiguous rectangles, not scattered dots."""
num_patches = grid_h * grid_w
num_masked = int(num_patches * mask_ratio)
mask = np.zeros(num_patches, dtype=bool)
while mask.sum() < num_masked:
# Random block dimensions
bh = np.random.randint(2, grid_h // 2)
bw = np.random.randint(2, grid_w // 2)
top = np.random.randint(0, grid_h - bh + 1)
left = np.random.randint(0, grid_w - bw + 1)
for r in range(top, top + bh):
for c in range(left, left + bw):
mask[r * grid_w + c] = True
return mask
def beit_mim_loss(vit, image, dvae_encoder, codebook):
"""One training step for BEiT's masked image modeling."""
# Step 1: Get visual tokens (frozen dVAE)
visual_tokens = tokenize_image(image, dvae_encoder, codebook) # (14,14)
# Step 2: Create blockwise mask
mask = blockwise_mask(14, 14, mask_ratio=0.4) # ~40% masked
# Step 3: Forward pass — masked patches use [MASK] embedding
hidden = vit.encode(image, mask) # (196, d_model)
# Step 4: Predict visual tokens at masked positions only
loss = 0
count = 0
for i in range(196):
if mask[i]:
logits = hidden[i] @ vit.token_head.T # (8192,)
target = visual_tokens.flatten()[i]
loss += cross_entropy(logits, target)
count += 1
return loss / count
# After pre-training: fine-tune with a task-specific head,
# exactly like BERT — same model, different output layer.Fine-tuning: one pre-trained model, many vision tasks
After pre-training, BEiT is fine-tuned exactly like BERT: keep the pre-trained encoder and add a task-specific head.
For on ImageNet: the [CLS] token's final representation is passed through an over all patch representations, then a linear classifier. The entire model is fine-tuned end-to-end.
For semantic segmentation on ADE20K: BEiT serves as the encoder, and a task-specific decoder (such as UperNet) is added on top. An intermediate fine-tuning step — first fine-tuning on ImageNet classification, then transferring to segmentation — further boosts performance.
No part of the architecture changes between tasks. The same pre-trained weights provide the foundation; only the output head differs. This is the pre-train → fine-tune paradigm that BERT established for language, now proven to work for vision.
Results: self-supervised beats supervised
BEiT demonstrated that self-supervised pre-training can outperform supervised pre-training for Vision Transformers — a landmark result:
- BEiT-Base on ImageNet-1K: 83.2% top-1 accuracy, vs. 81.8% for DeiT trained from scratch with the same architecture. A 1.4-point gain from pre-training alone.
- BEiT-Large on ImageNet-1K: 86.3% top-1, surpassing ViT-L supervised on ImageNet-22K (85.2%) — even though BEiT-Large was pre-trained on ImageNet-1K without labels.
- ADE20K semantic segmentation: BEiT-Large achieved 47.7 mIoU with intermediate fine-tuning (first fine-tune on ImageNet, then on ADE20K), improving on previous self-supervised methods.
The ablation studies were equally revealing: predicting visual tokens outperformed pixel regression, and blockwise masking outperformed random masking — especially on semantic segmentation where spatial reasoning matters most.
What BEiT unlocked
2021
BEiT — masked image modeling with visual tokens
First BERT-style self-supervised method to outperform supervised ViT pre-training. Used dVAE tokenizer and blockwise masking at 40%.
2021
MAE — drop the tokenizer, predict pixels at 75% masking
Masked Autoencoders removed the dVAE entirely and showed that predicting raw pixels works if you mask 75% of patches and use an asymmetric encoder-decoder. Simpler pipeline, no separate tokenizer training.
2022
SimMIM — simplify masking, predict pixels with a linear head
Showed that even simple random masking with direct pixel prediction and a lightweight head can match BEiT, questioning the necessity of discrete tokenizers.
2022
BEiT v2 — semantic tokenizer with VQ-KD
Replaced dVAE with a CLIP-aligned tokenizer trained via vector-quantized knowledge distillation. BEiT v2-Base achieved 85.5% on ImageNet — semantic-level tokens proved even more powerful.
BEiT's deepest impact is methodological. Before BEiT, the default approach for training Vision Transformers was supervised pre-training on large labeled datasets like ImageNet-22K or JFT-300M. BEiT showed that a self-supervised objective — one that needs no labels — can produce better representations. This opened the door to training on unlimited unlabeled image data, the same shift that BERT brought to NLP.
The idea that images can be "tokenized" into discrete codes also laid the groundwork for unified vision-language models. If images and text both live in discrete token spaces, they can share the same Transformer architecture and pre-training objectives — exactly the direction taken by BEiT-3 and other multimodal foundation models.
CitationBao, Dong, Piao, Wei. BEiT: BERT Pre-Training of Image Transformers. ICLR, 2022.
Terms in this paper
- Masked Image Modelingنمذجة الصورة المُقنَّعة
- Visual Tokensرموز بصرية
- Image Tokenizerمُرمِّز الصور
- Discrete VAEمرمّز تلقائي متغيّر منفصل
- Vector Quantizationالتكميم المتجهي
- Blockwise Maskingالتقنيع الكتلي
- Masking Ratioنسبة التقنيع
- Vision Transformer (ViT)محوِّل الرؤية (ViT)
- Patch Embeddingتضمين الرُّقع
- Codebookجدول الرموز
- Self-Supervised Learningالتعلم ذاتي الإشراف
- Fine-Tuningالضبط الدقيق
- Semantic Segmentationالتجزئة الدلالية للصورة
- Image Classificationتصنيف وفهرسة الصور
- Pretrainingالتدريب المسبق