Computer Vision2022intermediate13 min read

Masked Autoencoders Are Scalable Vision Learners

المُرمِّزات التلقائية المُقنَّعة: نماذج رؤية قابلة للتدريج

He, K. · Chen, X. · Xie, S. · Li, Y. · Dollár, P. · Girshick, R. — CVPR

The problem

By 2021, Vision Transformers (ViT) had proven that pure attention can match CNNs — but only when pre-trained on hundreds of millions of labeled images (JFT-300M). Without massive labeled data, large ViTs overfit badly. Meanwhile, NLP had solved the same scaling problem with self-supervised pre-training (BERT's masking, GPT's next- prediction) — learning from unlabeled text alone. Could a similar self-supervised recipe work for images? Previous attempts (BEiT, iGPT) either required complex tokenizers, were too slow, or didn't scale.

The contribution

MAE: a strikingly simple self-supervised method. Mask 75% of image patches at random, encode only the visible 25% with a ViT (no mask tokens — a critical efficiency trick), then feed both encoded patches and lightweight mask tokens into a small that reconstructs the missing pixels. Two design insights make this work: (1) the skips mask tokens in the heavy encoder, cutting compute by 3-4×, and (2) a very high forces the model to learn holistic, semantic understanding rather than local interpolation. A vanilla ViT-Huge pre-trained with MAE on ImageNet-1K alone achieves 87.8% accuracy — the best among methods using only ImageNet-1K — while being faster to train than contrastive alternatives.

The impact

MAE proved that BERT-style self-supervised pre-training works for vision, closing the gap with NLP. It showed that large ViTs can be trained on modest datasets (ImageNet-1K) without supervised labels, unlocking a scaling trajectory previously restricted to language. MAE spawned a family of masked modeling methods — VideoMAE, AudioMAE, MultiMAE, Point-MAE — and its representations now power detection, segmentation, and multimodal systems. It demonstrated that simple pixel reconstruction, without complex tokenizers, is sufficient for learning rich visual representations.

BERT taught language models to fill in the blanks: mask a few words in a sentence and predict them from context. It worked brilliantly — because language is dense with meaning, and every word matters.

But images aren't sentences. A missing of blue sky can be guessed from its neighbors without understanding anything. So how do you make a "fill-in-the-blank" game hard enough to teach real vision?

MAE's answer: throw away most of the puzzle. Mask 75% of the image — so much that local texture clues vanish. Now the model must understand objects, scenes, and spatial relationships to reconstruct what's missing. It's like solving a jigsaw puzzle where three out of four pieces are gone.

The gap: why BERT's recipe doesn't directly work for images

BERT masks 15% of tokens and predicts them. This works because language is information-dense: every word carries semantic weight, and even masking a few words creates a hard prediction task.

Images are different in three fundamental ways:

  • High . A missing patch can often be filled by copying textures from neighboring patches — no understanding required. Masking 15% (like BERT) would be too easy.
  • Pixels vs. words. BERT's decoder predicts words — rich semantic tokens. An image decoder predicts pixels — low-level values. The decoder must be designed carefully so the encoder learns semantic representations, not just pixel statistics.
  • Architecture mismatch. Until ViT, vision used CNNs, where mask tokens and positional embeddings don't integrate naturally. ViT removed this obstacle by treating image patches as tokens.

The idea: mask aggressively, encode efficiently, decode lightly

MAE's recipe has three steps, each with a design insight:

1. Mask 75% of patches. Take a 224×224 image, split it into 196 patches (as in ViT), and randomly remove 147 of them. Only 49 patches survive. This extreme masking ratio is critical: it eliminates local texture cues and forces the model to reason about objects, structures, and spatial relationships.

2. Encode only the visible patches. Here is the efficiency breakthrough. Unlike BERT, which feeds mask tokens through its entire encoder, MAE feeds only the 25% visible patches through the heavy ViT encoder. No mask tokens enter the encoder at all. Since cost is O(n²), processing 25% of tokens means roughly 6% of the attention cost — a massive speedup.

3. Decode with a lightweight network. After encoding, append learned mask tokens (one per missing patch), add positional embeddings so each token knows its location, and pass everything through a small decoder (8 blocks, 512-d width — only ~9% of the encoder's compute). The decoder reconstructs the missing pixels. After pre-training, the decoder is discarded — only the encoder is kept for downstream tasks.

Open in Lab
Drag the masking ratio slider. At 75%, the model must understand the scene holistically to reconstruct the missing patches.
The demo wakes as you arrive…

The asymmetric trick: encoder sees only real patches

The key to MAE's efficiency is an asymmetric encoder-decoder — the encoder is large but sees few tokens, the decoder is small but sees all tokens.

Think of it as a company with a senior analyst and a junior assistant. The analyst (encoder) studies only the visible evidence in depth, producing a rich understanding. The assistant (decoder) takes those deep insights, combines them with placeholder tokens for the missing pieces, and fills in the gaps. The heavy thinking happens on 25% of the data; the lightweight assembly happens on 100%.

This design has a powerful side effect: no train-test mismatch. BERT's encoder sees [MASK] tokens during pre-training but never during — creating a distribution gap. MAE's encoder never sees mask tokens, so it always processes real patches, whether during pre-training or deployment. This eliminates the mismatch and improves accuracy by up to 14% in .

Open in Lab
Click any stage to see how MAE processes an image from masking to reconstruction.
The demo wakes as you arrive…

Why 75%? The masking ratio sweet spot

The masking ratio is not arbitrary — it controls how much the model must reason vs. copy. The paper tested ratios from 10% to 90%:

  • At low ratios (10-40%), the task is too easy. Neighboring patches provide enough local information to fill gaps without understanding the scene. Like a crossword where most letters are already given.
  • At 75% (the sweet spot), local shortcuts fail. The model must recognize objects, understand spatial layouts, and grasp scene semantics. Fine-tuning accuracy peaks at 85.0%.
  • At 90%+, too little information remains. Even a human would struggle to reconstruct the image from 20 out of 196 patches.

This 75% is dramatically higher than BERT's 15%, confirming the key insight: images have far more spatial redundancy than text, so much more must be hidden to create a meaningful learning signal.

Open in Lab
Drag the slider to explore how masking ratio affects accuracy. Notice the sweet spot at 75%.
The demo wakes as you arrive…

What to reconstruct: pixels are enough

A key question for : what should the decoder predict? BEiT predicted discrete visual tokens from a pre-trained dVAE tokenizer — adding complexity and an extra pre-training stage. MAE shows that raw pixels work just as well, with two important details:

The function is (MSE) computed only on masked patches — the model is not penalized for visible patches, which keeps the focus on the reconstruction task.

A simple trick boosts quality: . Before computing the loss, normalize each patch to zero mean and unit variance. This forces the model to predict relative contrast and structure rather than absolute brightness, pushing it toward higher-frequency details. With normalized pixels, accuracy improves by ~0.5% over raw pixels — and matches tokenization-based targets, with far less complexity.

L=1∣M∣∑i∈M∥x^i−xi∥2\mathcal{L} = \frac{1}{|\mathcal{M}|} \sum_{i \in \mathcal{M}} \| \hat{x}_i - x_i \|^2
MAE reconstruction loss — MSE on masked patches only — M = set of masked patch indices · x̂ᵢ = decoder's predicted pixels for patch i · xᵢ = original pixels · The loss only penalizes predictions for patches the model never saw

Speed and scale: training ViT-Huge in hours, not weeks

MAE's asymmetric design yields dramatic efficiency gains. Compare the two approaches:

With mask tokens in encoder (like BERT): the encoder processes all 196 tokens including mask tokens. Self-attention computes 196² = 38,416 interactions. Training ViT-Huge for 800 epochs takes ~120 hours on 128 TPUs.

MAE's approach (mask tokens only in decoder): the encoder processes only 49 visible tokens. Self-attention computes 49² = 2,401 interactions — 6% of the cost. The same ViT-Huge trains in ~29 hours — a 4.1× speedup. Memory consumption drops proportionally, enabling even larger models.

This efficiency is why MAE could scale where previous methods could not. BEiT, which uses mask tokens in its encoder, took 36 hours for 300 epochs of ViT-L. MAE's ViT-L trains for 1600 epochs (5× more epochs) in only 31 hours — less total time for far more training.

Open in Lab
Compare the encoder workload with and without mask tokens. The speedup comes from processing 75% fewer tokens in the expensive self-attention layers.
The demo wakes as you arrive…

No sparse operations needed: the shuffle trick

A practical concern: if the encoder only processes visible patches, how do we efficiently select them? MAE uses an elegant shuffle-based implementation. No specialized sparse operations are needed:

  1. Generate position embeddings for all 196 patches.
  2. Randomly shuffle the list of patch tokens.
  3. Keep the first 25% (visible) and discard the rest.
  4. Encode only the kept tokens through the ViT encoder.
  5. After encoding, append mask tokens and unshuffle to restore the original order.
  6. Add positional embeddings and pass through the decoder.

The shuffle and unshuffle operations are extremely fast — they're just index permutations. The entire implementation adds negligible overhead and requires no custom GPU kernels.

MAE pre-training — simplified implementationpython

Simplified to show the idea — not the real implementation.

import numpy as np

def mae_pretrain_step(image, encoder, decoder, mask_ratio=0.75):
    """One MAE pre-training step — mask, encode, decode, compute loss."""
    # 1. Split image into patches (like ViT)
    patches = split_into_patches(image, patch_size=16)  # (196, 768)
    n = len(patches)

    # 2. Random masking via the shuffle trick
    indices = np.random.permutation(n)
    n_visible = int(n * (1 - mask_ratio))              # 49 patches
    visible_idx = sorted(indices[:n_visible])
    masked_idx  = sorted(indices[n_visible:])

    # 3. Encode ONLY visible patches (the key speedup!)
    visible_patches = patches[visible_idx]              # (49, 768)
    visible_emb = visible_patches + pos_embed[visible_idx]
    latent = encoder(visible_emb)                       # (49, 1024)

    # 4. Prepare decoder input: encoded + mask tokens
    mask_tokens = np.tile(learned_mask_token, (len(masked_idx), 1))
    full_tokens = np.zeros((n, decoder_dim))
    full_tokens[visible_idx] = project(latent)          # project to decoder dim
    full_tokens[masked_idx]  = mask_tokens
    full_tokens += decoder_pos_embed                    # add positional info

    # 5. Decode and compute loss on masked patches only
    reconstructed = decoder(full_tokens)                # (196, 768)
    loss = mse(reconstructed[masked_idx], patches[masked_idx])
    return loss

# After pre-training: discard the decoder, keep the encoder.
# Fine-tune the encoder on downstream tasks (classification, detection, etc).

A surprising finding: no augmentation needed

Contrastive self-supervised methods (SimCLR, BYOL, MoCo) depend heavily on — color jittering, random crops, Gaussian blur. Without augmentation, their accuracy drops by 13-28%.

MAE behaves entirely differently. It works well with no augmentation at all — not even horizontal flipping. Adding color jitter actually hurts performance. The reason is elegant: random masking is the augmentation. Each training iteration uses a different random mask, so the same image produces a different learning signal every time. The mask itself generates virtually infinite training diversity, making traditional augmentation redundant.

This is a fundamental advantage: MAE doesn't need to hand-design which transformations preserve semantic meaning. The masking naturally creates a self-supervisory signal without any assumptions about what transformations an image should be invariant to.

Results: state-of-the-art with just ImageNet-1K

MAE's results demonstrate its scalability across model sizes and downstream tasks:

ImageNet-1K classification — ViT-Huge pre-trained with MAE achieves 87.8% top-1 accuracy (with 448 input size), the best among all methods using only ImageNet-1K data. Even at 224 input, ViT-Huge reaches 86.9%. Crucially, MAE's gains grow with model size: the benefit over training from scratch is larger for ViT-Large than ViT-Base, and larger still for ViT-Huge — exactly the scaling behavior seen in NLP.

(COCO) — ViT-L with MAE pre-training achieves 53.3 AP^box, outperforming supervised pre-training by 4.0 points. The simple pixel target matches token-based BEiT.

(ADE20K) — MAE with ViT-L reaches 53.6 mIoU, beating supervised pre-training by 3.7 points and outperforming BEiT.

Robustness — MAE models show dramatically better robustness on ImageNet variants (adversarial, corrupted, sketch). ViT-Huge with MAE scores 68.2% on ImageNet-Adversarial, vs. 33.1% for supervised training — more than double.

Open in Lab
See how MAE's accuracy scales with model size, compared to supervised pre-training and other self-supervised methods.
The demo wakes as you arrive…

Linear probing vs. fine-tuning: the hidden strength

An interesting subtlety: MAE's linear probing accuracy (73.5% for ViT-L) is lower than contrastive methods like MoCo v3 (77.6%). Does this mean MAE learns worse features?

Not at all. When you fine-tune just one Transformer block (leaving the rest frozen), MAE's accuracy jumps from 73.5% to 81.0% — and it beats MoCo v3 at every level of . The gap is 2.6% when tuning just 4 blocks.

The explanation: MAE's representations are strongly non-linear. They encode rich, deep structure that a linear classifier cannot access but a small non-linear head easily unlocks. Contrastive methods produce more linearly separable features (useful for simple probing), but MAE's features are more powerful when even a tiny amount of adaptation is allowed.

This finding challenges the common practice of evaluating self-supervised methods by linear probing alone. is not the full picture — richness matters more for real downstream tasks.

Why it mattered: vision's BERT moment

  1. 2020

    iGPT — Image GPT

    Applied autoregressive prediction to image pixels. Proved the concept but was impractically slow — worked on low-resolution images only.

  2. 2021

    BEiT — BERT Pre-Training of Image Transformers

    Masked patch prediction with discrete visual tokens from a dVAE. Strong results but required a complex pre-trained tokenizer and was slower than MAE.

  3. 2022

    MAE — Masked Autoencoders

    Simple pixel reconstruction with asymmetric encoder-decoder. 87.8% on ImageNet-1K. Proved that tokenization is unnecessary and simple is better.

  4. 2022

    SimMIM — Simple Masked Image Modeling

    Concurrent work showing masked image modeling works with Swin Transformers too. Used mask tokens in the encoder (unlike MAE) but confirmed masking's broad effectiveness.

  5. 2022

    VideoMAE — Masked Autoencoders for Video

    Extended MAE to video with tube masking across frames. Showed that temporal redundancy in video allows even higher masking ratios (90-95%).

  6. 2023

    SAM — Segment Anything Model

    A foundation model for segmentation built on a MAE-pretrained ViT encoder. Trained on 1B masks, it segments any object from a single click. MAE pre-training was key to its data efficiency.

MAE's deepest contribution is a proof of concept: simple self-supervised methods that reconstruct raw signals scale beautifully in vision, just as they do in NLP. You don't need contrastive pairs, hand-designed augmentations, or complex tokenizers. A with pixel targets, trained on unlabeled images, produces representations that rival or exceed years of supervised pre-training research. Vision's path to scale is now clear: more data, more compute, same simple recipe.

CitationHe, Chen, Xie, Li, Dollár, Girshick. Masked Autoencoders Are Scalable Vision Learners. CVPR, 2022.

Terms in this paper