Computer Vision2023intermediate12 min read

Segment Anything

تجزئة أي شيء

Kirillov, A. · Mintun, E. · Ravi, N. · Mao, H. · Rolland, C. · Gustafson, L. · Xiao, T. · Whitehead, S. · Berg, A. C. · Lo, W.-Y. · Dollár, P. · Girshick, R. — ICCV

The problem

By 2023, image was fragmented: each domain — medical imaging, autonomous driving, satellite photos — trained its own specialist model on its own small labeled dataset. There was no "GPT moment" for segmentation: no single model that could segment any object in any image without task-specific training. Meanwhile, collecting segmentation masks at internet scale seemed impossible — masks are far more expensive to annotate than bounding boxes or text labels.

The contribution

Three interconnected innovations: (1) A task — give the model a point, a box, a mask, or text, and it returns a valid . (2) The Segment Anything Model (SAM) — a runs once per image, then a lightweight and produce masks in real time (~50 ms on CPU). The model is : when a prompt is ambiguous, it outputs multiple valid masks. (3) A that co-develops the model and dataset through three stages (assisted-manual → semi-automatic → fully automatic), yielding SA-1B: 1.1 billion masks on 11 million images — 400× more masks than any prior dataset.

The impact

SAM is the first for image segmentation. Its zero-shot generalization — segmenting objects it was never trained on — broke the paradigm of task-specific segmentation models. SAM sparked an ecosystem: SAM 2 extended it to video, HQ-SAM improved boundary quality, MedSAM adapted it for medical imaging, and Grounded-SAM combined it with language. SAM proved that the foundation-model recipe (massive data + promptable task + scale) works for dense prediction, not just language.

Traditional segmentation models are specialist doctors: each one trained for years in one narrow area — cardiology, dermatology, ophthalmology. You need the right specialist for each case, and they cannot help outside their training.

SAM is a general practitioner with X-ray vision: show it any patient (any image), point to where it hurts (a click, a box, or a description), and it instantly outlines the affected region (produces a segmentation mask) — even for conditions it has never studied. Its medical school was seeing a billion cases (masks) across every specialty imaginable.

The problem: segmentation is stuck in silos

Before SAM, the world of image segmentation looked like this: if you needed to segment tumors in CT scans, you trained a model on CT scans. If you needed to segment roads in satellite images, you trained a completely different model. Every new domain required new labeled data, new training, and new expertise. There was no transfer — a model that excelled at medical imaging was useless for autonomous driving.

In natural language processing, foundation models like GPT and BERT had already proven that a single model, pre-trained at massive scale, could transfer to any downstream task. In , ViT and CLIP showed that Transformers could learn general visual representations. But for segmentation — the pixel-level task of outlining every object — no foundation model existed.

The bottleneck was data. Segmentation masks are far more expensive to annotate than classification labels or bounding boxes: each mask requires tracing every pixel boundary of every object. The largest existing segmentation dataset at the time had about 2.5 million masks. To build a foundation model, the authors needed orders of magnitude more — and a way to collect that data efficiently.

Open in Lab
See how traditional segmentation requires a separate specialist model per domain, while SAM handles them all.
The demo wakes as you arrive…

The task: promptable segmentation

The key insight is borrowed from NLP: just as GPT takes a text prompt and generates a completion, SAM takes a segmentation prompt and generates a mask. The prompt can be:

  • A point (foreground or background click)
  • A (a rectangle around the object)
  • A rough mask (a coarse outline to refine)
  • Text (a description like "the red car")

For any prompt, the model must return a valid segmentation mask — one that makes sense even if the prompt is ambiguous. If you click the center of a person wearing a jacket, did you mean the whole person, the jacket, or the shirt underneath? SAM's answer: all three. It outputs multiple masks ranked by confidence, and the user picks the one they meant.

This ambiguity-awareness is crucial. A point on the pocket of a jacket could refer to the pocket, the jacket, or the person. Rather than guessing, SAM produces all plausible interpretations. This design makes SAM useful for interactive annotation, where a human iteratively refines the segmentation.

Open in Lab
Try the different prompt types — click a point, draw a box, or combine them to see how SAM responds.
The demo wakes as you arrive…

The model: image encoder + prompt encoder + mask decoder

SAM's architecture has three components, each with a distinct role — think of it as an assembly line with three stations:

Station 1 — Image Encoder (the heavy lifter). A ViT-H (Vision , Huge variant) pre-trained with MAE (). It processes a 1024×1024 image and produces a 64×64 grid of 256-dimensional embeddings — a dense that captures everything about the image. This is the slowest step, but it runs only once per image. The resulting embeddings can be cached and reused for many prompts.

Station 2 — Prompt Encoder (the translator). Converts whatever prompt the user gives into an the decoder can understand. Sparse prompts (points, boxes) become positional embeddings plus learned type embeddings. Dense prompts (masks) are embedded via convolutions. Text prompts use a CLIP text encoder. This component is extremely lightweight.

Station 3 — Mask Decoder (the fast finisher). A lightweight two-layer Transformer decoder that takes the image embeddings and prompt embeddings and produces the final mask. It uses in both directions — prompts attend to the image, and the image attends back to the prompts — then upsamples the result to pixel resolution. It outputs three masks (for ambiguity) plus a confidence score per mask. This runs in about 50 ms on a CPU — fast enough for interactive use in a web browser.

Open in Lab
Click each component to explore SAM's three-part architecture.
The demo wakes as you arrive…

Handling ambiguity: three masks, not one

A single point is inherently ambiguous — it could refer to the smallest object under the cursor, a medium-sized enclosing object, or the whole scene. SAM resolves this by predicting three masks at different granularity levels (whole, part, subpart), each with an IoU (intersection-over-union) confidence score. The model learns to rank them, and the highest-IoU mask is used by default.

This is inspired by how object hierarchies work in images: a click on a wheel could mean the tire, the wheel assembly, or the entire car. Rather than committing to one interpretation, SAM presents all three and lets the downstream application — or the human — choose.

Training for ambiguity awareness uses a simple trick: during training, only the mask with the lowest (best match to ) receives a . The other two predicted masks get no signal for that sample — this encourages each "slot" to specialize in a different granularity level.

Open in Lab
Click a point on the image and see SAM's three mask predictions at different granularity levels.
The demo wakes as you arrive…

The data engine: model and data grow together

Segmentation masks are not abundant on the internet — unlike text or even image-level labels, you cannot scrape a billion masks from the web. The authors needed a new strategy to build their dataset, so they invented a data engine: a feedback loop where the model helps annotators collect data, the new data improves the model, and the improved model accelerates further data collection. The engine has three stages:

Stage 1 — Assisted-Manual. Professional annotators use a browser-based tool powered by an early SAM (initially trained on public segmentation datasets). They click to indicate objects, SAM suggests masks, and annotators refine them. After six retraining cycles, per-mask annotation time dropped from 34 to 14 seconds, and masks per image grew from 20 to 44. Result: 4.3 million masks on 120,000 images.

Stage 2 — Semi-Automatic. SAM now detects some objects on its own. Annotators focus on the objects SAM missed, boosting diversity — especially for less prominent objects that annotators skipped in Stage 1. Result: 10.2 million masks on 300,000 images.

Stage 3 — Fully Automatic. The model is mature enough to work alone. A 32×32 grid of points is placed over each image, SAM predicts masks at each point, and post-processing (confidence thresholds, NMS, stability checks) filters them down to roughly 100 high-quality masks per image. Applied to all 11 million images, this yields 1.1 billion masks. This is SA-1B — 400× more masks than any previous dataset.

Open in Lab
Explore the three stages of the data engine and see how data volume and model quality co-evolved.
The demo wakes as you arrive…

Inside the mask decoder

The mask decoder is where image understanding meets prompt understanding. It is a modified Transformer decoder with two layers, and each layer performs four steps:

  1. on tokens — the prompt tokens attend to each other (a point can attend to a box, or to other points).
  2. Cross-attention from tokens to image — the prompt tokens query the image embedding to find the regions they refer to. This is the key step: the "where am I pointing?" question.
  3. Point-wise MLP — each token is independently transformed, similar to the feed-forward network in the standard Transformer.
  4. Cross-attention from image to tokens — the image embedding attends back to the prompt tokens, incorporating prompt information into every spatial position.

After two such layers, the decoder upsamples the output with two layers (from 64×64 to 256×256) and produces a between the upsampled features and the output tokens to generate the final mask at pixel resolution.

M=DotProduct(Upsample(D(FI,EP)),  MLP(tout))M = \text{DotProduct}\Big(\text{Upsample}\big(D(F_I, E_P)\big),\; \text{MLP}(t_{\text{out}})\Big)
SAM mask prediction — image meets prompt — F_I = image encoder output · E_P = prompt embeddings · D = two-layer Transformer decoder · t_out = output tokens · Upsample = transposed convolutions from 64×64 to 256×256 · M = final binary mask at pixel resolution

The idea in code

SAM forward pass — simplifiedpython

Simplified to show the idea — not the real implementation.

import numpy as np

def sam_forward(image, prompt_points=None, prompt_box=None):
    """Simplified SAM: image encoder → prompt encoder → mask decoder."""

    # 1. Image encoder: ViT-H processes the image ONCE
    image_embedding = vit_h_encoder(image)      # (64, 64, 256)

    # 2. Prompt encoder: convert clicks/boxes to embeddings
    prompt_tokens = []
    if prompt_points is not None:
        for (x, y, is_foreground) in prompt_points:
            pos_emb = positional_encoding(x, y)  # where?
            type_emb = FG_EMBEDDING if is_foreground else BG_EMBEDDING
            prompt_tokens.append(pos_emb + type_emb)

    if prompt_box is not None:
        x1, y1, x2, y2 = prompt_box
        prompt_tokens.append(positional_encoding(x1, y1) + TOP_LEFT_EMB)
        prompt_tokens.append(positional_encoding(x2, y2) + BOTTOM_RIGHT_EMB)

    # 3. Mask decoder: two-layer Transformer + upsampling
    tokens = np.stack(prompt_tokens + [output_token])
    for layer in decoder_layers:               # 2 layers
        tokens = self_attention(tokens)         # tokens attend to each other
        tokens = cross_attn(tokens, image_embedding)  # tokens → image
        tokens = mlp(tokens)
        image_embedding = cross_attn(image_embedding, tokens)  # image → tokens

    # 4. Upsample and predict 3 masks + confidence scores
    upsampled = transpose_conv(image_embedding)  # (256, 256, 32)
    masks = [dot_product(upsampled, mlp_head_k(tokens[-1])) for k in range(3)]
    iou_scores = iou_head(tokens[-1])           # which mask is best?

    return masks, iou_scores

# Key insight: step 1 runs ONCE. Steps 2-4 run per prompt in ~50ms.
# This is why SAM can be interactive in a web browser.

Zero-shot results: never trained, still competitive

SAM was evaluated on 23 diverse segmentation datasets — none of which it was trained on. The results demonstrated that a foundation model for segmentation is not only possible but can rival supervised specialists:

  • Edge detection (BSDS500): without any training for edge detection, SAM produced reasonable edge maps. It predicted more edges than the ground truth, but the extras were often sensible boundaries the annotators had missed.
  • Object proposals (LVIS v1): SAM outperformed the supervised ViTDet-H on medium and large objects, and on rare and common categories. Its ambiguity-aware multi-mask output was critical — an ablation removing ambiguity awareness significantly hurt recall.
  • (COCO & LVIS): when prompted with ViTDet bounding boxes, SAM's zero-shot masks were competitive with the fully supervised ViTDet masks. On the higher-quality LVIS annotations, the gap shrank further. A human quality study rated SAM masks higher than both ViTDet masks and the COCO ground truth itself.
Open in Lab
Compare SAM's zero-shot performance against supervised baselines across tasks.
The demo wakes as you arrive…

What SAM unlocked

  1. 2023

    SAM — Segment Anything Model

    The first foundation model for segmentation. 1.1 billion masks, zero-shot transfer to 23 datasets. Open-sourced model and SA-1B dataset.

  2. 2023

    Grounded-SAM

    Combined SAM with a grounding detector (Grounding DINO). Users describe objects in text, the grounding model provides boxes, and SAM segments them.

  3. 2023

    HQ-SAM — High-Quality SAM

    Improved SAM's boundary quality by adding a learnable HQ output token, achieving state of the art on fine-grained segmentation benchmarks.

  4. 2023

    MedSAM — Medical SAM

    Fine-tuned SAM on over 1 million medical images covering 11 modalities. Made foundation-model segmentation accessible to clinical applications.

  5. 2024

    SAM 2 — Segment Anything in Images and Videos

    Extended SAM to video segmentation with a memory bank module for tracking objects across frames. Used the more efficient Hiera architecture as image encoder.

SAM's most lasting contribution is not a number on a benchmark — it is the proof that the foundation-model paradigm works for dense visual prediction. Where GPT showed that a single language model could handle any text task, and ViT showed that Transformers work for images, SAM showed that a single segmentation model, trained at sufficient scale, can outline anything in any image. The data engine idea — model and data co-evolving in a virtuous loop — is now a template for building foundation models in domains where labeled data is scarce.

CitationKirillov, Mintun, Ravi, Mao, Rolland, Gustafson, Xiao, Whitehead, Berg, Lo, Dollár, Girshick. Segment Anything. ICCV, 2023.

Terms in this paper