Computer Vision2020intermediate13 min read

End-to-End Object Detection with Transformers

كشف الكائنات من البداية إلى النهاية باستخدام المحوِّلات

Carion, N. · Massa, F. · Synnaeve, G. · Usunier, N. · Kirillov, A. · Zagoruyko, S. — ECCV

The problem

By 2020, dominant object detectors — Faster R-CNN, RetinaNet, YOLO — all relied on hand-crafted components: anchor boxes defining candidate locations, (NMS) to remove duplicate predictions, and multi-stage pipelines stitched together with task-specific heuristics. These components required careful manual tuning for each dataset and weren't truly differentiable . The field lacked a clean formulation that could predict a set of objects directly, without any post-processing.

The contribution

DETR (DEtection ): the first fully end-to-end object detector built on a Transformer . It reformulates detection as a problem — no anchors, no NMS, no hand-designed assignment rules. A CNN extracts features, a Transformer encoder captures global context, and a Transformer decoder takes N learned object queries and outputs N predictions in parallel. A set-based loss using the for ensures each ground-truth object is uniquely assigned to one prediction. DETR matches Faster R-CNN on COCO (42 AP), excels on large objects (61.1 AP_L), and extends naturally to .

The impact

DETR proved that Transformers can replace the entire hand-engineered detection pipeline. It launched a family of successors — Deformable DETR, DAB-DETR, DN-DETR, DINO, and Grounding DINO — that fixed its slow convergence and small-object weakness while keeping the set-prediction paradigm. The idea of learned object queries became a standard interface for vision tasks. DETR's influence extends beyond detection into segmentation, tracking, and vision-language grounding, making it a pivotal bridge between CNN-era detectors and modern vision Transformers.

Traditional object detectors work like a metal detector on a beach: you sweep a sensor over the sand inch by inch, mark every beep, then go back and remove duplicate marks that are too close together. You need to choose the sensor size (anchor boxes), decide the sweep pattern, and manually resolve overlapping signals (NMS).

DETR works like a team of 100 scouts sitting in a watchtower overlooking the entire beach. Each scout is responsible for spotting exactly one item. They all look at the whole scene at once, communicate with each other to avoid claiming the same object, and a coordinator (the Hungarian algorithm) makes sure nobody doubles up. No sweeping, no cleanup — just one clean report.

The problem: detection pipelines are a patchwork of heuristics

Before DETR, every state-of-the-art detector followed the same script. First, define thousands of anchor boxes — predefined rectangles of various sizes and aspect ratios tiled across the image. Each anchor proposes "maybe there's an object here." Second, a neural network scores each anchor and refines its coordinates. Third, since many anchors overlap the same object, apply non-maximum suppression (NMS) to remove near-duplicate predictions, keeping only the highest-scoring one in each cluster.

This pipeline has three fundamental problems. (1) Anchor design is fragile: the number, sizes, and aspect ratios of anchors must be hand-tuned per dataset — get them wrong and the detector misses objects or wastes compute. (2) NMS is a heuristic, not a learned operation: it uses a fixed threshold to decide which boxes to discard, and this threshold interacts unpredictably with crowded scenes. (3) The pipeline is not end-to-end differentiable: NMS and anchor assignment break the gradient flow, so the system can't be optimized as a unified whole.

The core question DETR answers is: Can we replace this entire patchwork with a single differentiable model that directly outputs a set of detections?

Open in Lab
Compare the traditional anchor-based pipeline with DETR's streamlined approach.
The demo wakes as you arrive…

The key insight: detection as set prediction

DETR's breakthrough is a shift in problem formulation. Instead of asking "for each location, is there an object?", DETR asks: "What is the set of objects in this image?"

A set has no order — (cat, dog) is the same as (dog, cat). A set has no duplicates — each object appears exactly once. These two properties are exactly what we want from a detector, and they're exactly what traditional pipelines struggle with (ordering via anchor assignment, duplicates via NMS).

To predict a set, DETR makes N predictions (N=100 by default, larger than the typical number of objects in an image). Most predictions will be "no object" (∅). The must match predictions to ground-truth objects without caring about order. This is where bipartite matching enters.

Open in Lab
Drag predictions to ground-truth objects to find the optimal matching — then see how the Hungarian algorithm does it.
The demo wakes as you arrive…

Bipartite matching and the Hungarian loss

The training process has two steps every iteration. Step 1: Find the best matching. Given N predictions and M ground-truth objects (M ≤ N, padded with ∅ to N), find a one-to-one assignment that minimizes total matching cost. This is solved by the Hungarian algorithm in O(N³) time. The matching cost for pairing prediction σ(i) with ground-truth i combines class probability and box distance.

Step 2: Compute the loss on the matched pairs. Once matching is fixed, compute the standard loss (negative log-likelihood) and box regression loss (L1 + GIoU) only on matched pairs. The ∅ class is down-weighted by factor 10 to handle the class imbalance (most of the 100 predictions are "no object").

The key insight: matching is computed fresh every forward pass — it is not a fixed label assignment. As the model improves, the matching changes, creating a virtuous cycle.

σ^=arg⁡min⁡σ∈SN∑i=1NLmatch ⁣(yi,  y^σ(i))\hat{\sigma} = \underset{\sigma \in \mathfrak{S}_N}{\arg\min} \sum_{i=1}^{N} \mathcal{L}_{\text{match}}\!\bigl(y_i,\;\hat{y}_{\sigma(i)}\bigr)
Optimal bipartite matching — solved by the Hungarian algorithm — Find the permutation σ of N predictions that minimizes total matching cost against the N ground-truth targets (padded with ∅). The cost combines class probability and box similarity.
Lmatch(yi,y^σ(i))=−1{ci≠∅} p^σ(i)(ci)  +  1{ci≠∅} Lbox(bi,b^σ(i))\mathcal{L}_{\text{match}}(y_i, \hat{y}_{\sigma(i)}) = -\mathbb{1}_{\{c_i \neq \varnothing\}}\,\hat{p}_{\sigma(i)}(c_i) \;+\; \mathbb{1}_{\{c_i \neq \varnothing\}}\,\mathcal{L}_{\text{box}}(b_i, \hat{b}_{\sigma(i)})
Pair-wise matching cost — For each ground-truth target i with class cᵢ and box bᵢ, the cost is the negative class probability plus the box loss. Targets labeled ∅ contribute zero to the matching cost.
Lbox(bi,b^σ(i))=λiou Lgiou(bi,b^σ(i))  +  λL1 ∥bi−b^σ(i)∥1\mathcal{L}_{\text{box}}(b_i, \hat{b}_{\sigma(i)}) = \lambda_{\text{iou}}\,\mathcal{L}_{\text{giou}}(b_i, \hat{b}_{\sigma(i)}) \;+\; \lambda_{\text{L1}}\,\|b_i - \hat{b}_{\sigma(i)}\|_1
Box regression loss — GIoU + L1 — The box loss is a weighted combination of the Generalized IoU loss (scale-invariant) and the L1 loss (penalizes absolute coordinate error). Using both gives better convergence than either alone. In the paper, λ_iou = 2 and λ_L1 = 5.

The DETR architecture: CNN + Transformer + FFN

DETR has three components stacked in sequence:

1. CNN Backbone — A standard ResNet-50 (or ResNet-101) extracts a 2D feature map from the input image. An image of size (H₀, W₀) becomes a feature map of size (H₀/32, W₀/32, 2048). A 1×1 convolution reduces the channel dimension from 2048 to d (256 by default), and the spatial dimensions are flattened into a sequence of HW tokens, each with a fixed sinusoidal .

2. Transformer Encoder — The flattened feature sequence passes through 6 encoder layers of standard multi-head + . Every pixel token can attend to every other pixel, so the encoder builds global context — nearby pixels inform each other about object boundaries, and distant pixels share scene-level information. This is where DETR gains its strength on large objects: the receptive field is the entire image from layer 1.

3. Transformer Decoder — N=100 learned object queries (randomly initialized embeddings) are fed into 6 decoder layers. Each decoder layer has three sub-layers: (a) self- among queries — so queries communicate to avoid predicting the same object; (b) from queries to encoder output — each query "looks at" the image to find its object; (c) a feed-forward network that refines the representation. The N output embeddings are each independently decoded by a shared FFN into a class label + .

Open in Lab
Click on any component to explore how data flows through DETR.
The demo wakes as you arrive…

The idea in code

DETR forward pass — simplified PyTorch pseudocodepython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

class DETR(nn.Module):
    def __init__(self, backbone, transformer, num_queries=100, d_model=256, num_classes=91):
        super().__init__()
        self.backbone = backbone               # ResNet-50: image → feature map
        self.input_proj = nn.Conv2d(2048, d_model, 1)  # reduce channels
        self.transformer = transformer         # standard encoder-decoder
        self.query_embed = nn.Embedding(num_queries, d_model)  # 100 learned queries
        self.class_head = nn.Linear(d_model, num_classes + 1)  # +1 for "no object"
        self.bbox_head = nn.Sequential(         # predicts normalized (cx, cy, w, h)
            nn.Linear(d_model, d_model), nn.ReLU(),
            nn.Linear(d_model, d_model), nn.ReLU(),
            nn.Linear(d_model, 4), nn.Sigmoid()
        )

    def forward(self, images):
        # Step 1: CNN backbone → feature map
        features = self.backbone(images)                      # (B, 2048, H/32, W/32)
        src = self.input_proj(features)                       # (B, 256, H/32, W/32)
        pos = positional_encoding_2d(src)                     # sinusoidal, same shape

        # Step 2: flatten spatial dims → sequence for the transformer
        B, C, H, W = src.shape
        src = src.flatten(2).permute(2, 0, 1)                 # (H*W, B, 256)
        pos = pos.flatten(2).permute(2, 0, 1)

        # Step 3: transformer encoder-decoder
        queries = self.query_embed.weight.unsqueeze(1).repeat(1, B, 1)  # (100, B, 256)
        hs = self.transformer(src, queries, pos_embed=pos)    # (100, B, 256)

        # Step 4: predict class + box for each query — in parallel
        outputs_class = self.class_head(hs)                   # (100, B, num_classes+1)
        outputs_bbox = self.bbox_head(hs)                     # (100, B, 4)
        return outputs_class, outputs_bbox

# No anchors. No NMS. No post-processing.
# The Hungarian algorithm matches predictions ↔ ground truth during training.

2D positional encoding: teaching the Transformer about space

Transformers are permutation-equivariant — they don't inherently know the spatial position of each token. DETR uses fixed sinusoidal 2D positional encodings: for each spatial position (x, y) in the feature map, it generates d/2 sine/cosine values for the x-coordinate and d/2 for the y-coordinate, concatenated to form a d-dimensional vector. This encoding is added to the features before the encoder, and also supplied at every decoder cross-attention layer.

The 2D encoding means the model can distinguish "top-left" from "bottom-right" — essential for predicting bounding box coordinates. Unlike learned positional embeddings (as in BERT), sinusoidal encodings generalize to images of different sizes without retraining.

Results on COCO: matching Faster R-CNN, dominating large objects

On the COCO 2017 benchmark, DETR-R50 achieves 42.0 AP — on par with a highly tuned Faster R-CNN with the same ResNet-50 backbone (42.0 AP). With ResNet-101, DETR reaches 43.5 AP. The standout result is on large objects: DETR-R50 scores 61.1 AP_L vs Faster R-CNN's 53.4 — a massive 7.7-point gap, because the Transformer encoder's global attention gives DETR full-image context.

However, DETR struggles with small objects: 20.5 AP_S vs Faster R-CNN's 22.8. Small objects need high-resolution features, and DETR's encoder operates on features at 1/32 resolution. This limitation drove subsequent work like Deformable DETR.

Training is expensive — DETR needs 500 epochs to converge (vs ~36 for Faster R-CNN), which takes about 3 days on 16 V100 GPUs. The long training time is another limitation addressed by later variants.

Open in Lab
Compare DETR vs Faster R-CNN across object sizes — toggle between AP metrics.
The demo wakes as you arrive…

Extension: panoptic segmentation

DETR extends naturally to panoptic segmentation — the task of assigning every pixel in an image to either a "thing" instance (countable: car, person) or a "stuff" class (uncountable: sky, grass). The extension is elegant: add a small mask prediction head on top of the decoder outputs. Each object query already encodes one object; the mask head produces a binary mask for it using attention maps from the decoder's cross-attention, upsampled with a small FPN-like module. On COCO Panoptic, DETR achieves competitive results, especially on "things" categories, showing that the set prediction framework generalizes beyond bounding boxes.

What the attention sees: decoder attention maps

One of DETR's most compelling properties is interpretability. The cross-attention maps in the decoder show exactly which image regions each object query focuses on. For a query predicting "elephant," the attention map highlights the elephant's body. For a query predicting "person," the map lights up the person. This is far more interpretable than anchor-based detectors, where the assignment between anchors and objects is an opaque combinatorial process.

The self-attention among object queries reveals something equally interesting: queries that predict nearby objects attend to each other, forming implicit spatial reasoning. The model learns to separate overlapping objects through query-to-query communication, replacing the role NMS plays in traditional detectors.

Open in Lab
Click an object query to see which image regions it attends to across decoder layers.
The demo wakes as you arrive…

Ablation study: what matters most?

The paper provides thorough ablations:

  • Encoder self-attention is critical: removing the encoder drops AP by 3.9 points (from 42.0 to 38.1). The encoder's global reasoning separates nearby objects — without it, the model struggles with overlapping instances.

  • Number of decoder layers matters: going from 1 to 6 decoder layers improves AP from 33.8 to 42.0. With only 1 layer, the model cannot iteratively refine its predictions, and NMS actually helps (+2.1 AP) — showing that multi-layer decoding replaces NMS's function.

  • FFN in the decoder is important: removing FFNs from decoder layers drops AP by 2.3, confirming that the attention layers alone aren't sufficient for the coordinate regression task.

  • Positional encoding at every decoder layer: supplying positional encodings only at input degrades AP by 1.4. The model needs continuous spatial grounding throughout decoding.

What DETR unlocked

  1. 2020

    DETR

    First end-to-end Transformer-based detector. Set prediction with Hungarian matching, no anchors, no NMS. Matched Faster R-CNN on COCO.

  2. 2021

    Deformable DETR

    Replaced global attention with deformable attention — each query attends to a small set of learned key points. Converges 10× faster, handles multi-scale features, and improves small-object detection.

  3. 2022

    DAB-DETR & DN-DETR

    DAB-DETR replaced learned queries with dynamic anchor boxes as queries. DN-DETR added query denoising during training. Together they cut convergence from 500 to ~50 epochs.

  4. 2022

    DINO

    Combined deformable attention, denoising training, and contrastive learning for query initialization. Achieved 63.3 AP on COCO — best single-model result at the time.

  5. 2023

    Grounding DINO

    Extended the DETR family to open-set detection: detect any object described by text. Fused language and vision in the query mechanism, enabling detection without predefined category lists.

DETR's deepest legacy is proving that detection does not need to be a patchwork of heuristics. The set prediction formulation — learned queries, bipartite matching, parallel decoding — is now the dominant paradigm for Transformer-based detectors. Beyond detection, the same pattern of "learned queries ↔ Hungarian matching" has been adopted for instance segmentation, pose estimation, multi-object tracking, and vision-language grounding. DETR was the bridge between the CNN-era detection pipelines and the modern vision Transformer ecosystem.

CitationCarion, Massa, Synnaeve, Usunier, Kirillov, Zagoruyko. End-to-End Object Detection with Transformers. ECCV, 2020.

Terms in this paper