Computer Vision2020intermediate13 min read
End-to-End Object Detection with Transformers
كشف الكائنات من البداية إلى النهاية باستخدام المحوِّلات
Carion, N. · Massa, F. · Synnaeve, G. · Usunier, N. · Kirillov, A. · Zagoruyko, S. — ECCV
The problem
By 2020, dominant object detectors — Faster R-CNN, RetinaNet, YOLO — all relied on hand-crafted components: anchor boxes defining candidate locations, (NMS) to remove duplicate predictions, and multi-stage pipelines stitched together with task-specific heuristics. These components required careful manual tuning for each dataset and weren't truly differentiable . The field lacked a clean formulation that could predict a set of objects directly, without any post-processing.
The contribution
DETR (DEtection ): the first fully end-to-end object detector built on a Transformer . It reformulates detection as a problem — no anchors, no NMS, no hand-designed assignment rules. A CNN extracts features, a Transformer encoder captures global context, and a Transformer decoder takes N learned object queries and outputs N predictions in parallel. A set-based loss using the for ensures each ground-truth object is uniquely assigned to one prediction. DETR matches Faster R-CNN on COCO (42 AP), excels on large objects (61.1 AP_L), and extends naturally to .
The impact
DETR proved that Transformers can replace the entire hand-engineered detection pipeline. It launched a family of successors — Deformable DETR, DAB-DETR, DN-DETR, DINO, and Grounding DINO — that fixed its slow convergence and small-object weakness while keeping the set-prediction paradigm. The idea of learned object queries became a standard interface for vision tasks. DETR's influence extends beyond detection into segmentation, tracking, and vision-language grounding, making it a pivotal bridge between CNN-era detectors and modern vision Transformers.
Traditional object detectors work like a metal detector on a beach: you sweep a sensor over the sand inch by inch, mark every beep, then go back and remove duplicate marks that are too close together. You need to choose the sensor size (anchor boxes), decide the sweep pattern, and manually resolve overlapping signals (NMS).
DETR works like a team of 100 scouts sitting in a watchtower overlooking the entire beach. Each scout is responsible for spotting exactly one item. They all look at the whole scene at once, communicate with each other to avoid claiming the same object, and a coordinator (the Hungarian algorithm) makes sure nobody doubles up. No sweeping, no cleanup — just one clean report.
The problem: detection pipelines are a patchwork of heuristics
Before DETR, every state-of-the-art detector followed the same script. First, define thousands of anchor boxes — predefined rectangles of various sizes and aspect ratios tiled across the image. Each anchor proposes "maybe there's an object here." Second, a neural network scores each anchor and refines its coordinates. Third, since many anchors overlap the same object, apply non-maximum suppression (NMS) to remove near-duplicate predictions, keeping only the highest-scoring one in each cluster.
This pipeline has three fundamental problems. (1) Anchor design is fragile: the number, sizes, and aspect ratios of anchors must be hand-tuned per dataset — get them wrong and the detector misses objects or wastes compute. (2) NMS is a heuristic, not a learned operation: it uses a fixed threshold to decide which boxes to discard, and this threshold interacts unpredictably with crowded scenes. (3) The pipeline is not end-to-end differentiable: NMS and anchor assignment break the gradient flow, so the system can't be optimized as a unified whole.
The core question DETR answers is: Can we replace this entire patchwork with a single differentiable model that directly outputs a set of detections?
The key insight: detection as set prediction
DETR's breakthrough is a shift in problem formulation. Instead of asking "for each location, is there an object?", DETR asks: "What is the set of objects in this image?"
A set has no order — (cat, dog) is the same as (dog, cat). A set has no duplicates — each object appears exactly once. These two properties are exactly what we want from a detector, and they're exactly what traditional pipelines struggle with (ordering via anchor assignment, duplicates via NMS).
To predict a set, DETR makes N predictions (N=100 by default, larger than the typical number of objects in an image). Most predictions will be "no object" (∅). The must match predictions to ground-truth objects without caring about order. This is where bipartite matching enters.
Bipartite matching and the Hungarian loss
The training process has two steps every iteration. Step 1: Find the best matching. Given N predictions and M ground-truth objects (M ≤ N, padded with ∅ to N), find a one-to-one assignment that minimizes total matching cost. This is solved by the Hungarian algorithm in O(N³) time. The matching cost for pairing prediction σ(i) with ground-truth i combines class probability and box distance.
Step 2: Compute the loss on the matched pairs. Once matching is fixed, compute the standard loss (negative log-likelihood) and box regression loss (L1 + GIoU) only on matched pairs. The ∅ class is down-weighted by factor 10 to handle the class imbalance (most of the 100 predictions are "no object").
The key insight: matching is computed fresh every forward pass — it is not a fixed label assignment. As the model improves, the matching changes, creating a virtuous cycle.
The DETR architecture: CNN + Transformer + FFN
DETR has three components stacked in sequence:
1. CNN Backbone — A standard ResNet-50 (or ResNet-101) extracts a 2D feature map from the input image. An image of size (H₀, W₀) becomes a feature map of size (H₀/32, W₀/32, 2048). A 1×1 convolution reduces the channel dimension from 2048 to d (256 by default), and the spatial dimensions are flattened into a sequence of HW tokens, each with a fixed sinusoidal .
2. Transformer Encoder — The flattened feature sequence passes through 6 encoder layers of standard multi-head + . Every pixel token can attend to every other pixel, so the encoder builds global context — nearby pixels inform each other about object boundaries, and distant pixels share scene-level information. This is where DETR gains its strength on large objects: the receptive field is the entire image from layer 1.
3. Transformer Decoder — N=100 learned object queries (randomly initialized embeddings) are fed into 6 decoder layers. Each decoder layer has three sub-layers: (a) self- among queries — so queries communicate to avoid predicting the same object; (b) from queries to encoder output — each query "looks at" the image to find its object; (c) a feed-forward network that refines the representation. The N output embeddings are each independently decoded by a shared FFN into a class label + .
The idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
class DETR(nn.Module):
def __init__(self, backbone, transformer, num_queries=100, d_model=256, num_classes=91):
super().__init__()
self.backbone = backbone # ResNet-50: image → feature map
self.input_proj = nn.Conv2d(2048, d_model, 1) # reduce channels
self.transformer = transformer # standard encoder-decoder
self.query_embed = nn.Embedding(num_queries, d_model) # 100 learned queries
self.class_head = nn.Linear(d_model, num_classes + 1) # +1 for "no object"
self.bbox_head = nn.Sequential( # predicts normalized (cx, cy, w, h)
nn.Linear(d_model, d_model), nn.ReLU(),
nn.Linear(d_model, d_model), nn.ReLU(),
nn.Linear(d_model, 4), nn.Sigmoid()
)
def forward(self, images):
# Step 1: CNN backbone → feature map
features = self.backbone(images) # (B, 2048, H/32, W/32)
src = self.input_proj(features) # (B, 256, H/32, W/32)
pos = positional_encoding_2d(src) # sinusoidal, same shape
# Step 2: flatten spatial dims → sequence for the transformer
B, C, H, W = src.shape
src = src.flatten(2).permute(2, 0, 1) # (H*W, B, 256)
pos = pos.flatten(2).permute(2, 0, 1)
# Step 3: transformer encoder-decoder
queries = self.query_embed.weight.unsqueeze(1).repeat(1, B, 1) # (100, B, 256)
hs = self.transformer(src, queries, pos_embed=pos) # (100, B, 256)
# Step 4: predict class + box for each query — in parallel
outputs_class = self.class_head(hs) # (100, B, num_classes+1)
outputs_bbox = self.bbox_head(hs) # (100, B, 4)
return outputs_class, outputs_bbox
# No anchors. No NMS. No post-processing.
# The Hungarian algorithm matches predictions ↔ ground truth during training.2D positional encoding: teaching the Transformer about space
Transformers are permutation-equivariant — they don't inherently know the spatial position of each token. DETR uses fixed sinusoidal 2D positional encodings: for each spatial position (x, y) in the feature map, it generates d/2 sine/cosine values for the x-coordinate and d/2 for the y-coordinate, concatenated to form a d-dimensional vector. This encoding is added to the features before the encoder, and also supplied at every decoder cross-attention layer.
The 2D encoding means the model can distinguish "top-left" from "bottom-right" — essential for predicting bounding box coordinates. Unlike learned positional embeddings (as in BERT), sinusoidal encodings generalize to images of different sizes without retraining.
Results on COCO: matching Faster R-CNN, dominating large objects
On the COCO 2017 benchmark, DETR-R50 achieves 42.0 AP — on par with a highly tuned Faster R-CNN with the same ResNet-50 backbone (42.0 AP). With ResNet-101, DETR reaches 43.5 AP. The standout result is on large objects: DETR-R50 scores 61.1 AP_L vs Faster R-CNN's 53.4 — a massive 7.7-point gap, because the Transformer encoder's global attention gives DETR full-image context.
However, DETR struggles with small objects: 20.5 AP_S vs Faster R-CNN's 22.8. Small objects need high-resolution features, and DETR's encoder operates on features at 1/32 resolution. This limitation drove subsequent work like Deformable DETR.
Training is expensive — DETR needs 500 epochs to converge (vs ~36 for Faster R-CNN), which takes about 3 days on 16 V100 GPUs. The long training time is another limitation addressed by later variants.
Extension: panoptic segmentation
DETR extends naturally to panoptic segmentation — the task of assigning every pixel in an image to either a "thing" instance (countable: car, person) or a "stuff" class (uncountable: sky, grass). The extension is elegant: add a small mask prediction head on top of the decoder outputs. Each object query already encodes one object; the mask head produces a binary mask for it using attention maps from the decoder's cross-attention, upsampled with a small FPN-like module. On COCO Panoptic, DETR achieves competitive results, especially on "things" categories, showing that the set prediction framework generalizes beyond bounding boxes.
What the attention sees: decoder attention maps
One of DETR's most compelling properties is interpretability. The cross-attention maps in the decoder show exactly which image regions each object query focuses on. For a query predicting "elephant," the attention map highlights the elephant's body. For a query predicting "person," the map lights up the person. This is far more interpretable than anchor-based detectors, where the assignment between anchors and objects is an opaque combinatorial process.
The self-attention among object queries reveals something equally interesting: queries that predict nearby objects attend to each other, forming implicit spatial reasoning. The model learns to separate overlapping objects through query-to-query communication, replacing the role NMS plays in traditional detectors.
Ablation study: what matters most?
The paper provides thorough ablations:
-
Encoder self-attention is critical: removing the encoder drops AP by 3.9 points (from 42.0 to 38.1). The encoder's global reasoning separates nearby objects — without it, the model struggles with overlapping instances.
-
Number of decoder layers matters: going from 1 to 6 decoder layers improves AP from 33.8 to 42.0. With only 1 layer, the model cannot iteratively refine its predictions, and NMS actually helps (+2.1 AP) — showing that multi-layer decoding replaces NMS's function.
-
FFN in the decoder is important: removing FFNs from decoder layers drops AP by 2.3, confirming that the attention layers alone aren't sufficient for the coordinate regression task.
-
Positional encoding at every decoder layer: supplying positional encodings only at input degrades AP by 1.4. The model needs continuous spatial grounding throughout decoding.
What DETR unlocked
2020
DETR
First end-to-end Transformer-based detector. Set prediction with Hungarian matching, no anchors, no NMS. Matched Faster R-CNN on COCO.
2021
Deformable DETR
Replaced global attention with deformable attention — each query attends to a small set of learned key points. Converges 10× faster, handles multi-scale features, and improves small-object detection.
2022
DAB-DETR & DN-DETR
DAB-DETR replaced learned queries with dynamic anchor boxes as queries. DN-DETR added query denoising during training. Together they cut convergence from 500 to ~50 epochs.
2022
DINO
Combined deformable attention, denoising training, and contrastive learning for query initialization. Achieved 63.3 AP on COCO — best single-model result at the time.
2023
Grounding DINO
Extended the DETR family to open-set detection: detect any object described by text. Fused language and vision in the query mechanism, enabling detection without predefined category lists.
DETR's deepest legacy is proving that detection does not need to be a patchwork of heuristics. The set prediction formulation — learned queries, bipartite matching, parallel decoding — is now the dominant paradigm for Transformer-based detectors. Beyond detection, the same pattern of "learned queries ↔ Hungarian matching" has been adopted for instance segmentation, pose estimation, multi-object tracking, and vision-language grounding. DETR was the bridge between the CNN-era detection pipelines and the modern vision Transformer ecosystem.
CitationCarion, Massa, Synnaeve, Usunier, Kirillov, Zagoruyko. End-to-End Object Detection with Transformers. ECCV, 2020.
Terms in this paper
- Object Detectionرصد وتحديد الكائنات
- Transformerالمحوِّل
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Bipartite Matchingالمطابقة الثنائية
- Hungarian Algorithmخوارزمية المجرية
- Non-Maximum Suppressionكبت غير أعظمي
- Anchor Boxمربعات الإحاطة المرجعية (المرساة)
- Set Predictionالتنبؤ بالمجموعات
- Object Queryاستعلام الكائن
- Positional Encodingالترميز الموضعي
- Cross-Attentionالانتباه التبادلي
- Self-Attentionالانتباه الذاتي
- Feed Forward Network (FFN)شبكة التغذية الأمامية
- Bounding Boxمربع الإحاطة
- Panoptic Segmentationالتجزئة الشاملة