Computer Vision2023advanced14 min read
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Grounding DINO: دمج DINO مع التدريب المسبق المُوجَّه لغوياً لكشف الأجسام المفتوح
Liu, S. · Zeng, Z. · Ren, T. · Li, F. · Zhang, H. · Yang, J. · Li, C. · Yang, J. · Su, H. · Zhu, J. · Zhang, L. — ECCV
The problem
Traditional object detectors are "closed-set": they can only find objects from a fixed list of categories defined at time. If you train on 80 COCO classes, the model is blind to anything outside those 80. Extending detection to open-set scenarios requires injecting language understanding so the model can generalize to arbitrary concepts described in natural language. Previous attempts only fused language and vision at one or two stages of the detection pipeline, leaving cross-modal alignment incomplete. Moreover, most open-set detectors relied on frozen embeddings, which are pretrained on image-text pairs and are suboptimal for region-level detection tasks.
The contribution
Grounding DINO extends the DINO detector with three-phase language fusion: a Feature Enhancer that uses bidirectional between image and text in the , Language-Guided Query Selection that picks the most text-relevant image features as queries, and a Cross-Modality Decoder that adds text cross- layers for fine-grained alignment. A sub-sentence level text representation removes unrelated category interactions during feature extraction. Pre-trained on detection, grounding, and caption data, Grounding DINO achieves 52.5 AP on COCO detection and 26.1 mean AP on ODinW zero-shot — both state-of-the-art — while also excelling at .
The impact
Grounding DINO became a foundational building block for multimodal vision systems. It powers the detection stage in Grounded-SAM (combining it with Segment Anything for open-set segmentation), image editing pipelines with Stable Diffusion, and robotic manipulation systems. Its three-phase fusion design influenced subsequent open-set detectors and vision-language models. With over 1.8 million monthly downloads on Hugging Face, it is one of the most widely deployed open-set detection models in both research and production.
Imagine a security guard at a building entrance. Traditionally, the guard has a photo album of exactly 80 people — if your face is in the album, you're recognized; if not, you're invisible. Now replace the photo album with an earpiece connected to a dispatcher who can say anything: "the person in the red hat," "a delivery box," or "coffee mug on the counter."
The guard now listens to the dispatcher at every step — while scanning the crowd (feature enhancer), while choosing who to focus on (query selection), and while making the final identification (decoder). This deep, three-stage conversation between eyes and ears is what makes Grounding DINO so powerful: it doesn't just glance at the text once; it weaves language into every level of its visual processing.
The problem: closed-set detection hits a wall
Standard models — from Faster R-CNN to DINO — are trained on a fixed set of categories. The COCO benchmark, for instance, defines exactly 80 classes: person, bicycle, car, ... , toothbrush. During training, each detected region is assigned to one of these categories. During , the model can only output those same categories. This is called closed-set detection.
The problem is obvious: the real world has far more than 80 objects. A closed-set detector cannot find a "fire hydrant" if that class wasn't in its training vocabulary. This makes such detectors brittle in open-world applications — autonomous driving, robotics, image editing — where the set of interesting objects changes constantly.
The goal of is to break this limitation. Instead of a fixed class list, the user provides a natural language description — category names like "coffee mug, laptop" or full referring expressions like "the red car on the left" — and the model finds matching objects in the image. This requires the detector to understand language, not just visual patterns.
Architecture: three-phase fusion of vision and language
To understand Grounding DINO's design, first recall how a typical closed-set detector is structured. It has three main stages: a that extracts image features, a (or encoder) that refines and enhances those features, and a head (or decoder) that produces the final detections. In a closed-set detector, language never touches any of these stages — is just a linear projection into a fixed set of classes.
Previous open-set detectors injected language at only one or two stages. GLIP fused features only in the neck (phase A). OV-DETR injected language only at the decoder input (phase B). Grounding DINO's key insight is that fusing language at all three phases produces much better cross-modal alignment.
The architecture is a dual-encoder-single-decoder system. An image backbone () produces visual features. A text backbone () produces per- text features. These two streams then interact in three phases:
Phase A — Feature Enhancer: Bidirectional cross-attention between image and text features in the encoder. Phase B — Language-Guided Query Selection: Selects the most text-relevant image features as queries for the decoder. Phase C — Cross-Modality Decoder: Adds text cross-attention layers inside each decoder layer so queries continuously absorb language information.
Phase A: the Feature Enhancer — teaching eyes and ears to talk
The Feature Enhancer is the first place where image and text features interact. Think of it as a conference room where visual and textual delegates sit together and exchange notes. Each Feature Enhancer layer contains four sub-layers stacked in sequence:
- Deformable on image features — each image token attends to a sparse, learned set of positions across all feature scales, efficiently capturing multi-scale context.
- Vanilla self-attention on text features — text tokens attend to each other to build richer contextual representations.
- Image-to-text cross-attention — image features query text features to absorb language semantics. Each spatial location in the image asks: "which words describe me?"
- Text-to-image cross-attention — text features query image features to ground themselves visually. Each word asks: "which regions in the image correspond to me?"
There are six such layers stacked sequentially. After passing through all six, the image features are language-enriched and the text features are visually grounded — both sides have learned about each other before the decoder even begins.
Phase B: Language-Guided Query Selection — asking the text what to focus on
In DETR-like detectors, the decoder receives a set of "queries" — learnable vectors that each try to detect one object. In the original DINO, these queries are selected based on visual saliency alone. Grounding DINO changes this: it asks the text to guide which image locations become queries.
The idea is simple but powerful. We have image features (over 10,000 tokens from multi-scale feature maps) and text features (up to 256 tokens). We compute the dot product between every image token and every text token, take the maximum text score for each image token, and select the top 900 image tokens as decoder queries.
Intuitively, image locations that strongly match any word in the text prompt become candidates for detection. A location that correlates highly with "cat" or "red" gets promoted; a background patch that matches nothing gets filtered out. This focuses the decoder's attention on the most text-relevant regions from the very start.
Simplified to show the idea — not the real implementation.
# image_feat: (bs, num_img_tokens, dim) # text_feat: (bs, num_text_tokens, dim)
# Step 1: compute pairwise similarity logits = torch.einsum("bic,btc->bit", image_feat, text_feat)
# Step 2: max text relevance per image token logits_per_img = logits.max(dim=-1)[0] # (bs, num_img_tokens)
# Step 3: select top-900 image tokens topk_idx = torch.topk(logits_per_img, num_query, dim=1)[1]Phase C: the Cross-Modality Decoder — continuous language refinement
The decoder in standard DINO has six layers. Each layer contains self-attention among queries, image cross-attention to probe visual features, and an FFN. Grounding DINO adds one more sub-layer to every decoder layer: text cross-attention. The full sequence in each decoder layer is:
- Self-attention among cross-modality queries
- Image cross-attention — queries attend to image features (using for efficiency)
- Text cross-attention — queries attend to text features to absorb language cues
- FFN — feedforward transformation
This extra text cross-attention layer is the key addition. Without it, the queries must carry all language information from the query initialization step alone. With it, queries continuously check back with the text at every layer: "Am I still looking for what the text describes?" This iterative refinement produces tighter alignment between detected regions and language concepts.
At the final decoder layer, each query output is dot-producted with text features to produce per-token logits. A is applied for classification, and L1 + GIoU losses handle regression — all matched via Hungarian , just like standard DETR.
Sub-sentence text representation: blocking unwanted interactions
When using a detector for open-set detection, category names are concatenated into a single string: "cat . dog . mouse ." and fed to the text encoder. But this creates a problem: during self-attention, the word "cat" attends to the word "dog" even though they are unrelated categories. The arbitrary ordering of categories can introduce spurious dependencies.
Previous approaches handled this in two extremes. Sentence-level representations encode the entire string as one , losing per-word granularity. Word-level representations keep per-word features but allow full cross-attention between all words, including unrelated categories.
Grounding DINO introduces a sub-sentence level approach: it uses attention masks to block attention between different category phrases while keeping full attention within each phrase. The word "sleeping" can attend to "cat" within "a cat is sleeping," but "cat" cannot attend to "baseball glove" in a different phrase. For referring expressions, the full sentence remains connected since all words describe one concept.
This simple masking trick eliminates noise from inter-category interactions while preserving fine-grained per-word features — delivering the best of both worlds.
Training: contrastive loss and grounded pre-training
Grounding DINO replaces the standard classification head with a between detected regions and language tokens. Instead of projecting a query into 80 logits, the query is dot-producted with all text token features to produce per-token scores. A focal loss is applied to each score. This formulation naturally supports any number of categories or free-text descriptions — the text defines the class space.
The bounding box regression uses L1 loss and GIoU loss, following standard DETR practice. Predictions and ground truths are matched using Hungarian bipartite matching with costs combining classification (weight 2.0), L1 (weight 5.0), and GIoU (weight 2.0). Auxiliary losses are added after each decoder layer and after encoder outputs.
The model is pre-trained on three types of data. Detection data (Objects365, COCO, OpenImages) where category names are concatenated as text prompts. Grounding data (GoldG, RefCOCO/+/g) with natural language descriptions and region annotations. Caption data (pseudo-labeled from GLIP) that enriches vocabulary by generating detection labels from captions. This heterogeneous training recipe gives the model exposure to both structured category vocabularies and free-form language descriptions.
Results: state-of-the-art across three settings
Grounding DINO is evaluated across three complementary settings that together paint a complete picture of open-set detection capability.
Zero-shot COCO detection: Without seeing any COCO training images, Grounding DINO Large achieves 52.5 AP — surpassing all prior models including GLIP-L (49.8 AP). With , it reaches 63.0 AP on COCO test-dev, outperforming even the original DINO.
Zero-shot ODinW: The ODinW benchmark tests detection across 35 diverse real-world datasets. Grounding DINO Large sets a new record with 26.1 mean AP, outperforming the giant Florence model (25.8 AP) despite using fewer parameters.
Referring Expression Comprehension (RefCOCO/+/g): When fine-tuned with referring expression data, Grounding DINO achieves top-1 accuracy of 90.56% on RefCOCO val, setting state-of-the-art across all splits. Even without RefCOCO training data, it outperforms GLIP under the same settings.
Ablations: which components matter most?
The (Table 7 of the paper) reveals the contribution of each component. All experiments use a Swin-T backbone pre-trained on Objects365.
Encoder fusion (Phase A) has the largest impact: removing it drops COCO zero-shot AP from 46.7 to 45.8 and LVIS AP from 16.1 to 13.1 — a 3-point collapse showing that early cross-modal interaction is critical.
Language-guided query selection (Phase B) contributes +0.4 AP on COCO and a large +2.5 AP on LVIS, showing that text-guided initialization is especially valuable for detecting rare or uncommon objects.
Text cross-attention in decoder (Phase C) adds +0.6 AP on COCO and +1.8 AP on LVIS. While less impactful than encoder fusion, it introduces very few parameters and consistently improves alignment.
Sub-sentence text representation provides +0.3 AP on COCO and +0.5 AP on LVIS — a small but meaningful gain from simply masking attention between unrelated categories.
The ablations confirm that tight three-phase fusion is not just a design principle — each phase provides measurable, complementary gains.
Ecosystem: from detection to segmentation and editing
Grounding DINO's impact extends far beyond detection benchmarks. Its ability to detect arbitrary objects from text makes it a universal "pointing" module that other systems build on.
Grounded-SAM combines Grounding DINO with Segment Anything to create an open-set segmentation pipeline: Grounding DINO detects bounding boxes from text, and SAM segments the detected regions into precise masks. This pipeline requires zero training and generalizes to any object described in language.
Image editing pipelines pair Grounding DINO with Stable Diffusion: detect an object by text, generate a mask, and inpaint the region with a new prompt. The paper demonstrates this with examples like replacing "pandas" with "dogs and birthday cakes."
Robotic manipulation systems use Grounding DINO to locate task-relevant objects from natural language instructions, enabling robots to understand commands like "pick up the red cup" without pre-defined object lists.
This ecosystem effect — where one model becomes a building block for many downstream systems — is a hallmark of foundational AI models.
Timeline: from closed-set detection to open-set grounding
2020
DETR — end-to-end detection with Transformers
Introduced Transformer-based object detection with bipartite matching, eliminating hand-crafted components like non-maximum suppression and anchor generation.
2021
CLIP — connecting vision and language
Trained on 400M image-text pairs with contrastive learning. Showed that visual representations can be aligned with language for zero-shot recognition.
2022
GLIP — grounded language-image pre-training
Reformulated object detection as phrase grounding and performed early fusion in the neck module. Introduced contrastive training between regions and language phrases on large-scale data.
2022
DINO — improved denoising anchor boxes for detection
Advanced DETR with contrastive de-noising, mixed query selection, and anchor box improvements. Set new records on COCO closed-set detection.
2023
Grounding DINO (this paper)
Married DINO's detector with three-phase language fusion and grounded pre-training. Achieved 52.5 AP zero-shot on COCO and 26.1 AP on ODinW, becoming the foundation for Grounded-SAM and multimodal pipelines.
2023
Grounded-SAM — open-set segmentation
Combined Grounding DINO with Segment Anything to create a zero-training pipeline for detecting and segmenting any object described in text.
CitationLiu, Zeng, Ren, Li, Zhang, Yang, Li, Yang, Su, Zhu, Zhang. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. ECCV, 2023.
Terms in this paper
- Open-Set Object Detectionالكشف المفتوح عن الأجسام
- Visual Groundingالتأريض البصري
- Cross-Attentionالانتباه التبادلي
- Contrastive Lossالخسارة التبايُنية
- Deformable Attentionالانتباه القابل للتشوّه
- Object Detectionرصد وتحديد الكائنات
- Referring Expression Comprehensionفهم التعبيرات الإشارية
- Zero-Shot Transferنقل بدون تدريب
- Anchor Boxمربعات الإحاطة المرجعية (المرساة)
- Bounding Boxمربع الإحاطة
- Feature Pyramidهرم السِّمات
- Encoder-Decoderمرمِّز-فاكّ ترميز
- BERTبيرت
- Swin Transformerمحوِّل سوين
- Bipartite Matchingالمطابقة الثنائية