Glossary

Visual Grounding

التأريض البصري

قدرة النموذج على ربط المفاهيم اللغوية بمناطق محددة في الصورة. في VQA، التأريض يعني أن النموذج ينظر إلى المنطقة الصحيحة عند الإجابة عن سؤال حول جزء معين من المشهد.

The ability of a model to link linguistic concepts to specific regions in an image. In VQA, grounding means the model attends to the correct region when answering a question about a specific part of the scene.

Also translated asالإرساء البصري، الربط بالمشهد

First appears in this corpus in: Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection (2023)

Appears in these papers