Vision-Language Model
نموذج الرؤية واللغة
شبكة عصبية مُدرَّبة على معالجة الصور والنصوص معاً، عادةً على بيانات بحجم الإنترنت. تستطيع وصف الصور والإجابة عن أسئلة بصرية وأداء مهام متعددة الوسائط.
A neural network trained to jointly process images and text, typically on internet-scale data. VLMs can caption images, answer visual questions, and perform various multimodal tasks.
Also translated asالمعالج البصري اللغوي الـمدمج، النموذج البصري اللغوي، بنية الربط الدلالي بين الصور والنصوص، نموذج بصري-لغوي، نموذج رؤية-لغة، نموذج مُشترك للرؤية واللغة
First appears in this corpus in: VQA: Visual Question Answering (2015)
Appears in these papers
- ALIGN: Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision2021in the sky ✦
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- Learning Transferable Visual Models from Natural Language Supervision2021in the sky ✦
- A Generalist Agent2022in the sky ✦
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2024in the sky ✦
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2024in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦