Visual Question Answering
الإجابة البصرية عن الأسئلة
مهمة رؤية ولغة يتلقى فيها النموذج صورة وسؤالاً بلغة طبيعية ويجب عليه توليد إجابة نصية. يصوغ RT-2 التحكم الروبوتي كمسألة إجابة بصرية.
A vision-language task where the model receives an image and a natural language question and must generate a textual answer. RT-2 formats robot control as a VQA problem.
Also translated asالإجابة على الأسئلة المرئية، الاستفسار البصري
First appears in this corpus in: VQA: Visual Question Answering (2015)
Appears in these papers
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models2023in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- HuggingGPT: Solving AI Tasks with ChatGPT and Its Friends in Hugging Face2023in the sky ✦
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control2023in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦