Computer Vision2015intermediate10 min read
VQA: Visual Question Answering
VQA: الإجابة البصرية عن الأسئلة
Antol, S. · Agrawal, A. · Lu, J. · Mitchell, M. · Batra, D. · Zitnick, C. L. · Parikh, D. — ICCV
The problem
By 2015, models could describe a scene in a generic sentence, but there was no rigorous way to test whether a system truly understood the image. A caption like "a man on a field" reveals little about whether the model knows what sport he plays, how many people are watching, or what the weather is. There was no large-scale, open-ended that probed fine-grained visual understanding through natural language.
The contribution
VQA: a new task, a large-scale , and a suite of baselines. The task requires answering free-form natural language questions about images. The dataset contains ~0.25M COCO images, ~0.76M questions, and ~10M human answers. Baselines combine a (VGGNet) for image features with Bag-of-Words or LSTMs for question encoding, fused via element-wise multiplication and classified over the top-1000 answers. The best baseline (deeper Q + norm I) reached 58.16% open-ended — far below human performance at 83.30%.
The impact
VQA defined the understanding benchmark that drove a decade of research. It revealed that simple fusion of CNN and LSTM features is far from sufficient — models need , grounding, and reasoning. The VQA Challenge became the standard competition, spurring Stacked Attention Networks, Bottom-Up/Top-Down Attention, and eventually vision-language models like CLIP and LLaVA that approach human-level performance.
Imagine a tour guide at a museum. A generic audio guide just says "This is a painting of a garden" — that's image captioning. But a real tour guide can answer any question: "What season is it?", "How many people are in the scene?", "Is the painter using warm or cool colors?"
VQA turns an AI into that tour guide: it must look at any image and answer any question about it. The questions probe things a generic caption would never mention — background details, spatial relationships, common-sense reasoning, and even reading text in the image.
Why VQA? The limits of image captioning
Before VQA, the main way to evaluate image understanding was captioning: generate a sentence describing the image. But captioning has two blind spots:
- It's shallow. A model can say "a dog on a couch" without knowing the dog's breed, color, or what it's doing.
- Evaluation is hard. Two correct captions can use completely different words, making automatic scoring unreliable.
VQA flips the paradigm: instead of asking the model to volunteer information, you probe it with specific questions. The answers are short (usually 1–3 words), so evaluation is straightforward. And because different questions target different aspects of the image, VQA tests a much wider range of visual understanding abilities than captioning ever could.
The dataset: 0.76 million questions, 10 million answers
The VQA dataset is built on MS-COCO images — real photographs of everyday scenes. For each image, crowdworkers on Amazon Mechanical Turk wrote three open-ended questions, and ten different annotators answered each question. This redundancy is critical: it captures the natural ambiguity of language (is the answer "two" or "2" or "a couple"?) and enables a soft accuracy metric.
The dataset also includes abstract cartoon scenes — clipart images of people and objects in various configurations — to isolate reasoning from low-level visual recognition. A model that works on cartoons but not photos might be good at reasoning but weak at visual features; the reverse suggests feature strength but reasoning weakness.
What kinds of questions does VQA ask?
The questions span a rich taxonomy of visual reasoning skills:
- Yes/No — "Is there a cat in the image?" Tests basic .
- Counting — "How many cars are parked?" Requires localization and enumeration.
- Color/Attribute — "What color is the umbrella?" Tests fine-grained recognition.
- Spatial — "What is to the left of the person?" Tests relational understanding.
- Activity — "What sport is being played?" Requires scene-level reasoning.
- Common sense — "Is this a good day for a picnic?" Goes beyond what's directly visible.
This diversity is what makes VQA hard. A model that only recognizes objects will fail on spatial and commonsense questions; a model that only reads text will fail on everything visual. Solving VQA demands genuine multimodal intelligence.
The evaluation metric: measuring soft agreement
Because VQA answers are open-ended, the paper introduces a soft accuracy metric that accounts for human disagreement. Each question has 10 human answers. A predicted answer gets credit based on how many humans gave the same answer:
The baselines: how do you fuse an image and a question?
The core architectural challenge in VQA is : how do you combine a visual representation of the image with a linguistic representation of the question to predict an answer?
The paper explores several baselines, each progressively stronger:
- Image only — ignore the question entirely, just predict the most common answer for images that look similar. This tests pure visual bias.
- Question only — ignore the image, predict from the question text alone. This tests (e.g., "What color is the ___?" → "white" is often correct).
- Q + I (Bag-of-Words) — encode the question as a bag of word embeddings, extract image features from a pretrained VGGNet, and combine them.
- LSTM Q + I — replace the bag-of-words with an LSTM that reads the question sequentially, capturing word order and context.
- Deeper LSTM Q + norm I — a two-layer LSTM with 2048-dimensional , combined with L2-normalized image features via element-wise multiplication.
Fusion: element-wise multiplication
The key fusion operation is surprisingly simple. The image (from VGGNet's last , 4096 dimensions) and the question feature vector (from the LSTM's final hidden state, also projected to 1024 or 2048 dimensions) are combined by element-wise multiplication.
Think of it as a selective filter: the question vector acts like a mask that amplifies the image features relevant to the question and suppresses the rest. If the question is about color, the "color-sensitive" dimensions of the image vector get boosted; if it's about counting, the spatial dimensions get more weight.
The fused vector is then passed through a fully connected network (two hidden layers of 1000 units each, with and tanh activation) and a final over the top 1000 most frequent answers.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def softmax(x):
e = np.exp(x - x.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
def vqa_predict(image_features, question_tokens, W_embed, W_lstm, W1, b1, W2, b2):
"""
image_features: (4096,) from VGGNet last hidden layer
question_tokens: list of word indices
"""
# Step 1: Encode the question with an LSTM
h = np.zeros(1024) # LSTM hidden state
for token in question_tokens:
word_vec = W_embed[token] # (300,) word embedding
h = lstm_step(h, word_vec, W_lstm) # update hidden state
q_features = h # final hidden state = question repr.
# Step 2: Normalize image features
img_norm = image_features / (np.linalg.norm(image_features) + 1e-8)
# Step 3: Fuse by element-wise multiplication
fused = img_norm * q_features # the "selective filter"
# Step 4: Classify over top-1000 answers
hidden = np.tanh(W1 @ fused + b1) # first MLP layer
logits = W2 @ hidden + b2 # (1000,)
probs = softmax(logits)
return np.argmax(probs) # predicted answer indexResults: how far from human?
The results reveal a striking gap between machines and humans. The best baseline reached 58.16% open-ended accuracy vs. human performance at 83.30%. But the breakdown by question type tells a more nuanced story:
On Yes/No questions, models perform reasonably well (around 75%) because the answer space is binary and visual bias helps. On Number questions, performance drops significantly (~33%) because counting requires precise localization. On "Other" questions (colors, activities, objects), models achieve ~42% — better than random but far from understanding.
The most revealing finding: the question-only baseline scores 48.09% overall, meaning nearly half the questions can be "answered" without ever looking at the image. This exposed a fundamental dataset bias that later led to VQA v2.0, where complementary images with opposite answers were added to force models to actually look.
The language bias problem
Abstract scenes: isolating reasoning from recognition
One of VQA's most clever design choices is the inclusion of abstract (clipart) scenes alongside real photographs. In a real photo, a model must handle lighting, occlusion, texture, and countless low-level visual challenges. In an , objects are clean, unambiguous icons.
By comparing performance on real vs. abstract scenes, the paper can disentangle two sources of error: visual recognition failure (the model can't see what's there) vs. reasoning failure (the model sees the objects but can't answer the question). This diagnostic split was influential — it showed that even with perfect object recognition, reasoning remained the harder challenge.
What VQA revealed about AI understanding
VQA exposed three critical gaps in the AI systems of 2015:
- No attention mechanism. The CNN produces a single global feature vector for the entire image. It has no way to focus on the region relevant to the question. This directly inspired Stacked Attention Networks (2016) and Bottom-Up/Top-Down Attention (2018).
- No spatial reasoning. Element-wise multiplication treats image features as a flat bag of properties, discarding spatial structure. The model literally can't tell "left" from "right."
- No compositional reasoning. Questions like "What color is the small object next to the red box?" require chaining multiple reasoning steps — something a single multiplication cannot express.
Each of these gaps spawned a major research direction. VQA didn't solve visual understanding — it gave the field a scoreboard that made progress measurable.
The legacy: from VQA to vision-language models
2015
VQA v1.0
The original dataset and task definition. 0.25M images, 0.76M questions. Baselines reach 58% vs. 83% human accuracy.
2016
Stacked Attention Networks (SAN)
Multi-step attention over image regions, guided by the question. Showed that looking at the right place matters as much as recognizing objects.
2017
VQA v2.0 — Making the V matter
Paired each question with complementary images yielding opposite answers, forcing models to actually look at the image instead of guessing from language priors.
2018
Bottom-Up / Top-Down Attention
Used object detection (Faster R-CNN) to propose salient regions, then attention selected which regions to focus on. Won the VQA Challenge.
2021
CLIP
Contrastive pre-training on 400M image-text pairs learned visual concepts from natural language supervision — enabling zero-shot transfer to VQA-like tasks.
2023
LLaVA & GPT-4V
Large vision-language models that combine powerful LLMs with visual encoders, approaching and sometimes matching human-level VQA performance.
VQA's baselines were simple by today's standards. But the dataset and evaluation framework the paper established became the testbed on which the entire field of vision-language AI sharpened its tools — from attention mechanisms to -based multimodal models to today's large vision-language models.
CitationAntol, Agrawal, Lu, Mitchell, Batra, Zitnick, Parikh. VQA: Visual Question Answering. ICCV, 2015.
Terms in this paper
- Visual Question Answeringالإجابة البصرية عن الأسئلة
- Multimodal Modelالنموذج متعدد الأنماط
- Image Classificationتصنيف وفهرسة الصور
- Convolutional Neural Network (CNN)الشبكة العصبية الالتفافية
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Embeddingالتضمين
- Softmaxسوفت ماكس
- Ground Truthالحقيقة الأرضية
- Open Vocabularyالقاموس المفتوح
- Crowdsourcingالتعهيد الجماعي