NLP2016beginner11 min read
SQuAD: 100,000+ Questions for Machine Comprehension of Text
SQuAD: أكثر من 100,000 سؤال لفهم الآلة للنصوص
Rajpurkar, P. · Zhang, J. · Lopyrev, K. · Liang, P. — EMNLP
The problem
By 2016, datasets fell into two camps: small, high-quality sets like MCTest (2,640 questions) that were too tiny for modern data-hungry neural networks, and massive -style datasets like CNN/Daily Mail (1.4M) that only required filling in a single entity — not genuine comprehension. There was no large-scale where answers were free-form spans requiring real reasoning over a passage.
The contribution
SQuAD: 107,785 question-answer pairs on 536 Wikipedia articles, where every answer is an exact from the passage. Crowdworkers wrote the questions in their own words, producing diverse reasoning types — lexical variation, syntactic paraphrase, multi-sentence inference. A baseline reached 51.0% F1 versus 86.8% human F1, establishing a large and measurable gap for the community to close. Two evaluation metrics — and F1 — became the standard for .
The impact
SQuAD became the single most important for reading comprehension and a standard evaluation task for every major language model from BiDAF to BERT to GPT. It proved that large, carefully crowdsourced datasets can drive rapid progress: within two years, neural models surpassed human performance. SQuAD's extractive format and F1/EM metrics became the blueprint for dozens of successor benchmarks including SQuAD 2.0, Natural Questions, and the QA components of GLUE and SuperGLUE.
Imagine a reading exam for machines. A teacher hands out a paragraph from an encyclopedia and asks five questions about it. The student must answer each question by highlighting the exact phrase in the text — no paraphrasing, no guessing from outside knowledge, just pointing to the right words.
Before SQuAD, machine "exams" were either too easy (fill in one missing word) or too small (a few hundred questions). SQuAD created the first exam that was both large enough to train powerful models and hard enough to separate good readers from bad ones.
The gap: too small or too shallow
Reading comprehension is a foundational test of language understanding: read a passage, then answer questions about it. By 2016, two kinds of datasets existed, and neither was satisfactory:
Small, high-quality datasets like MCTest had only 2,640 questions — far too few to train the deep neural networks that were beginning to dominate NLP. Models would simply memorize the .
Large, synthetic datasets like CNN/Daily Mail contained 1.4 million examples but used a cloze format: a single entity was removed from a summary, and the model had to guess which one. This tested pattern matching more than comprehension — Chen et al. (2016) showed performance was nearly saturated with simple techniques.
The field needed a dataset that was both large enough for modern models and required genuine understanding: multi-word answers, paraphrased questions, and reasoning across sentences.
Building SQuAD: three-stage crowdsourcing pipeline
SQuAD was built in three careful stages, like constructing a high-quality exam from scratch:
Stage 1 — Passage curation. The authors selected 536 high-quality Wikipedia articles using PageRank scores from the top 10,000 English articles, then extracted paragraphs longer than 500 characters. This yielded 23,215 paragraphs covering everything from musical celebrities to abstract physics concepts. Articles were split 80/10/10 into train, development, and test sets.
Stage 2 — Question-answer collection. Crowdworkers on Amazon Mechanical Turk (via the Daemo platform) read each paragraph and wrote up to 5 questions, highlighting the answer span directly in the text. Workers were required to use their own words — copy-paste from the paragraph was disabled. This design choice was critical: it forced natural paraphrasing, creating the lexical and syntactic diversity that makes SQuAD challenging.
Stage 3 — Additional answers. To measure human performance and make evaluation robust, the authors collected at least 2 additional answers per question on the development and test sets. Only 2.6% of questions were marked unanswerable, confirming the quality of the original annotations.
Why SQuAD is hard: diverse answers and reasoning types
SQuAD's difficulty comes from two dimensions of diversity:
Answer type diversity. Unlike cloze datasets where answers are single entities, SQuAD answers span a wide range: dates (8.9%), numbers (10.9%), person names (12.9%), locations (4.4%), other entities (15.3%), common noun phrases (31.8%), adjective phrases (3.9%), verb phrases (5.5%), and full clauses (3.7%). Nearly half of all answers are not named entities — they require understanding meaning, not just recognizing entity types.
Reasoning type diversity. The authors manually analyzed 192 questions and found multiple reasoning types: lexical variation through synonymy (33.3%), world knowledge (9.1%), syntactic variation where the question's structure differs substantially from the answer sentence (64.1%), multi-sentence reasoning requiring coreference or fusion across sentences (13.6%), and ambiguous cases (6.1%).
The authors also introduced a measure of syntactic divergence — the edit distance between paths from the question to the answer. Think of it as measuring how much the question has been reworded relative to the passage. A question with zero divergence uses nearly the same structure as the passage; a question with high divergence (6+) is a complete paraphrase.
The key finding: the logistic regression model's performance degrades steadily as syntactic divergence increases, while human performance stays flat. This gap reveals that the model is matching surface patterns rather than truly understanding, and closing this gap became the central challenge for subsequent work.
How to score: Exact Match and F1
SQuAD introduced two evaluation metrics that became the gold standard for extractive :
Exact Match (EM) measures the percentage of predictions that match a ground truth answer exactly — after removing punctuation and articles (a, an, the). It is strict: one extra or missing word means a score of zero.
F1 Score treats the prediction and the ground truth as bags of tokens and computes their -level and . It is more forgiving: a prediction that captures most of the answer still gets partial credit. The maximum F1 over all ground truth answers is taken per question, then averaged.
Both metrics ignore articles and punctuation. The human performance baseline is 77.0% EM and 86.8% F1 — the gap between EM and F1 reflects how humans often include or exclude non-essential phrases (e.g., "monsoon trough" vs "movement of the monsoon trough").
The baseline: logistic regression with handcrafted features
To establish how hard SQuAD really is, the authors built a logistic regression model — think of it as the best a "traditional" NLP system could do before the revolution.
The model works in two stages. First, candidate generation: every constituent in the of the passage is a candidate answer (77.3% of correct answers are constituents — this sets an effective ceiling). Second, scoring: each candidate is scored using ~180 million features across several groups.
The most important features, as revealed by ablation, were lexicalized features (matching lemmas between the question and nearby words in the passage) and dependency tree path features (structural paths through the parse tree connecting anchor words to the answer span). Removing both of these dropped F1 from 51.0% to 35.8%.
The model achieved 40.4% EM and 51.0% F1 — dramatically better than the baseline (20% F1) but far below human performance (86.8% F1). Crucially, the model could select the correct sentence 79.3% of the time — the difficulty was in finding the exact span within that sentence.
Simplified to show the idea — not the real implementation.
import re, string
def normalize(text):
"""Remove articles, punctuation, and extra whitespace."""
text = text.lower()
text = re.sub(r'\b(a|an|the)\b', ' ', text)
text = ''.join(ch for ch in text if ch not in string.punctuation)
return ' '.join(text.split())
def f1_score(prediction, ground_truth):
"""Token-level F1 between a single prediction and ground truth."""
pred_tokens = normalize(prediction).split()
gold_tokens = normalize(ground_truth).split()
common = set(pred_tokens) & set(gold_tokens)
if len(common) == 0:
return 0.0
precision = len(common) / len(pred_tokens) # how many predicted tokens are correct?
recall = len(common) / len(gold_tokens) # how many gold tokens were predicted?
return 2 * precision * recall / (precision + recall)
# Example from the paper
pred = "gravity"
gold = "gravity"
print(f"F1: {f1_score(pred, gold):.1%}") # → F1: 100.0%
# Partial match — still gets credit
pred = "falls under gravity"
gold = "gravity"
print(f"F1: {f1_score(pred, gold):.1%}") # → F1: 50.0%The task: extractive question answering
SQuAD formalized a specific flavor of question answering called extractive QA. The task is simple to state: given a passage and a question , find the start index and end index of the answer span in .
Unlike multiple-choice QA (where you pick from 4 options) or abstractive QA (where you generate a free-form answer), extractive QA constrains the answer to be a contiguous subsequence of the passage. This constraint has a powerful advantage: evaluation is straightforward and unambiguous. There is no need for subjective judgment about whether a generated answer is "close enough."
The candidate space is large — for a passage of words, there are possible spans. The paper reduces this by only considering constituency parse constituents as candidates, but even then, each passage produces dozens of plausible candidates. The model must learn to distinguish the correct span from many wrong ones that may share significant lexical overlap with the question.
Results: a measurable gap between machines and humans
The results told a clear story. The sliding window baseline — which simply matched overlapping words — achieved just 20% F1. The logistic regression model with its 180 million handcrafted features reached 51.0% F1 and 40.4% EM. Human performance stood at 86.8% F1 and 77.0% EM.
The performance gap widened along two axes. By answer type: the model performed best on dates (72.1% F1) and numbers (62.5% F1), where candidates are few and easily identified by POS tags. It struggled most with verb phrases (31.2% F1) and clauses (34.3% F1), which require deeper understanding. By syntactic divergence: model F1 dropped from ~60% at zero divergence to ~35% at divergence 6+, while human F1 stayed flat at ~90% throughout.
ablation revealed a critical insight: removing both lexicalized and dependency path features dropped F1 from 51.0% to 35.8%, but the model still massively overfit the training data (91.7% F1 on train vs 51.0% on dev), suggesting that traditional features couldn't generalize the way neural representations would.
The legacy: how SQuAD changed NLP
2016
SQuAD v1.0 released
107,785 question-answer pairs on 536 Wikipedia articles. Logistic regression baseline: 51.0% F1. Human performance: 86.8% F1. The race begins.
2016
Match-LSTM & BiDAF
Neural attention models began closing the gap. Wang & Jiang's Match-LSTM reached 70.3% F1 within months. BiDAF introduced bidirectional attention flow.
2018
SQuAD 2.0 — unanswerable questions
Rajpurkar et al. added 50,000+ unanswerable questions to test whether models know when to say "I don't know." This addressed SQuAD v1's key limitation.
2018
BERT surpasses human performance
BERT achieved 93.2% F1 on SQuAD v1.1, surpassing the 86.8% human baseline. Pre-training on massive text corpora then fine-tuning on SQuAD proved devastatingly effective.
2019
SQuAD format adopted everywhere
Natural Questions, TriviaQA, HotpotQA, and the QA components of GLUE and SuperGLUE all adopted SQuAD's extractive format and EM/F1 metrics.
SQuAD's deepest contribution is not the dataset itself but the paradigm it established. It demonstrated that a well-designed crowdsourced benchmark, combined with a public , can channel the entire community's energy toward a single measurable goal. This "ImageNet moment for NLP" pattern has been replicated by every major benchmark since. The BERT paper lists SQuAD as one of its primary evaluation tasks, and SQuAD remains a standard benchmark to this day.
CitationRajpurkar, Zhang, Lopyrev, Liang. SQuAD: 100,000+ Questions for Machine Comprehension of Text. EMNLP, 2016.
Terms in this paper
- Reading Comprehensionالفهم القرائي
- Question Answeringالإجابة الحوسبية عن الأسئلة
- Extractive QAالإجابة الاستخلاصية
- Crowdsourcingالتعهيد الجماعي
- F1-Scoreمقياس إف 1 (المتوسط التوافقي للدقة والاستدعاء)
- Exact Matchالتطابق التام
- Spanمقطع نصّي
- Logistic Regressionالانحدار اللوجستي الاحتمالي
- Benchmarkالمعيار المرجعي
- Datasetمجموعة البيانات