Information Retrieval2020intermediate10 min read
Dense Passage Retrieval for Open-Domain Question Answering
الاسترجاع الكثيف للفقرات للإجابة عن الأسئلة في النطاق المفتوح
Karpukhin, V. · Oğuz, B. · Min, S. · Lewis, P. · Wu, L. · Edunov, S. · Chen, D. · Yih, W. — EMNLP
The problem
needs to find relevant passages from millions of documents. By 2020, the standard approach was — a sparse keyword-matching algorithm dating to the 1990s. BM25 counts word overlaps: if your question says "climate" but the passage says "weather", BM25 misses it. Prior attempts at were complex, slow, or underperformed BM25.
The contribution
DPR: a dual- framework using two independent BERT models — one encodes the question, one encodes the passage. Both are mapped to the same 768-dimensional , and similarity is a simple . Training uses plus one BM25 hard negative per question — surprisingly effective with only ~59K question-passage pairs. indexes all 21M Wikipedia passages offline, enabling sub-second retrieval. DPR achieved 78.4% top-20 accuracy on Natural Questions — 9–19% above BM25 — and when paired with a reader, set a new state-of-the-art on multiple open-domain QA benchmarks.
The impact
DPR proved that dense retrieval can decisively beat sparse methods with a simple architecture. It became the retrieval backbone of RAG (Retrieval-Augmented Generation), inspired ColBERT's late interaction approach, led to Contriever's unsupervised pre-training, and powered FiD's multi-passage reading. Every modern retrieval system — from search engines to AI assistants — traces a lineage back to DPR's dual-encoder blueprint.
BM25 is a card-catalogue search: you look up your exact keywords and find documents whose index cards contain the same words. If you ask about "solar energy" but the best article is filed under "photovoltaic power", you walk out empty-handed.
DPR replaces the card catalogue with a mind-reading librarian. She reads your question, understands what you mean, walks into the stacks, and pulls out the passages whose meaning is closest — even if they share zero words with your question.
The problem: keyword matching fails when meaning matters
Open-domain question answering has two stages: a retriever finds candidate passages from a massive (like all of Wikipedia), then a reader extracts the answer from those passages. The retriever is the bottleneck — if it misses the right passage, the reader never gets a chance.
By 2020, BM25 was the universal retriever. It works by counting how many question words appear in a passage, weighting rare words higher. This sparse approach is fast and surprisingly good for exact keyword matches, but it has a fundamental blind spot: synonymy. The question "Who is the head of government in Germany?" and a passage containing "The Chancellor of the Federal Republic" share almost no words, yet they answer the same query.
Prior dense retrieval attempts used complex pipelines — pre-training with Inverse Cloze Tasks, multi-step training, or latent variable models. They added significant complexity but still couldn't consistently beat BM25. The field was stuck: everyone knew semantic matching should help, but nobody had made it work simply enough to be practical.
The idea: two encoders, one shared space
DPR's insight is radical simplicity. Take two copies of BERT, one for questions and one for passages. Each encoder reads its input and compresses it into a single 768-dimensional — the [CLS] token's final hidden state. The question and passage now live in the same vector space. Similarity is just a dot product: if two vectors point in the same direction, their dot product is high.
Think of it as giving the question and each passage a GPS coordinate in meaning space. Close coordinates = similar meaning. The question "What year was penicillin discovered?" and the passage "Alexander Fleming discovered penicillin in 1928" should land near each other, even though they share only one word.
This architecture is called a (or ): two independent encoders that never see each other's input, connected only by the shared vector space they project into. This independence is what makes DPR fast — all passage vectors can be computed offline.
Training: the art of choosing negatives
DPR trains with : push question-passage pairs with correct answers closer in vector space, and push irrelevant pairs apart. The is the negative log-likelihood of the positive passage — essentially asking: "Among all passages in this batch, can you identify the right one?"
The critical insight is which negatives to train against. DPR tested three types:
- Random negatives: any random passage from the corpus. Easy for the model — too easy. The model barely learns.
- BM25 : passages that BM25 ranks highly but don't actually contain the answer. These are tricky because they share keywords with the question. One BM25 hard negative per question dramatically improved results.
- In-batch negatives: the positive passages of other questions in the same training batch become negatives for this question. With a batch size of 128, each question gets 127 free negatives — no extra computation needed.
The winning recipe was in-batch negatives + 1 BM25 hard negative. Gold (annotated) negatives actually hurt performance because human annotators select passages that are close to correct, creating confusing training signals.
Retrieval at scale: offline indexing with FAISS
The dual-encoder architecture has a crucial property: the passage encoder and the question encoder are independent. This means all 21 million Wikipedia passages can be encoded once and stored. At query time, only the question needs encoding — a single through a small BERT.
But searching 21 million 768-dimensional vectors by brute force is slow. This is where FAISS comes in. FAISS (Facebook AI Similarity Search) builds a hierarchical index that supports search. Instead of comparing against all 21M vectors, FAISS partitions the space and searches only the relevant clusters.
The pipeline at looks like this: (1) encode the question into a 768-d vector, (2) query FAISS to retrieve the top-k most similar passage vectors, (3) pass those passages to a reader model that extracts the answer span. The entire retrieval step takes milliseconds.
This offline-index paradigm is the reason DPR became the backbone of production retrieval systems. You pay the encoding cost once, then serve millions of queries against the same index. Updating the index only requires re-encoding changed passages.
Inside the vector space: where questions meet passages
The magic of DPR is what happens in the shared . Before training, BERT's [CLS] vectors for questions and passages are scattered randomly — a question about penicillin might be far from a passage about Fleming. After training, the contrastive loss has reshaped the space: question vectors cluster near their answer-bearing passages.
This is fundamentally different from what Sentence-BERT does. Sentence-BERT creates a symmetric space where similar sentences are close. DPR creates an asymmetric space where questions are close to their answers, not to similar questions. "Who discovered penicillin?" is far from "Who invented the telephone?" in DPR's space, even though they're structurally similar — because their answer passages are in completely different regions.
The dot product as a similarity function is a deliberate choice. Unlike , it allows vectors of different magnitudes to express different confidence levels. A passage encoder that produces a longer vector is "louder" — it will be retrieved more often. This gives the model an extra degree of freedom to express how strongly a passage relates to a topic.
Ablation: why negative selection is everything
The paper's ablation study revealed a surprising hierarchy of negative types. Using only random negatives gave ~54% top-20 accuracy on Natural Questions. Adding in-batch negatives jumped to ~68%. Adding a single BM25 hard negative per question pushed it to ~78%.
Why does one hard negative matter so much? Because it forces the model to solve the hard discrimination problem. Random negatives are obviously wrong — a passage about cooking won't confuse a model asking about physics. BM25 negatives share surface-level keywords with the question, so the model must learn to look beyond words at actual meaning. It's like studying for an exam with trick questions: they teach you more than easy ones.
Interestingly, gold (human-annotated) negatives hurt performance. The authors hypothesize that gold negatives are too close to positives — they might be alternative correct answers or near-misses that confuse the training signal.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def encode_question(question_text, bert_q):
"""Encode question -> 768-d vector via question-BERT's [CLS]."""
tokens = tokenize("[CLS] " + question_text + " [SEP]")
hidden = bert_q(tokens) # (seq_len, 768)
return hidden[0] # [CLS] vector -> (768,)
def encode_passage(passage_text, bert_p):
"""Encode passage -> 768-d vector via passage-BERT's [CLS]."""
tokens = tokenize("[CLS] " + passage_text + " [SEP]")
hidden = bert_p(tokens)
return hidden[0] # [CLS] vector -> (768,)
def similarity(q_vec, p_vec):
"""Dot product -- the only distance metric DPR needs."""
return q_vec @ p_vec # scalar
def nll_loss(q_vec, pos_vec, neg_vecs):
"""Negative log-likelihood: push positives up, negatives down."""
pos_score = similarity(q_vec, pos_vec)
all_scores = [pos_score] + [similarity(q_vec, n) for n in neg_vecs]
log_sum_exp = np.log(sum(np.exp(s) for s in all_scores))
return log_sum_exp - pos_score # minimize this
def retrieve_top_k(question, faiss_index, bert_q, k=20):
"""At inference: encode question, search FAISS, return passages."""
q_vec = encode_question(question, bert_q)
scores, indices = faiss_index.search(q_vec.reshape(1, -1), k)
return indices[0], scores[0] # top-k passage IDs & scores
# That's the full DPR retriever. Two BERTs, a dot product, and FAISS.
# The reader (answer extractor) is a separate model on top.Why it mattered
2019
Sentence-BERT
Siamese BERT networks for sentence similarity. Created efficient symmetric embeddings but wasn't optimized for the asymmetric question-passage retrieval task.
2020
DPR — Dense Passage Retrieval
Proved dense retrieval beats BM25 with a simple dual-encoder framework. In-batch negatives + BM25 hard negatives = 78.4% top-20 accuracy.
2020
RAG — Retrieval-Augmented Generation
Used DPR as the retrieval backbone inside a generative model. Instead of extracting answer spans, RAG generates free-form answers conditioned on retrieved passages.
2020
ColBERT — Late Interaction
Kept DPR's independent encoding but preserved token-level embeddings instead of collapsing to one [CLS] vector. Late interaction gives finer-grained matching at modest extra cost.
2021
FiD — Fusion-in-Decoder
Retrieved 100 passages with DPR, encoded each independently, then fused them in the decoder. Showed more passages = better answers, validating DPR's recall strength.
2022
Contriever — Unsupervised Dense Retrieval
Extended DPR's idea with unsupervised pre-training via contrastive learning on unlabeled data. No question-passage pairs needed for pre-training, only fine-tuning.
DPR's deepest legacy is the paradigm shift it proved: you don't need hand-crafted features or complex pipelines to beat decades-old systems. Two pre-trained language models, a dot product, and smart negative sampling are enough. Every retrieval system in modern AI — from RAG pipelines to enterprise search to the retrieval layer inside AI assistants — inherits this blueprint.
CitationKarpukhin, Oguz, Min, Lewis, Wu, Edunov, Chen, Yih. Dense Passage Retrieval for Open-Domain Question Answering. EMNLP, 2020.
Terms in this paper
- Dense Retrievalالاسترجاع الكثيف
- Dual Encoderالمُرمِّز الثنائي
- Passage Retrievalاسترجاع الفقرات
- BM25BM25
- FAISSفايس (FAISS)
- In-Batch Negativesسلبيات داخل الدُّفعة
- Hard Negativesالسلبيات الصعبة
- Sparse Retrievalالاسترجاع المتناثر
- Approximate Nearest Neighborالجار الأقرب التقريبي
- Open-Domain Question Answeringالإجابة المفتوحة عن الأسئلة