Information Retrieval2020intermediate10 min read

Dense Passage Retrieval for Open-Domain Question Answering

الاسترجاع الكثيف للفقرات للإجابة عن الأسئلة في النطاق المفتوح

Karpukhin, V. · Oğuz, B. · Min, S. · Lewis, P. · Wu, L. · Edunov, S. · Chen, D. · Yih, W. — EMNLP

The problem

needs to find relevant passages from millions of documents. By 2020, the standard approach was — a sparse keyword-matching algorithm dating to the 1990s. BM25 counts word overlaps: if your question says "climate" but the passage says "weather", BM25 misses it. Prior attempts at were complex, slow, or underperformed BM25.

The contribution

DPR: a dual- framework using two independent BERT models — one encodes the question, one encodes the passage. Both are mapped to the same 768-dimensional , and similarity is a simple . Training uses plus one BM25 hard negative per question — surprisingly effective with only ~59K question-passage pairs. indexes all 21M Wikipedia passages offline, enabling sub-second retrieval. DPR achieved 78.4% top-20 accuracy on Natural Questions — 9–19% above BM25 — and when paired with a reader, set a new state-of-the-art on multiple open-domain QA benchmarks.

The impact

DPR proved that dense retrieval can decisively beat sparse methods with a simple architecture. It became the retrieval backbone of RAG (Retrieval-Augmented Generation), inspired ColBERT's late interaction approach, led to Contriever's unsupervised pre-training, and powered FiD's multi-passage reading. Every modern retrieval system — from search engines to AI assistants — traces a lineage back to DPR's dual-encoder blueprint.

BM25 is a card-catalogue search: you look up your exact keywords and find documents whose index cards contain the same words. If you ask about "solar energy" but the best article is filed under "photovoltaic power", you walk out empty-handed.

DPR replaces the card catalogue with a mind-reading librarian. She reads your question, understands what you mean, walks into the stacks, and pulls out the passages whose meaning is closest — even if they share zero words with your question.

The problem: keyword matching fails when meaning matters

Open-domain question answering has two stages: a retriever finds candidate passages from a massive (like all of Wikipedia), then a reader extracts the answer from those passages. The retriever is the bottleneck — if it misses the right passage, the reader never gets a chance.

By 2020, BM25 was the universal retriever. It works by counting how many question words appear in a passage, weighting rare words higher. This sparse approach is fast and surprisingly good for exact keyword matches, but it has a fundamental blind spot: synonymy. The question "Who is the head of government in Germany?" and a passage containing "The Chancellor of the Federal Republic" share almost no words, yet they answer the same query.

Prior dense retrieval attempts used complex pipelines — pre-training with Inverse Cloze Tasks, multi-step training, or latent variable models. They added significant complexity but still couldn't consistently beat BM25. The field was stuck: everyone knew semantic matching should help, but nobody had made it work simply enough to be practical.

Open in Lab
Type a question and compare BM25 keyword matching with DPR's semantic matching.
The demo wakes as you arrive…

The idea: two encoders, one shared space

DPR's insight is radical simplicity. Take two copies of BERT, one for questions and one for passages. Each encoder reads its input and compresses it into a single 768-dimensional — the [CLS] token's final hidden state. The question and passage now live in the same vector space. Similarity is just a dot product: if two vectors point in the same direction, their dot product is high.

Think of it as giving the question and each passage a GPS coordinate in meaning space. Close coordinates = similar meaning. The question "What year was penicillin discovered?" and the passage "Alexander Fleming discovered penicillin in 1928" should land near each other, even though they share only one word.

This architecture is called a (or ): two independent encoders that never see each other's input, connected only by the shared vector space they project into. This independence is what makes DPR fast — all passage vectors can be computed offline.

sim(q,p)=EQ(q)⊤⋅EP(p)\text{sim}(q, p) = E_Q(q)^\top \cdot E_P(p)
Similarity score — the full retrieval engine in one line — E_Q(q) = question encoder (BERT + [CLS] pooling) · E_P(p) = passage encoder (separate BERT + [CLS] pooling) · The dot product measures how aligned the two vectors are in meaning space
Open in Lab
Click each component to see how the question and passage flow through their independent BERT encoders.
The demo wakes as you arrive…

Training: the art of choosing negatives

DPR trains with : push question-passage pairs with correct answers closer in vector space, and push irrelevant pairs apart. The is the negative log-likelihood of the positive passage — essentially asking: "Among all passages in this batch, can you identify the right one?"

The critical insight is which negatives to train against. DPR tested three types:

  • Random negatives: any random passage from the corpus. Easy for the model — too easy. The model barely learns.
  • BM25 : passages that BM25 ranks highly but don't actually contain the answer. These are tricky because they share keywords with the question. One BM25 hard negative per question dramatically improved results.
  • In-batch negatives: the positive passages of other questions in the same training batch become negatives for this question. With a batch size of 128, each question gets 127 free negatives — no extra computation needed.

The winning recipe was in-batch negatives + 1 BM25 hard negative. Gold (annotated) negatives actually hurt performance because human annotators select passages that are close to correct, creating confusing training signals.

L(qi,pi+,pi,1−,…,pi,n−)=−log⁡esim(qi,pi+)esim(qi,pi+)+∑j=1nesim(qi,pi,j−)L(q_i, p_i^+, p_{i,1}^-, \ldots, p_{i,n}^-) = -\log \frac{e^{\text{sim}(q_i, p_i^+)}}{e^{\text{sim}(q_i, p_i^+)} + \sum_{j=1}^{n} e^{\text{sim}(q_i, p_{i,j}^-)}}
Negative log-likelihood loss — the training signal — The numerator is the similarity of the correct pair · the denominator adds all negative pairs · minimizing this pushes the positive score up and all negative scores down
Open in Lab
Watch how in-batch negatives work: each question's positive passage becomes a negative for every other question.
The demo wakes as you arrive…

Retrieval at scale: offline indexing with FAISS

The dual-encoder architecture has a crucial property: the passage encoder and the question encoder are independent. This means all 21 million Wikipedia passages can be encoded once and stored. At query time, only the question needs encoding — a single through a small BERT.

But searching 21 million 768-dimensional vectors by brute force is slow. This is where FAISS comes in. FAISS (Facebook AI Similarity Search) builds a hierarchical index that supports search. Instead of comparing against all 21M vectors, FAISS partitions the space and searches only the relevant clusters.

The pipeline at looks like this: (1) encode the question into a 768-d vector, (2) query FAISS to retrieve the top-k most similar passage vectors, (3) pass those passages to a reader model that extracts the answer span. The entire retrieval step takes milliseconds.

This offline-index paradigm is the reason DPR became the backbone of production retrieval systems. You pay the encoding cost once, then serve millions of queries against the same index. Updating the index only requires re-encoding changed passages.

Open in Lab
Follow a question through the full DPR pipeline — from encoding to FAISS search to answer extraction.
The demo wakes as you arrive…

Inside the vector space: where questions meet passages

The magic of DPR is what happens in the shared . Before training, BERT's [CLS] vectors for questions and passages are scattered randomly — a question about penicillin might be far from a passage about Fleming. After training, the contrastive loss has reshaped the space: question vectors cluster near their answer-bearing passages.

This is fundamentally different from what Sentence-BERT does. Sentence-BERT creates a symmetric space where similar sentences are close. DPR creates an asymmetric space where questions are close to their answers, not to similar questions. "Who discovered penicillin?" is far from "Who invented the telephone?" in DPR's space, even though they're structurally similar — because their answer passages are in completely different regions.

The dot product as a similarity function is a deliberate choice. Unlike , it allows vectors of different magnitudes to express different confidence levels. A passage encoder that produces a longer vector is "louder" — it will be retrieved more often. This gives the model an extra degree of freedom to express how strongly a passage relates to a topic.

Open in Lab
Drag questions and passages to see how training reshapes the vector space — before and after DPR fine-tuning.
The demo wakes as you arrive…

Ablation: why negative selection is everything

The paper's ablation study revealed a surprising hierarchy of negative types. Using only random negatives gave ~54% top-20 accuracy on Natural Questions. Adding in-batch negatives jumped to ~68%. Adding a single BM25 hard negative per question pushed it to ~78%.

Why does one hard negative matter so much? Because it forces the model to solve the hard discrimination problem. Random negatives are obviously wrong — a passage about cooking won't confuse a model asking about physics. BM25 negatives share surface-level keywords with the question, so the model must learn to look beyond words at actual meaning. It's like studying for an exam with trick questions: they teach you more than easy ones.

Interestingly, gold (human-annotated) negatives hurt performance. The authors hypothesize that gold negatives are too close to positives — they might be alternative correct answers or near-misses that confuse the training signal.

Open in Lab
Compare the three negative strategies and see how each impacts retrieval accuracy.
The demo wakes as you arrive…

The same idea in code

DPR dual-encoder retrieval, completepython

Simplified to show the idea — not the real implementation.

import numpy as np

def encode_question(question_text, bert_q):
    """Encode question -> 768-d vector via question-BERT's [CLS]."""
    tokens = tokenize("[CLS] " + question_text + " [SEP]")
    hidden = bert_q(tokens)          # (seq_len, 768)
    return hidden[0]                 # [CLS] vector -> (768,)

def encode_passage(passage_text, bert_p):
    """Encode passage -> 768-d vector via passage-BERT's [CLS]."""
    tokens = tokenize("[CLS] " + passage_text + " [SEP]")
    hidden = bert_p(tokens)
    return hidden[0]                 # [CLS] vector -> (768,)

def similarity(q_vec, p_vec):
    """Dot product -- the only distance metric DPR needs."""
    return q_vec @ p_vec             # scalar

def nll_loss(q_vec, pos_vec, neg_vecs):
    """Negative log-likelihood: push positives up, negatives down."""
    pos_score = similarity(q_vec, pos_vec)
    all_scores = [pos_score] + [similarity(q_vec, n) for n in neg_vecs]
    log_sum_exp = np.log(sum(np.exp(s) for s in all_scores))
    return log_sum_exp - pos_score   # minimize this

def retrieve_top_k(question, faiss_index, bert_q, k=20):
    """At inference: encode question, search FAISS, return passages."""
    q_vec = encode_question(question, bert_q)
    scores, indices = faiss_index.search(q_vec.reshape(1, -1), k)
    return indices[0], scores[0]     # top-k passage IDs & scores

# That's the full DPR retriever. Two BERTs, a dot product, and FAISS.
# The reader (answer extractor) is a separate model on top.

Why it mattered

  1. 2019

    Sentence-BERT

    Siamese BERT networks for sentence similarity. Created efficient symmetric embeddings but wasn't optimized for the asymmetric question-passage retrieval task.

  2. 2020

    DPR — Dense Passage Retrieval

    Proved dense retrieval beats BM25 with a simple dual-encoder framework. In-batch negatives + BM25 hard negatives = 78.4% top-20 accuracy.

  3. 2020

    RAG — Retrieval-Augmented Generation

    Used DPR as the retrieval backbone inside a generative model. Instead of extracting answer spans, RAG generates free-form answers conditioned on retrieved passages.

  4. 2020

    ColBERT — Late Interaction

    Kept DPR's independent encoding but preserved token-level embeddings instead of collapsing to one [CLS] vector. Late interaction gives finer-grained matching at modest extra cost.

  5. 2021

    FiD — Fusion-in-Decoder

    Retrieved 100 passages with DPR, encoded each independently, then fused them in the decoder. Showed more passages = better answers, validating DPR's recall strength.

  6. 2022

    Contriever — Unsupervised Dense Retrieval

    Extended DPR's idea with unsupervised pre-training via contrastive learning on unlabeled data. No question-passage pairs needed for pre-training, only fine-tuning.

DPR's deepest legacy is the paradigm shift it proved: you don't need hand-crafted features or complex pipelines to beat decades-old systems. Two pre-trained language models, a dot product, and smart negative sampling are enough. Every retrieval system in modern AI — from RAG pipelines to enterprise search to the retrieval layer inside AI assistants — inherits this blueprint.

CitationKarpukhin, Oguz, Min, Lewis, Wu, Edunov, Chen, Yih. Dense Passage Retrieval for Open-Domain Question Answering. EMNLP, 2020.

Terms in this paper