Language Models2020intermediate10 min read
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
التوليد المعزَّز بالاسترجاع للمهام اللغوية كثيفة المعرفة
Lewis, P. · Perez, E. · Piktus, A. · Petroni, F. · Karpukhin, V. · Goyal, N. · Küttler, H. · Lewis, M. · Yih, W. · Rocktäschel, T. · Riedel, S. · Kiela, D. — NeurIPS
The problem
Large pre-trained language models store knowledge implicitly in their parameters, but this knowledge is fixed after , hard to update, and the model can't cite where facts come from. On tasks — open-domain QA, , knowledge-grounded generation — purely parametric models hallucinate, can't access fresh information, and expanding their knowledge requires full retraining.
The contribution
RAG: a hybrid architecture that combines a pre-trained parametric generator (BART) with a dense retriever (DPR). Given a query, DPR retrieves the top-k documents from a Wikipedia index; these documents are treated as latent variables and marginalized out during generation. Two variants — RAG-Sequence (one document per entire output) and RAG- (different documents per token) — set state-of-the-art on open-domain QA, fact verification, and Jeopardy question generation, while providing interpretable provenance through the retrieved passages.
The impact
RAG established the retrieval-augmented paradigm that now underpins virtually every production LLM system — from chatbots citing sources to enterprise search to AI agents consulting knowledge bases. Its insight that generation models should consult external evidence rather than rely solely on memorized parameters launched an entire research direction including FiD, RETRO, Atlas, and Self-RAG, and made "RAG" a standard term in both research and industry.
A purely parametric model is like a student taking a closed-book exam: everything it knows must be memorized inside its head (parameters). It may sound confident, but when pressed on specifics it sometimes invents facts — hallucinating answers it was never taught.
RAG turns the exam into open-book: the student (generator) keeps a massive library (Wikipedia index) on the desk. When a question arrives, a research assistant (the retriever) flips through the library, pulls the most relevant pages, and places them in front of the student — who then writes a polished answer grounded in actual evidence.
The problem: memorized knowledge is fragile and opaque
By 2020, large language models like GPT-2 and BART had shown impressive generation abilities, but they stored all knowledge in their parameters — a purely parametric memory. This design has three fundamental problems:
- Staleness. Knowledge is frozen at training time. If a fact changes tomorrow, the model is wrong until retrained — and retraining costs millions of dollars.
- . When the model lacks knowledge, it generates plausible-sounding but fabricated answers. There is no mechanism to say "I don't know."
- Opacity. You cannot ask the model where it learned a fact. There is no provenance, no citation, no way to verify.
For knowledge-intensive tasks — where getting the facts right is the whole point — these problems are deal-breakers.
The idea: retrieve then generate
RAG's insight is deceptively simple: don't make the generator memorize everything — let it look things up. Combine two systems:
- A retriever (DPR) that maps queries and documents into a shared . Given a question, it finds the top-k most relevant passages from a huge (all of Wikipedia, ~21 million passages) using Maximum Inner Product Search (MIPS).
- A generator (BART), a pre-trained seq2seq model that takes the question plus each retrieved passage as input and generates an answer.
The retrieved documents are treated as latent variables — hidden evidence the model consults but doesn't directly output. The final answer marginalizes (averages) over these documents, so the model considers multiple sources before committing to an answer.
The architecture: DPR retriever + BART generator
The RAG architecture has two components that work together like a research team:
The Retriever (DPR) uses two BERT encoders — one for the query, one for documents. Both produce dense representations. Retrieval is a nearest-neighbor search in this shared vector space: the query is compared against a pre-computed FAISS index of all Wikipedia passages using Maximum Inner Product Search (MIPS). The top-k passages (typically k=5 or 10) with the highest similarity are returned.
The Generator (BART) is a -decoder. For each retrieved passage, BART receives the concatenation of the original question and the passage as input, then generates the answer autoregressively. The crucial insight: BART sees each document independently — it generates a separate probability distribution over answers for each document, then these are combined via .
Two variants: RAG-Sequence vs. RAG-Token
RAG models documents as latent variables: the model retrieves them but doesn't output them directly. The key question is when to marginalize (average) over documents. This gives two variants:
RAG-Sequence uses the same single document to generate the entire output sequence. Think of it as: pick one reference book, write your whole essay from it, then average the results across different book choices. This is ideal for tasks where coherence matters — like writing a complete, consistent answer.
RAG-Token can use a different document for each output token. Think of it as: for every word you write, you can flip to a different page in a different book. This gives finer-grained control and is ideal for synthesis tasks where different parts of the answer draw on different facts.
Training: the retriever and generator learn together
RAG is trained on pairs of (question, answer) — no document labels needed. The training signal flows like this:
- The retriever finds documents it thinks are relevant.
- The generator tries to produce the correct answer using those documents.
- If the answer is wrong, the updates both the generator (to better use documents) and the retriever's query encoder (to fetch better documents next time).
Crucially, the document encoder is frozen — only the query encoder is updated. This avoids the need to re-embed and re-index all 21 million Wikipedia passages after every gradient step. The document index is pre-built once using FAISS and stays fixed during training.
The is simply the negative log-likelihood of the correct answer, marginalized over retrieved documents. No retrieval supervision is needed — the model learns which documents are useful purely from the end task signal.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def rag_forward(question, retriever, generator, index, k=5):
"""RAG forward pass: retrieve top-k documents, then generate."""
# Step 1: Encode the question into a dense vector
q_emb = retriever.encode_query(question) # (d,)
# Step 2: Retrieve top-k passages via MIPS (Maximum Inner Product Search)
doc_scores, doc_ids = index.search(q_emb, k) # FAISS index
documents = [index.get_passage(did) for did in doc_ids]
# Step 3: For each document, generate answer probabilities
all_probs = []
for doc, score in zip(documents, doc_scores):
input_text = f"{question} [SEP] {doc}" # concatenate q + doc
gen_prob = generator.generate_probs(input_text) # p(y | x, z)
retrieval_prob = np.exp(score) # p(z | x)
all_probs.append(retrieval_prob * gen_prob)
# Step 4: Marginalize — average over documents
# RAG-Token: sum at each position; RAG-Sequence: sum full sequences
final_probs = sum(all_probs) / sum(np.exp(doc_scores))
return final_probs
# Key insight: the retrieved documents are LATENT VARIABLES.
# The model marginalizes over them — it doesn't output them directly.
# Training: gradient flows through generator + query encoder.
# Document encoder stays frozen. No re-indexing needed.Results: knowledge without memorization
RAG set new state-of-the-art on four knowledge-intensive benchmarks:
- Open-Domain QA (Natural Questions): RAG achieved 44.5 Exact Match, outperforming all purely extractive models and previous generative approaches. Unlike extractive models that can only copy spans from retrieved passages, RAG synthesizes answers.
- Open-Domain QA (CuratedTrec & WebQuestions): RAG set new records on both benchmarks without task-specific architecture modifications.
- Abstractive QA (MSMARCO NLG): RAG achieved a BLEU score of 15.1, showing it can generate more natural language answers than purely extractive systems.
- Jeopardy Question Generation: RAG generated more factual and specific questions than BART alone, demonstrating that retrieval grounds even creative generation tasks.
- Fact Verification (FEVER): Within 4.3% of state-of-the-art models that use much more complex pipelines with task-specific components.
The hidden advantage: updatable knowledge
Perhaps RAG's most profound advantage isn't on any benchmark — it's the ability to update knowledge without retraining. Because the knowledge lives in the external index rather than in the model's parameters, you can:
- Swap the index to change the knowledge domain. Point RAG at medical papers instead of Wikipedia and it becomes a medical QA system — same model, different knowledge.
- Update facts by re-indexing new documents. When a world leader changes, you update the index, not the model.
- Trace provenance by examining which passages were retrieved. Every answer comes with receipts — you can verify claims against the source documents.
This separation of knowledge from computation is what makes RAG practical for production systems where facts change and trust matters.
What RAG unlocked
2019
REALM — Retrieval-Enhanced Language Model
First to treat documents as latent variables in language model pre-training. REALM showed retrieval could improve pre-training itself, not just downstream tasks.
2020
DPR — Dense Passage Retrieval
Replaced sparse (BM25) retrieval with learned dense representations. DPR became RAG's retriever backbone and proved that neural retrievers outperform keyword matching.
2020
RAG — Retrieval-Augmented Generation
Combined DPR + BART into a single trainable system. Documents as latent variables, marginalized during generation. State-of-the-art on knowledge-intensive tasks.
2021
FiD — Fusion-in-Decoder
Encoded each document separately but fused them in the decoder's cross-attention. Scaled to 100+ documents efficiently by avoiding encoder-side cross-attention.
2022
RETRO — Retrieval-Enhanced Transformers
Integrated retrieval into the Transformer architecture itself with chunked cross-attention. Showed retrieval augmentation benefits even models with billions of parameters.
2023
Atlas — Few-Shot Retrieval-Augmented LM
Extended RAG with jointly trained retriever, achieving strong few-shot performance by learning to retrieve task-relevant documents with minimal examples.
2023
Self-RAG — Self-Reflective RAG
Trained the model to decide *when* to retrieve and to critique its own outputs with reflection tokens. Adaptive retrieval — only search when the model's own knowledge is insufficient.
Today, the retrieval-augmented paradigm RAG introduced has become the industry standard. Every major LLM system — from ChatGPT with browsing to Claude with search to enterprise knowledge assistants — uses some form of retrieve-then-generate. The term "RAG" itself has entered the everyday vocabulary of AI practitioners, engineers, and product managers worldwide.
CitationLewis, Perez, Piktus, Petroni, Karpukhin, Goyal, Küttler, Lewis, Yih, Rocktäschel, Riedel, Kiela. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS, 2020.
Terms in this paper
- Retrieval-Augmented Generationالتوليد المعزَّز بالاسترجاع
- Dense Passage Retrievalالاسترجاع الكثيف للفقرات
- Non-Parametricلا مَعلَمي
- Marginalizationالتهميش
- Latent Variableالمتغير الكامن
- Knowledge-Intensiveكثيف المعرفة
- Open-Domain Question Answeringالإجابة المفتوحة عن الأسئلة
- Abstractive Question Answeringالإجابة التجريدية عن الأسئلة
- Fact Verificationالتحقّق من الحقائق
- Hallucinationالهلوسة الرقمية
- Groundingالتأسيس