NLP2021intermediate11 min read

Fusion-in-Decoder: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering

الدمج في المفكِّك: الاستفادة من استرجاع المقاطع مع النماذج التوليدية للإجابة عن الأسئلة المفتوحة

Izacard, G. · Grave, E. — EACL

The problem

By 2020, large generative models like T5-11B had shown they could answer open-domain questions by storing facts in their parameters alone — but this required billions of parameters and was expensive to train and query. On the other hand, retrieval-based extractive models could pull in external evidence, but they struggled to aggregate information scattered across multiple passages. Extractive readers predict a single span from one passage and have no natural mechanism to combine clues from different documents. The question was whether generative models could efficiently leverage retrieved passages to match or surpass both approaches — with far fewer parameters.

The contribution

Fusion-in- (FiD): a method that encodes each retrieved passage independently with the question using a T5 , then concatenates all encoded representations and passes them to the decoder. The decoder attends over all passages jointly via to generate the answer. This simple architecture achieves state-of-the-art on NaturalQuestions (51.4 EM) and TriviaQA (67.6 EM), and scales remarkably — performance keeps improving from 10 to 100 passages, unlike extractive models that plateau around 10–20.

The impact

FiD established the retriever–generative-reader paradigm that became the backbone of modern . It proved that the decoder's cross-attention is a powerful evidence fusion mechanism, more effective than extractive span selection when evidence is distributed. The architecture directly influenced Atlas, RETRO, and a generation of knowledge-intensive NLP systems. Its insight that encoder cost scales linearly while the decoder fuses information made processing 100+ passages practical.

Imagine a detective investigating a crime. The old approach () was like reading one witness statement and highlighting the single most suspicious sentence. If the key clue was split across three different witnesses, you would miss it.

Fusion-in-Decoder is like interviewing each witness in a separate room, writing a summary of each interview, then spreading all summaries on one big table and reading across them to write the final report. Each witness is heard independently — no crosstalk, no confusion — but the final conclusion draws on every interview at once.

The result: a detective that gets sharper with every additional witness, instead of getting overwhelmed.

The problem: answering questions when no single passage has the full answer

is the task of answering factual questions using a large corpus like Wikipedia, without knowing in advance which document contains the answer. By 2020, two families of approaches had emerged.

The first family, closed-book generative models, memorized facts directly into model parameters during pretraining. T5-11B with 11 billion parameters could answer questions without any retrieval — but storing all of Wikipedia in weights is expensive. The second family, retriever–extractive-reader pipelines, first retrieved relevant passages using sparse () or dense (DPR) retrieval, then extracted an answer span from a single passage using a BERT-based reader.

The extractive approach had a critical blind spot: if the answer required combining information from multiple passages — for example, a birth date from one document and a location name from another — extractive models had no natural way to merge these clues. Techniques like global span normalization helped, but were fundamentally limited by the single-span prediction bottleneck.

Open in Lab
See how an extractive reader is limited to one passage, while a generative reader can fuse evidence from multiple passages to produce the correct answer.
The demo wakes as you arrive…

Step 1: Retrieving passages — BM25 vs Dense Passage Retrieval

Before the model can answer, it needs to find the right passages. FiD uses two retrieval strategies depending on the dataset.

BM25 is a classic sparse retrieval method. It represents both questions and passages as bags of words and scores them using term frequency and inverse document frequency. It is fast, requires no , and works well when the question contains the exact words that appear in the answer passage. FiD uses BM25 for SQuAD Open.

(DPR) encodes questions and passages as dense vectors using two separate BERT encoders. Retrieval is performed by finding the passages whose vectors have the highest dot product with the question vector, using search via FAISS. DPR captures semantic similarity — it can match "birthplace" in the question to "born in" in the passage — and is used for NaturalQuestions and TriviaQA.

In both cases, the retriever returns the top 100 passages from Wikipedia (split into non-overlapping chunks of 100 words each). These passages, along with their titles, are then passed to the reader model.

Open in Lab
Compare how BM25 (keyword matching) and DPR (semantic matching) retrieve passages for the same question. Toggle between them to see the difference.
The demo wakes as you arrive…

Step 2: The Fusion-in-Decoder architecture

The core insight of Fusion-in-Decoder is a division of labor between the encoder and the decoder. Think of it as a two-phase process: independent reading followed by joint reasoning.

Phase 1 — Independent encoding. Each retrieved passage is concatenated with the question using special tokens: question: [q] title: [t] context: [p]. Each of these question–passage pairs is processed independently by the T5 encoder. This is the key efficiency trick: because each passage is encoded separately, the encoder performs only within a single passage at a time. The computational cost grows linearly with the number of passages, not quadratically as it would if all passages were concatenated into one long sequence.

Phase 2 — Joint decoding. The encoded representations of all passages are concatenated into one large sequence of hidden states. The decoder then attends over this entire concatenated representation via cross-attention while generating the answer token-by-token. This is where evidence fusion happens: the decoder's cross-attention can look at any position from any passage, freely combining information from different sources to produce each output token.

The name Fusion-in-Decoder captures exactly where the magic happens: evidence is fused not in the encoder (where passages are isolated), but in the decoder (where everything comes together).

Open in Lab
The full Fusion-in-Decoder pipeline. Click on each phase to see how passages flow from independent encoding to joint decoding.
The demo wakes as you arrive…

How FiD differs from RAG and extractive readers

FiD, RAG, and extractive readers all use a retriever, but they differ in how the reader processes retrieved passages.

An extractive reader (like DPR's reader) uses a BERT-style encoder. It processes each passage with the question, scores every token to find start and end positions, and extracts a span. It picks the best span across all passages. The problem: the model never sees multiple passages simultaneously, so it cannot combine evidence from two different passages.

RAG (Lewis et al., 2020) is a that processes each passage with the question through an encoder, then generates the answer with a decoder. However, RAG marginalizes over passages: it generates a separate answer for each passage and then combines the token probabilities using the retriever scores as weights. Each passage contributes independently — the decoder never attends to multiple passages at the same time.

FiD is the critical shift: the decoder attends to all encoded passages jointly. Cross-attention ranges over the concatenated encoder outputs from all passages, so the decoder can freely combine a date from passage 3, a name from passage 17, and a location from passage 42 — all in a single forward pass. This joint attention is what makes FiD scale so well with more passages: every additional passage adds new evidence the decoder can access.

Open in Lab
Compare three approaches: Extractive (single span), RAG (marginalize per passage), and FiD (joint decoder attention). See how each processes the same retrieved passages.
The demo wakes as you arrive…

Input formatting: how passages are prepared

Each retrieved passage is formatted as a single text sequence using special prefix tokens. The structure looks like this:

FiD input format for one passagetext

Simplified to show the idea — not the real implementation.

question: Where was Alan Turing born?
title: Alan Turing
context: Alan Turing was a British computer scientist.
Born in Maida Vale, London, he is widely considered
to be the father of theoretical computer science.

This format is applied to each of the 100 retrieved passages independently. The passages are truncated to 250 word pieces each. The encoder processes all 100 formatted inputs in parallel, producing 100 sequences of hidden states. These are then concatenated along the sequence dimension and passed to the decoder.

Results: state-of-the-art performance and passage scaling

FiD achieves state-of-the-art results across three major open-domain QA benchmarks. With the large model (770M parameters), it reaches 51.4 on NaturalQuestions, 67.6 on TriviaQA, and 56.7 on SQuAD Open — surpassing all prior models including T5-11B (which has 14× more parameters), RAG, DPR, and REALM.

But the most striking finding is how FiD scales with the number of passages. Increasing from 10 to 100 passages gives a 6% EM improvement on TriviaQA and 3.5% on NaturalQuestions. The performance curve does not plateau — it keeps climbing. In contrast, extractive models like Multi-Passage BERT typically peak at 10–20 passages and then stall or even degrade, because extractive readers cannot meaningfully aggregate signal from too many passages.

This scaling behavior is strong evidence that the decoder's cross-attention mechanism is genuinely fusing information across passages, not just picking the best one.

Open in Lab
Drag the slider to see how FiD's Exact Match score improves as the number of retrieved passages increases from 5 to 100.
The demo wakes as you arrive…

Training efficiency: fewer passages, then fine-tune

Training with 100 passages is computationally expensive. The authors explored a practical shortcut: train with fewer passages (e.g., 5 or 10), then fine-tune briefly with 100 passages for the last 1000 steps.

Training with 5 passages and with 100 reaches 45.0 EM on NaturalQuestions — only 1.5 points below the full 100-passage training (46.5 EM) — while using roughly 3× fewer GPU hours (147 vs 425). This works because the decoder learns the general skill of evidence fusion from a few passages, and the short fine-tuning phase adapts it to the longer attention patterns needed for 100 passages.

This finding is practically important: it means teams with limited compute budgets can still train effective FiD models.

Open in Lab
Compare training cost vs performance: train with few passages, then fine-tune with 100.
The demo wakes as you arrive…

Technical details

FiD is initialized from pretrained T5 weights (base: 220M parameters, large: 770M). It is fine-tuned on each dataset independently using Adam with a constant of 10−410^{-4} and a of 10%. Training runs for 10,000 gradient steps with a of 64 on 64 Tesla V100 GPUs. The best checkpoint is selected based on validation Exact Match score, evaluated every 500 steps.

Answer generation uses — no needed. Each passage is truncated to 250 word pieces. For NaturalQuestions and TriviaQA, passages are retrieved using DPR; for SQuAD Open, BM25 is used. The Wikipedia corpus is split into non-overlapping passages of 100 words each, following the DPR preprocessing pipeline.

L=−∑t=1Tlog⁡P(yt∣y<t, concat(E1,E2,…,EN))\mathcal{L} = -\sum_{t=1}^{T} \log P(y_t \mid y_{<t}, \, \text{concat}(E_1, E_2, \ldots, E_N))
FiD training objective — standard sequence-to-sequence cross-entropy loss — The model is trained with the standard auto-regressive cross-entropy loss. At each step tt, it predicts the next answer token yty_t conditioned on previously generated tokens y<ty_{<t} and the concatenated encoder outputs from all NN passages (E1,E2,…,EN)(E_1, E_2, \ldots, E_N). The encoder outputs are fixed per input; only the decoder generates tokens sequentially.

Legacy: the decoder-fusion paradigm

  1. 2017

    DrQA (Chen et al.)

    Established the retrieve-then-read pipeline. Used TF-IDF retrieval and a single-passage LSTM reader. First modern open-domain QA system.

  2. 2020

    DPR (Karpukhin et al.)

    Replaced sparse TF-IDF retrieval with dense dual-encoder retrieval using BERT. Dramatically improved retrieval recall and downstream QA performance.

  3. 2020

    RAG (Lewis et al.)

    Combined DPR retrieval with a BART generator. Marginalized over passages but did not fuse them jointly in the decoder — each passage contributed independently.

  4. 2021

    Fusion-in-Decoder (this paper)

    Encode passages independently, fuse in the decoder via cross-attention over all passages jointly. Scales to 100 passages with continued improvement.

  5. 2022

    RETRO (Borgeaud et al.)

    Extended retrieval-augmented generation to language modeling, retrieving chunks at every layer. Scaled to trillions of tokens using a frozen retrieval database.

  6. 2023

    Atlas (Izacard et al.)

    Built on FiD with end-to-end retriever training, few-shot learning, and joint retriever-reader optimization. Achieved strong performance with only 770M parameters.

FiD's insight — that independent encoding plus joint decoding is both efficient and powerful — became a design pattern adopted across the field. It showed that you do not need end-to-end differentiable retrieval (like REALM or RAG) to get strong results; sometimes a fixed retriever plus a powerful reader is all you need. The architecture is simple enough to implement in an afternoon, yet it held the state of the art for years.

CitationIzacard, Grave. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. EACL, 2021.

Terms in this paper