Information Retrieval2022intermediate11 min read

HyDE: Precise Zero-Shot Dense Retrieval without Relevance Labels

HyDE: استرجاع كثيف دقيق دون أمثلة ودون تسميات صِلة

Gao, L. · Ma, X. · Lin, J. · Callan, J. — ACL

The problem

encodes queries and documents into vectors and matches them by . But learning two separate encoders — one for queries and one for documents — that land in the same requires massive amounts of human-labeled relevance judgments. Without those labels, the system has no signal for what counts as "relevant." Unsupervised dense retrievers like Contriever exist, but they still struggle because queries and documents look fundamentally different: a query is short and interrogative, while a document is long and declarative. Bridging that gap without supervision is the core challenge.

The contribution

HyDE decomposes the retrieval problem into two easier tasks. First, an instruction-following (InstructGPT) generates a that "answers" the query — fictional but relevance-patterned. Second, an unsupervised contrastive (Contriever) embeds that hypothetical document and searches for real documents nearby in space. No model is trained or fine-tuned. The key insight is that document-document similarity is far easier to capture without supervision than query-document similarity. HyDE significantly outperforms Contriever alone and rivals supervised fine-tuned systems across web search, QA, fact verification, and multilingual retrieval.

The impact

HyDE introduced a paradigm shift: instead of training retrieval models, use generative language models to bridge the query-document gap. This "generate-then-retrieve" pattern became foundational in RAG pipelines, where LLMs are augmented with retrieved knowledge. HyDE demonstrated that relevance can be modeled through language generation rather than numerical scoring, influencing subsequent work like Self-RAG, query rewriting, and hypothetical document expansion in production search systems.

Imagine you walk into a massive library and ask the librarian: "What causes rainbows?" The librarian has two choices. She can take your five-word question and scan every book's index for those exact words — but questions and book passages don't look alike, so the matches are poor.

Or she can do something clever: she quickly writes a rough paragraph about light refraction, water droplets, and the visible spectrum — not a perfect answer, but it sounds like what a relevant passage would say. Then she searches the shelves for pages that resemble her draft. Because her paragraph looks like a book passage, the matches are far better.

HyDE is that clever librarian. It uses a language model to draft a hypothetical answer, then uses that draft — not the question — to find real documents.

The gap: queries and documents speak different languages

Dense retrieval works by encoding both queries and documents into vectors in a shared embedding space, then finding documents whose vectors are closest to the query vector. The similarity is measured by the inner product:

sim(q,d)=⟨encq(q),  encd(d)⟩=⟨vq,  vd⟩\text{sim}(q, d) = \langle \text{enc}_q(q),\; \text{enc}_d(d) \rangle = \langle v_q,\; v_d \rangle
Dense retrieval similarity — inner product between query and document embeddings — Two separate encoders map the query qq and document dd into vectors vqv_q and vdv_d. Their inner product serves as the relevance score. The challenge: learning encq\text{enc}_q and encd\text{enc}_d to place relevant pairs near each other requires relevance labels.

The fundamental difficulty is this: queries and documents are structurally different. A query like "how long does wisdom tooth removal take?" is short, interrogative, and lacks context. A relevant document passage is long, declarative, and packed with details: "Wisdom tooth extraction typically takes 30 to 60 minutes depending on the complexity of the case..."

Training two encoders to map these structurally different inputs into a shared space where their inner product captures relevance requires massive amounts of labeled query-document pairs — datasets like MS-MARCO with hundreds of thousands of human judgments. Without such labels, the system has no gradient signal for what "relevant" means.

Unsupervised approaches like Contriever learn document-document similarity through , but they still struggle when the input is a query rather than a document — the gap between the question form and the passage form remains.

Open in Lab
See how queries and documents land in different regions of embedding space. The gap makes matching hard without supervised training.
The demo wakes as you arrive…

The HyDE insight: don't match queries to documents — match documents to documents

HyDE's key insight is deceptively simple: instead of trying to bridge the gap between queries and documents, transform the query into a document first, then search in document-only space.

The system works in two steps. First, an instruction-following language model (like InstructGPT) receives the query along with a task-specific instruction such as "Write a passage that answers the question" and generates a hypothetical document. This document is fictional — it may contain factual errors and hallucinations — but it captures the pattern of what a relevant answer looks like: the right vocabulary, the right structure, the right level of detail.

Second, an unsupervised contrastive encoder (like Contriever) encodes this hypothetical document into a vector. This vector now lives in the same space as all the real documents in the . The search becomes document-to-document matching — a problem that contrastive learning handles well without any relevance labels.

Open in Lab
Follow the HyDE pipeline step by step: query enters, hypothetical document is generated, encoded, and used to retrieve real documents.
The demo wakes as you arrive…

The mechanism: from query to hypothetical embedding

The formal setup begins with a key design choice: use a single encoder for everything. Instead of separate query and document encoders, HyDE uses only the document encoder from an unsupervised contrastive model. Denote this encoder ff:

f=encd=encconf = \text{enc}_d = \text{enc}_{\text{con}}
Single encoder — the document encoder is the contrastive encoder — By using a single encoder, the search happens entirely in document embedding space. Both the hypothetical document and real corpus documents are encoded by the same function.

Now consider an instruction-following language model gg. Given a query qq and a task-specific instruction INST\text{INST} (such as "write a passage to answer the question"), gg generates a hypothetical document:

d^=g(q,  INST)\hat{d} = g(q,\; \text{INST})
Hypothetical document generation — The instruction-following LM generates a hypothetical document d^\hat{d}. It captures relevance patterns — vocabulary, structure, detail level — but may be factually wrong.

The hypothetical document is then encoded by ff to produce the query vector. To reduce variance, HyDE samples NN hypothetical documents and averages their embeddings:

v^q=1N∑k=1Nf(d^k),d^k∼g(q,INST)\hat{v}_q = \frac{1}{N} \sum_{k=1}^{N} f(\hat{d}_k), \quad \hat{d}_k \sim g(q, \text{INST})
Averaged hypothetical embedding — reducing variance by sampling multiple documents — Each sampled document d^k\hat{d}_k is encoded independently, then the vectors are averaged. The averaged vector captures the central tendency of what a relevant document looks like, smoothing out individual hallucinations.

The inner product is then computed between v^q\hat{v}_q and all document vectors {f(d)∣d∈D}\{f(d) \mid d \in D\}. The most similar real documents are returned. The encoder ff acts as a lossy compressor: it maps the hypothetical document into a dense vector, discarding specific (possibly hallucinated) details and retaining only the semantic neighborhood. This is what grounds the retrieval — the vector identifies the right region of embedding space, even when the generated text is factually wrong.

Think of it as a postal system: the hypothetical document is an envelope with the right zip code but wrong street address. The encoder extracts the zip code (the semantic neighborhood) and delivers you to the right area, where the real documents live.

Open in Lab
Watch the encoder's dense bottleneck filter out hallucinated details while preserving the correct semantic neighborhood.
The demo wakes as you arrive…

Task adaptation through instructions

A critical design feature of HyDE is how it adapts to different retrieval tasks without any training. The same underlying models (InstructGPT + Contriever) serve web search, scientific fact-checking, financial QA, and multilingual retrieval — all by changing only the text instruction given to the language model.

For web search: "Please write a passage to answer the question." For scientific claims: "Please write a scientific paper passage to support/refute the claim." For financial QA: "Please write a financial article passage to answer the question." For Korean retrieval: "Please write a passage in Korean to answer the question in detail."

This is the power of instruction-following models: the same frozen model can generate documents in the right style for each task simply by reading a different instruction. No , no task-specific training data, no architectural changes.

Open in Lab
Switch between tasks and see how the same query generates different hypothetical documents through different instructions.
The demo wakes as you arrive…

Results: unsupervised performance that rivals supervised systems

HyDE was evaluated across 11 query sets covering web search (TREC DL19, DL20), low-resource retrieval (6 BEIR datasets), and multilingual retrieval (Mr.TyDi in Swahili, Korean, Japanese, and Bengali).

On web search, HyDE improved Contriever's nDCG@10 from 44.5 to 61.3 on DL19 — a 38% relative improvement. It matched or exceeded the performance of ContrieverFT, a version of Contriever fine-tuned on MS-MARCO's hundreds of thousands of labeled pairs.

On low-resource tasks, HyDE outperformed (the classical lexical baseline) on 5 out of 6 datasets, and outperformed supervised models like DPR and ANCE that were fine-tuned on MS-MARCO.

On multilingual retrieval, HyDE improved mContriever across all four languages, even though GPT-3's generation quality varies across languages. HyDE outperformed fine-tuned models like mDPR and mBERT, though it trailed the fully fine-tuned mContrieverFT.

Open in Lab
Compare HyDE against baselines across web search, low-resource, and multilingual tasks. Toggle between metrics.
The demo wakes as you arrive…

Scaling: bigger language models, better retrieval

The authors tested HyDE with three instruction-following language models of different sizes: FLAN-T5 (11B parameters), Cohere (52B), and InstructGPT (175B). The results revealed a clear scaling trend: larger language models produced better hypothetical documents, which led to better retrieval.

On DL19, nDCG@10 climbed from 48.9 (FLAN-T5) to 53.8 (Cohere) to 61.3 (InstructGPT). Even the smallest model improved over the Contriever baseline (44.5), confirming that the HyDE approach works across scales — but benefits significantly from stronger generation.

This has a practical implication: as language models continue to improve, HyDE will automatically improve too, without any retraining of the retrieval system.

Why hallucinations don't break retrieval

At first glance, it seems paradoxical: the language model generates text that is factually unreliable, yet the retrieval results are precise. The resolution lies in the encoder's dense bottleneck.

When the encoder maps a hypothetical document to a vector, it compresses a variable-length text into a fixed-size vector (typically 768 dimensions). This compression is lossy by design — it cannot preserve every detail. What it does preserve is the overall semantic neighborhood: the topic, the domain, the type of content.

Specific factual claims — "wisdom tooth removal takes exactly 47 minutes" — are details that get averaged out in the embedding. But the general semantic signal — "this is about dental procedures and surgery duration" — is robustly captured. The retrieval finds documents in the right neighborhood, regardless of the exact (hallucinated) numbers.

This is why HyDE works: the bottleneck acts as a filter between and retrieval. It keeps the signal (semantic relevance) and discards the noise (specific fabrications).

The bigger question: is relevance just language understanding?

The authors close the paper with a provocative reflection. Traditional retrieval systems model relevance as a numerical score — a learned function that maps query-document pairs to numbers. HyDE replaces this with language generation: the concept of relevance is captured by a language model producing an example of what a relevant document looks like.

This raises a deep question: is numerical relevance scoring just a statistical artifact of language understanding? If a strong enough language model can generate text that captures relevance patterns, and a simple encoder can ground that text in a corpus, do we need to learn relevance scores at all?

The authors do not claim to have the answer, but the results suggest that as language models grow stronger, the need for explicit relevance modeling may diminish. The boundary between information retrieval and natural language understanding is blurring.

Context: the generate-then-retrieve timeline

  1. 2020

    Dense Passage Retrieval (DPR)

    Karpukhin et al. showed that dual-encoder dense retrieval trained on labeled question-passage pairs outperforms BM25 for open-domain QA. But it requires hundreds of thousands of labeled pairs.

  2. 2021

    Contriever — unsupervised contrastive retrieval

    Izacard et al. trained dense retrievers with unsupervised contrastive learning, removing the need for relevance labels. Strong but still limited by the query-document gap.

  3. 2022

    InstructGPT — instruction following at scale

    Ouyang et al. showed GPT-3 models aligned with human intent can zero-shot follow diverse instructions. This unlocked the "generate hypothetical documents" strategy.

  4. 2022

    HyDE (this paper)

    Combined InstructGPT generation with Contriever encoding to achieve supervised-level zero-shot retrieval. No model training, no relevance labels.

  5. 2023

    RAG pipelines go mainstream

    Retrieval-Augmented Generation became the standard architecture for grounding LLMs in external knowledge. HyDE's generate-then-retrieve pattern influenced query expansion and rewriting in production RAG systems.

HyDE's most lasting contribution is not a specific model but a conceptual reframing. It showed that retrieval does not require explicitly modeling relevance scores. Instead, a language model that understands relevance implicitly — through its ability to generate relevant-looking text — can be combined with simple unsupervised matching to achieve state-of-the-art results. This insight has reshaped how we think about the interface between language generation and information retrieval.

CitationGao, Ma, Lin, Callan. Precise Zero-Shot Dense Retrieval without Relevance Labels. ACL, 2023.

Terms in this paper