Information Retrieval2022intermediate11 min read
Contriever: Unsupervised Dense Information Retrieval with Contrastive Learning
Contriever: استرجاع كثيف للمعلومات بدون إشراف عبر التعلّم التبايُني
Izacard, G. · Caron, M. · Hosseini, L. · Riedel, S. · Bojanowski, P. · Joulin, A. · Grave, E. — TMLR
The problem
Dense retrievers based on neural networks had achieved strong results on benchmarks like MS MARCO, but only when large labeled datasets were available. When transferred to new domains with no training data (), they were consistently outperformed by classical unsupervised methods like , which simply count word overlaps. This meant that the power of neural retrieval was locked behind expensive human annotation. Additionally, labeled retrieval datasets barely existed outside English, making multilingual impractical.
The contribution
Contriever: a dense retriever trained entirely without supervision using . The model uses a architecture initialized from BERT, trained with the framework on Wikipedia and CCNet data. Positive pairs are created by randomly cropping two spans from the same document, with no human labels needed. On the BEIR , the unsupervised Contriever outperforms BM25 on 11 of 15 datasets for @100. When used as before on MS MARCO, it achieves state-of-the-art results among bi-encoders, and with re-ranking, it sets new state-of-the-art on 8 BEIR datasets. The multilingual variant, mContriever, enables — even between different scripts like Arabic queries to English documents.
The impact
Contriever proved that dense retrievers do not need labeled data to match or beat BM25, breaking the assumption that neural retrieval requires supervision. It became the retrieval backbone for Atlas (a retrieval-augmented language model) and inspired HyDE (hypothetical document embeddings). The contrastive pre-training recipe — random cropping + MoCo on raw text — established a new baseline for how to bootstrap dense retrievers from scratch, influencing every subsequent unsupervised and semi-supervised retrieval system.
Imagine you move to a new city and need to find the best restaurant for your taste. One approach: look for restaurants whose menu uses the exact words you typed — "spicy noodles." That is BM25: fast, reliable, but if the menu says "fiery ramen" instead, you miss it entirely.
Contriever takes a different approach. It wanders the city, tasting dishes and building an internal flavor map. Two dishes from the same kitchen end up close on the map; dishes from different kitchens end up far apart. After enough wandering — with no one ever telling it what is "good" — it can match your craving to the right kitchen, even when the words do not overlap at all.
The problem: dense retrievers need labels, BM25 does not
By 2021, (DPR) had shown that neural bi-encoders can dramatically outperform keyword-based search — but only on tasks with large supervised training sets like NaturalQuestions or MS MARCO. The moment you moved to a new domain — biomedical literature, financial reports, legal documents — where no labeled query–document pairs existed, the neural retriever collapsed. Plain BM25, which simply counts how often query words appear in the document (weighted by inverse document frequency), beat every supervised dense retriever in zero-shot settings.
The core issue is that models like DPR learn their from human-labeled positive pairs: "this question goes with this passage." Without those labels, the model has no signal to organize its space. Meanwhile, BM25 needs no training at all — it is a fixed formula applied to raw text.
This created an uncomfortable gap: the most powerful retrieval paradigm (dense vectors + approximate nearest neighbors) was also the most brittle when labels were absent. The question Contriever asks is: can we train the space itself without any labels, using only the structure of raw text?
Architecture: a shared bi-encoder with mean pooling
Contriever uses the bi- architecture. A single encoder (BERT-base) processes queries and documents independently. Each text is fed through the encoder, and the output is computed by averaging the hidden states of the last layer — a simple mean pooling over all tokens. This produces a single dense vector for each input.
The relevance score between a query and a document is the of their representations. Crucially, Contriever uses the same encoder for both queries and documents, unlike DPR which uses separate encoders. The authors found that sharing weights improves robustness in zero-shot and few-shot settings, because the model learns a single unified embedding space rather than two potentially misaligned ones.
Because documents are encoded independently of queries, the entire document collection can be pre-encoded once and stored. At query time, retrieval reduces to a fast maximum inner product search (MIPS), implemented with libraries like FAISS.
Training signal: contrastive learning from raw text
The key insight is that every document is, in some way, unique. Two excerpts from the same document should be more similar to each other than to excerpts from different documents. This is the only supervision signal Contriever needs — no human labels, no relevance judgments.
The model learns by discrimination: given a query , it must identify the correct positive key (from the same document) among a pool of negative keys (from other documents). This is formalized with the InfoNCE loss, borrowed from contrastive learning in computer vision.
Building pairs: random cropping beats Inverse Cloze Task
How you create positive pairs from a single document is the most critical design decision. Previous work used the Inverse Cloze Task (ICT): randomly select a sentence from the document as the "query," and the remaining text becomes the "key." The query and key are mutually exclusive — they have zero overlap.
Contriever instead uses independent random cropping: sample two random spans from the document, independently. The spans may overlap partially, fully, or not at all. This is the same strategy that powered SimCLR and MoCo in computer vision, adapted to text.
Why does cropping work better than ICT? Two reasons. First, partial overlap between the two views encourages the model to learn something akin to lexical matching — it discovers that shared words are a relevance signal, similar to what BM25 does naturally. Second, cropping produces symmetric distributions for queries and keys (both are random spans), which stabilizes training. ICT creates an asymmetry: the query is a single sentence while the key is an entire paragraph minus that sentence.
The authors also add random word deletion (10% of tokens) on top of cropping. This teaches the model not to rely too heavily on exact word matching and to build more robust representations.
Scaling negatives: MoCo momentum queue
Contrastive learning benefits enormously from more negative examples. With in-batch negatives, the number of negatives equals the , so you need huge GPU memory. Contriever uses MoCo ( Contrast) to decouple the number of negatives from the batch size.
MoCo maintains two networks: a query encoder updated by normal , and a key encoder updated as an exponential moving average of the query encoder. Keys computed by the key encoder are stored in a queue. At each step, the oldest keys are dequeued and fresh keys are enqueued. This gives Contriever access to 131,072 negatives — far more than any batch size could support.
The momentum update ensures the key encoder changes slowly, so representations in the queue remain consistent. Without it, old representations would be computed by a very different version of the model, creating noisy gradients.
Results: unsupervised Contriever matches BM25
The headline result: on the BEIR benchmark (15 diverse retrieval datasets), the fully unsupervised Contriever outperforms BM25 on 11 out of 15 datasets for Recall@100. This is remarkable because Contriever has never seen a single labeled query–document pair — its entire training signal comes from random cropping of raw text.
On NaturalQuestions, Contriever achieves 82.1% Recall@100 versus 78.3% for BM25. On TriviaQA, both tie at 83.2%. Contriever beats all previous unsupervised dense retrievers (ICT, REALM, SimCSE) by wide margins.
When Contriever is used as pre-training before fine-tuning on MS MARCO, it achieves the best average Recall@100 (67.1) among all bi-encoder methods on BEIR, beating DPR (48.3), ANCE (60.1), TAS-B (65.0), and Splade v2 (64.8). With cross-encoder re-ranking on top, it reaches state-of-the-art on 8 of 15 BEIR datasets.
In the few-shot setting (a few hundred to a few thousand labeled examples), the contrastive pre-training gives Contriever a massive advantage: it outperforms even BERT fine-tuned on the large MS MARCO dataset as an intermediate step.
Multilingual: cross-lingual retrieval without parallel data
Labeled retrieval datasets are almost exclusively in English. For low-resource languages, supervised dense retrieval is not an option. This is where unsupervised pre-training shines.
The authors train mContriever, a multilingual variant initialized from mBERT and pre-trained with contrastive learning on CCNet data from 29 languages. On the Mr. TyDi benchmark (11 languages), the unsupervised mContriever already beats BM25 for Recall@100. After fine-tuning on English-only MS MARCO data, it achieves state-of-the-art Recall@100, with improvements across all languages — even ones not seen during MS MARCO fine-tuning.
Most impressively, mContriever enables cross-lingual retrieval: given a query in Arabic, Japanese, or Korean, it retrieves relevant documents from an English Wikipedia collection. This is fundamentally impossible for BM25 or any term-matching method, since the query and document scripts share no vocabulary. On the MKQA benchmark, mContriever fine-tuned on English MS MARCO outperforms the CORA retriever, which was specifically trained on cross-lingual data with translation augmentation.
Ablations: what matters and what does not
The paper includes careful ablations that isolate the contribution of each design choice. The key findings are as follows.
Cropping vs ICT: Random cropping (nDCG@10 avg: 32.2) outperforms the Inverse Cloze Task (25.9) when used without fine-tuning. Adding random word deletion to cropping (33.8) further improves performance.
Number of negatives: Increasing the MoCo queue from 2,048 to 131,072 consistently improves performance, especially in the unsupervised setting. More negatives give the model harder discrimination problems, which builds better representations.
Training data: Wikipedia alone is good for Wikipedia-like tasks (FEVER), while CCNet alone excels on diverse domains (FiQA, Quora). A 50/50 mix of both gives the best overall performance.
MoCo vs in-batch negatives: The two methods perform similarly, but MoCo scales to many more negatives without requiring enormous batch sizes.
Impact of contrastive pre-training: When fine-tuned on MS MARCO, Contriever (avg nDCG@10: 46.5) significantly outperforms BERT with the same fine-tuning recipe (42.0), proving the value of the contrastive stage.
Code: Contriever in PyTorch
Simplified to show the idea — not the real implementation.
import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
# ── Encoder with mean pooling ──────────────────────────
class Contriever(torch.nn.Module):
def __init__(self, model_name="bert-base-uncased"):
super().__init__()
self.encoder = AutoModel.from_pretrained(model_name)
def forward(self, input_ids, attention_mask):
# Run through BERT; get last hidden states
out = self.encoder(input_ids=input_ids,
attention_mask=attention_mask)
hidden = out.last_hidden_state # (B, L, D)
# Mean pooling over non-padding tokens
mask = attention_mask.unsqueeze(-1) # (B, L, 1)
pooled = (hidden * mask).sum(1) / mask.sum(1)
return pooled # (B, D)
# ── InfoNCE contrastive loss ───────────────────────────
def info_nce_loss(queries, keys, temperature=0.05):
"""queries, keys: (B, D) — positive pairs at same index"""
# Similarity matrix: (B, B)
logits = torch.mm(queries, keys.T) / temperature
labels = torch.arange(len(queries), device=queries.device)
return F.cross_entropy(logits, labels)Timeline: from sparse to unsupervised dense retrieval
2009
BM25
The gold standard of unsupervised retrieval. A term-frequency weighting scheme that needs no training. Remained dominant for over a decade.
2019
Sentence-BERT
Adapted BERT for sentence embeddings using siamese networks with supervised NLI data. Made dense sentence similarity practical.
2020
DPR (Dense Passage Retrieval)
First bi-encoder retriever to outperform BM25 on open-domain QA. Required large supervised datasets with hard negatives from BM25.
2020
SimCLR / MoCo (Computer Vision)
Contrastive learning frameworks for images. SimCLR used large in-batch negatives; MoCo used a momentum queue. Both showed unsupervised features can match supervised ones.
2022
Contriever (this paper)
Combined MoCo + random cropping for text to train a dense retriever with zero labels. Matched BM25 unsupervised, set new state-of-the-art with fine-tuning.
2022
Atlas
A retrieval-augmented language model using Contriever as its retriever. Showed that few-shot learning with retrieval can match much larger models.
2023
HyDE
Used an LLM to generate a hypothetical document for the query, then retrieved with Contriever. Zero-shot retrieval without any task-specific training.
CitationIzacard, Caron, Hosseini, Riedel, Bojanowski, Joulin, Grave. Unsupervised Dense Information Retrieval with Contrastive Learning. Transactions on Machine Learning Research (TMLR), 2022.
Terms in this paper
- Contrastive Learningالتعلم التبايُني
- Dense Retrievalالاسترجاع الكثيف
- Bi-Encoderالمُرمِّز الثنائي
- Unsupervised Learningالتعلّم غير الخاضع للإشراف
- BM25BM25
- Information Retrievalاسترجاع المعلومات
- MoCoMoCo
- Data Augmentationتعزيز البيانات
- Fine-Tuningالضبط الدقيق
- Cross-Lingual Retrievalالاسترجاع عبر اللغات
- Embeddingالتضمين
- Dot Productالضرب النقطي
- Negative Samplingالتعيين السلبي
- Transfer Learningنقل التعلم
- Zero-Shotالنمط الصفري