Information Retrieval2022intermediate11 min read

Contriever: Unsupervised Dense Information Retrieval with Contrastive Learning

Contriever: استرجاع كثيف للمعلومات بدون إشراف عبر التعلّم التبايُني

Izacard, G. · Caron, M. · Hosseini, L. · Riedel, S. · Bojanowski, P. · Joulin, A. · Grave, E. — TMLR

The problem

Dense retrievers based on neural networks had achieved strong results on benchmarks like MS MARCO, but only when large labeled datasets were available. When transferred to new domains with no training data (), they were consistently outperformed by classical unsupervised methods like , which simply count word overlaps. This meant that the power of neural retrieval was locked behind expensive human annotation. Additionally, labeled retrieval datasets barely existed outside English, making multilingual impractical.

The contribution

Contriever: a dense retriever trained entirely without supervision using . The model uses a architecture initialized from BERT, trained with the framework on Wikipedia and CCNet data. Positive pairs are created by randomly cropping two spans from the same document, with no human labels needed. On the BEIR , the unsupervised Contriever outperforms BM25 on 11 of 15 datasets for @100. When used as before on MS MARCO, it achieves state-of-the-art results among bi-encoders, and with re-ranking, it sets new state-of-the-art on 8 BEIR datasets. The multilingual variant, mContriever, enables — even between different scripts like Arabic queries to English documents.

The impact

Contriever proved that dense retrievers do not need labeled data to match or beat BM25, breaking the assumption that neural retrieval requires supervision. It became the retrieval backbone for Atlas (a retrieval-augmented language model) and inspired HyDE (hypothetical document embeddings). The contrastive pre-training recipe — random cropping + MoCo on raw text — established a new baseline for how to bootstrap dense retrievers from scratch, influencing every subsequent unsupervised and semi-supervised retrieval system.

Imagine you move to a new city and need to find the best restaurant for your taste. One approach: look for restaurants whose menu uses the exact words you typed — "spicy noodles." That is BM25: fast, reliable, but if the menu says "fiery ramen" instead, you miss it entirely.

Contriever takes a different approach. It wanders the city, tasting dishes and building an internal flavor map. Two dishes from the same kitchen end up close on the map; dishes from different kitchens end up far apart. After enough wandering — with no one ever telling it what is "good" — it can match your craving to the right kitchen, even when the words do not overlap at all.

The problem: dense retrievers need labels, BM25 does not

By 2021, (DPR) had shown that neural bi-encoders can dramatically outperform keyword-based search — but only on tasks with large supervised training sets like NaturalQuestions or MS MARCO. The moment you moved to a new domain — biomedical literature, financial reports, legal documents — where no labeled query–document pairs existed, the neural retriever collapsed. Plain BM25, which simply counts how often query words appear in the document (weighted by inverse document frequency), beat every supervised dense retriever in zero-shot settings.

The core issue is that models like DPR learn their from human-labeled positive pairs: "this question goes with this passage." Without those labels, the model has no signal to organize its space. Meanwhile, BM25 needs no training at all — it is a fixed formula applied to raw text.

This created an uncomfortable gap: the most powerful retrieval paradigm (dense vectors + approximate nearest neighbors) was also the most brittle when labels were absent. The question Contriever asks is: can we train the space itself without any labels, using only the structure of raw text?

Open in Lab
Compare how BM25 (sparse) and a dense retriever find documents. BM25 matches exact words; the dense model matches meaning — but needs training signal.
The demo wakes as you arrive…

Architecture: a shared bi-encoder with mean pooling

Contriever uses the bi- architecture. A single encoder (BERT-base) processes queries and documents independently. Each text is fed through the encoder, and the output is computed by averaging the hidden states of the last layer — a simple mean pooling over all tokens. This produces a single dense vector for each input.

The relevance score between a query qq and a document dd is the of their representations. Crucially, Contriever uses the same encoder for both queries and documents, unlike DPR which uses separate encoders. The authors found that sharing weights improves robustness in zero-shot and few-shot settings, because the model learns a single unified embedding space rather than two potentially misaligned ones.

Because documents are encoded independently of queries, the entire document collection can be pre-encoded once and stored. At query time, retrieval reduces to a fast maximum inner product search (MIPS), implemented with libraries like FAISS.

s(q,d)=⟨fθ(q),  fθ(d)⟩s(q, d) = \langle f_\theta(q),\; f_\theta(d) \rangle
Relevance score — dot product of shared encoder representations — The same encoder fθf_\theta maps both the query qq and document dd to dense vectors. The dot product measures their similarity: high when the texts are semantically related, low when they are not. Mean pooling over the last transformer layer produces each vector.
Open in Lab
The shared bi-encoder pipeline: both query and document pass through the same BERT encoder, then mean pooling, then dot product for scoring.
The demo wakes as you arrive…

Training signal: contrastive learning from raw text

The key insight is that every document is, in some way, unique. Two excerpts from the same document should be more similar to each other than to excerpts from different documents. This is the only supervision signal Contriever needs — no human labels, no relevance judgments.

The model learns by discrimination: given a query qq, it must identify the correct positive key k+k^+ (from the same document) among a pool of KK negative keys (from other documents). This is formalized with the InfoNCE loss, borrowed from contrastive learning in computer vision.

L(q,k+)=−log⁡exp⁡(s(q,k+)/τ)exp⁡(s(q,k+)/τ)+∑i=1Kexp⁡(s(q,ki)/τ)\mathcal{L}(q, k^+) = -\log \frac{\exp\bigl(s(q, k^+)/\tau\bigr)} {\exp\bigl(s(q, k^+)/\tau\bigr) + \sum_{i=1}^{K} \exp\bigl(s(q, k_i)/\tau\bigr)}
InfoNCE contrastive loss — The loss pushes the score s(q,k+)s(q, k^+) of the positive pair up, while pushing scores of all KK negative pairs down. The temperature τ\tau controls sharpness: lower τ\tau makes the model more selective. This is exactly a softmax classification over K+1K+1 candidates — "which key came from the same document as the query?"

Building pairs: random cropping beats Inverse Cloze Task

How you create positive pairs from a single document is the most critical design decision. Previous work used the Inverse Cloze Task (ICT): randomly select a sentence from the document as the "query," and the remaining text becomes the "key." The query and key are mutually exclusive — they have zero overlap.

Contriever instead uses independent random cropping: sample two random spans from the document, independently. The spans may overlap partially, fully, or not at all. This is the same strategy that powered SimCLR and MoCo in computer vision, adapted to text.

Why does cropping work better than ICT? Two reasons. First, partial overlap between the two views encourages the model to learn something akin to lexical matching — it discovers that shared words are a relevance signal, similar to what BM25 does naturally. Second, cropping produces symmetric distributions for queries and keys (both are random spans), which stabilizes training. ICT creates an asymmetry: the query is a single sentence while the key is an entire paragraph minus that sentence.

The authors also add random word deletion (10% of tokens) on top of cropping. This teaches the model not to rely too heavily on exact word matching and to build more robust representations.

Open in Lab
Compare how ICT and random cropping create positive pairs from the same document. Notice the overlap (highlighted) in cropping but not in ICT.
The demo wakes as you arrive…

Scaling negatives: MoCo momentum queue

Contrastive learning benefits enormously from more negative examples. With in-batch negatives, the number of negatives equals the , so you need huge GPU memory. Contriever uses MoCo ( Contrast) to decouple the number of negatives from the batch size.

MoCo maintains two networks: a query encoder θq\theta_q updated by normal , and a key encoder θk\theta_k updated as an exponential moving average of the query encoder. Keys computed by the key encoder are stored in a queue. At each step, the oldest keys are dequeued and fresh keys are enqueued. This gives Contriever access to 131,072 negatives — far more than any batch size could support.

The momentum update ensures the key encoder changes slowly, so representations in the queue remain consistent. Without it, old representations would be computed by a very different version of the model, creating noisy gradients.

θk←m θk+(1−m) θq\theta_k \leftarrow m\,\theta_k + (1-m)\,\theta_q
MoCo momentum update for the key encoder — At each step, the key encoder parameters θk\theta_k are nudged toward the query encoder parameters θq\theta_q by the momentum mm (set to 0.9995). When mm is close to 1, the key encoder evolves very slowly, keeping the queue consistent. This avoids the staleness problem of storing representations from a fast-changing network.
Open in Lab
Watch how MoCo's queue provides a large pool of negatives without increasing batch size. The momentum encoder drifts slowly to keep the queue consistent.
The demo wakes as you arrive…

Results: unsupervised Contriever matches BM25

The headline result: on the BEIR benchmark (15 diverse retrieval datasets), the fully unsupervised Contriever outperforms BM25 on 11 out of 15 datasets for Recall@100. This is remarkable because Contriever has never seen a single labeled query–document pair — its entire training signal comes from random cropping of raw text.

On NaturalQuestions, Contriever achieves 82.1% Recall@100 versus 78.3% for BM25. On TriviaQA, both tie at 83.2%. Contriever beats all previous unsupervised dense retrievers (ICT, REALM, SimCSE) by wide margins.

When Contriever is used as pre-training before fine-tuning on MS MARCO, it achieves the best average Recall@100 (67.1) among all bi-encoder methods on BEIR, beating DPR (48.3), ANCE (60.1), TAS-B (65.0), and Splade v2 (64.8). With cross-encoder re-ranking on top, it reaches state-of-the-art on 8 of 15 BEIR datasets.

In the few-shot setting (a few hundred to a few thousand labeled examples), the contrastive pre-training gives Contriever a massive advantage: it outperforms even BERT fine-tuned on the large MS MARCO dataset as an intermediate step.

Open in Lab
BEIR benchmark Recall@100: Contriever (unsupervised) vs BM25 across 15 datasets. Contriever wins on 11 of 15.
The demo wakes as you arrive…

Multilingual: cross-lingual retrieval without parallel data

Labeled retrieval datasets are almost exclusively in English. For low-resource languages, supervised dense retrieval is not an option. This is where unsupervised pre-training shines.

The authors train mContriever, a multilingual variant initialized from mBERT and pre-trained with contrastive learning on CCNet data from 29 languages. On the Mr. TyDi benchmark (11 languages), the unsupervised mContriever already beats BM25 for Recall@100. After fine-tuning on English-only MS MARCO data, it achieves state-of-the-art Recall@100, with improvements across all languages — even ones not seen during MS MARCO fine-tuning.

Most impressively, mContriever enables cross-lingual retrieval: given a query in Arabic, Japanese, or Korean, it retrieves relevant documents from an English Wikipedia collection. This is fundamentally impossible for BM25 or any term-matching method, since the query and document scripts share no vocabulary. On the MKQA benchmark, mContriever fine-tuned on English MS MARCO outperforms the CORA retriever, which was specifically trained on cross-lingual data with translation augmentation.

Ablations: what matters and what does not

The paper includes careful ablations that isolate the contribution of each design choice. The key findings are as follows.

Cropping vs ICT: Random cropping (nDCG@10 avg: 32.2) outperforms the Inverse Cloze Task (25.9) when used without fine-tuning. Adding random word deletion to cropping (33.8) further improves performance.

Number of negatives: Increasing the MoCo queue from 2,048 to 131,072 consistently improves performance, especially in the unsupervised setting. More negatives give the model harder discrimination problems, which builds better representations.

Training data: Wikipedia alone is good for Wikipedia-like tasks (FEVER), while CCNet alone excels on diverse domains (FiQA, Quora). A 50/50 mix of both gives the best overall performance.

MoCo vs in-batch negatives: The two methods perform similarly, but MoCo scales to many more negatives without requiring enormous batch sizes.

Impact of contrastive pre-training: When fine-tuned on MS MARCO, Contriever (avg nDCG@10: 46.5) significantly outperforms BERT with the same fine-tuning recipe (42.0), proving the value of the contrastive stage.

Code: Contriever in PyTorch

Contriever bi-encoder with mean pooling and InfoNCE losspython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer

# ── Encoder with mean pooling ──────────────────────────
class Contriever(torch.nn.Module):
    def __init__(self, model_name="bert-base-uncased"):
        super().__init__()
        self.encoder = AutoModel.from_pretrained(model_name)

    def forward(self, input_ids, attention_mask):
        # Run through BERT; get last hidden states
        out = self.encoder(input_ids=input_ids,
                           attention_mask=attention_mask)
        hidden = out.last_hidden_state          # (B, L, D)
        # Mean pooling over non-padding tokens
        mask = attention_mask.unsqueeze(-1)      # (B, L, 1)
        pooled = (hidden * mask).sum(1) / mask.sum(1)
        return pooled                            # (B, D)


# ── InfoNCE contrastive loss ───────────────────────────
def info_nce_loss(queries, keys, temperature=0.05):
    """queries, keys: (B, D) — positive pairs at same index"""
    # Similarity matrix: (B, B)
    logits = torch.mm(queries, keys.T) / temperature
    labels = torch.arange(len(queries), device=queries.device)
    return F.cross_entropy(logits, labels)

Timeline: from sparse to unsupervised dense retrieval

  1. 2009

    BM25

    The gold standard of unsupervised retrieval. A term-frequency weighting scheme that needs no training. Remained dominant for over a decade.

  2. 2019

    Sentence-BERT

    Adapted BERT for sentence embeddings using siamese networks with supervised NLI data. Made dense sentence similarity practical.

  3. 2020

    DPR (Dense Passage Retrieval)

    First bi-encoder retriever to outperform BM25 on open-domain QA. Required large supervised datasets with hard negatives from BM25.

  4. 2020

    SimCLR / MoCo (Computer Vision)

    Contrastive learning frameworks for images. SimCLR used large in-batch negatives; MoCo used a momentum queue. Both showed unsupervised features can match supervised ones.

  5. 2022

    Contriever (this paper)

    Combined MoCo + random cropping for text to train a dense retriever with zero labels. Matched BM25 unsupervised, set new state-of-the-art with fine-tuning.

  6. 2022

    Atlas

    A retrieval-augmented language model using Contriever as its retriever. Showed that few-shot learning with retrieval can match much larger models.

  7. 2023

    HyDE

    Used an LLM to generate a hypothetical document for the query, then retrieved with Contriever. Zero-shot retrieval without any task-specific training.

CitationIzacard, Caron, Hosseini, Riedel, Bojanowski, Joulin, Grave. Unsupervised Dense Information Retrieval with Contrastive Learning. Transactions on Machine Learning Research (TMLR), 2022.

Terms in this paper