Language Models2023advanced12 min read
Atlas: Few-shot Learning with Retrieval Augmented Language Models
Atlas: التعلُّم بأمثلة قليلة عبر نماذج لغوية مُعزَّزة بالاسترجاع
Izacard, G. · Lewis, P. · Lomeli, M. · Hosseini, L. · Petroni, F. · Schick, T. · Dwivedi-Yu, J. · Joulin, A. · Riedel, S. · Grave, E. — JMLR
The problem
Large language models like PaLM (540B parameters) achieve impressive few-shot results by storing enormous amounts of knowledge in their weights. But this brute-force memorization is extremely expensive: hundreds of billions of parameters, weeks of on thousands of accelerators, and the knowledge becomes stale the moment training ends. Retrieval-augmented models had shown they can offload knowledge to an external index, but it was unclear whether they could also learn new tasks from just a handful of examples — the few-shot setting that made large LLMs so exciting.
The contribution
Atlas: a retrieval-augmented language model that jointly pre-trains a Contriever-based dense retriever with a T5-based language model. The system uses four novel retriever training objectives — Attention Distillation, EMDR², , and Leave-One-Out — to align retriever and reader. Atlas-11B reaches 42.4% accuracy on NaturalQuestions with only 64 examples, outperforming PaLM-540B by 3 points while using 50× fewer parameters. The paper also demonstrates that Atlas's knowledge can be updated by simply swapping the , without retraining.
The impact
Atlas proved that retrieval augmentation and are not competing paradigms but complementary ones. It showed that a carefully designed retrieval-augmented model can match or beat models 50× its size on tasks, fundamentally changing the cost equation for deploying capable AI systems. The methodology influenced subsequent work like Self-RAG, and the insight that knowledge can live in a swappable index rather than frozen weights became a design principle for modern RAG systems.
Imagine two students preparing for a quiz. Student A memorized the entire textbook — all 10,000 pages. Student B brought a much smaller brain but has a brilliant research assistant who instantly finds the three most relevant pages for any question. Student A is impressive but slow, expensive to train, and stuck with whatever edition of the textbook they memorized. Student B is lean, fast, and — if the textbook gets updated — just hands the assistant a new copy.
Atlas is Student B. Its "brain" is a T5 language model, and its "research assistant" is a Contriever retriever. The key breakthrough: the brain and the assistant went to school together, so the assistant learned exactly which pages the brain finds most helpful.
The problem: memorization is expensive and fragile
By 2022, large language models had demonstrated a remarkable ability: show them a handful of examples in the prompt — "few-shot learning" — and they could perform tasks they were never explicitly trained for. GPT-3 could answer trivia, classify sentiment, or translate, all from a few demonstrations.
But this ability came at a cost. To answer factual questions, the model needed to have the knowledge stored inside its parameters. This meant hundreds of billions of weights encoding everything from historical dates to scientific facts. PaLM used 540 billion parameters. The training cost was immense — thousands of TPUs for weeks — and the knowledge was frozen at training time.
Retrieval-augmented models offered an alternative: instead of memorizing, look things up. Models like RAG and RETRO showed that a smaller model with access to a document index could match much larger models on knowledge-intensive tasks. But there was a gap: nobody had shown that retrieval-augmented models could also be strong few-shot learners. Could a model that looks things up also learn new tasks from just a few examples?
The architecture: a retriever and a reader, trained as one
Atlas is built from two components that work in concert, like a researcher and a writer collaborating on every answer.
The Retriever (Contriever) is a dual- model based on BERT. Given a query, it encodes the query and every document in the index into dense vectors, then finds the top- documents by . Unlike keyword-based search (like BM25), captures semantic meaning — "What is the capital of France?" matches a document that says "Paris is where the French government sits" even without the word "capital."
The Reader (Fusion-in- / T5) is a model. It receives the query concatenated with each retrieved document, processes them independently through the encoder, then fuses all encoded representations in the decoder via . This Fusion-in-Decoder design scales linearly with the number of documents, unlike naive concatenation which would be quadratic.
The crucial innovation is joint training: the retriever does not operate in isolation. Its parameters are updated based on signals from the language model, so it learns to retrieve documents that actually help the reader produce correct answers.
Joint training: teaching the librarian what the reader needs
The hardest challenge in retrieval-augmented models is the chicken-and-egg problem: the retriever needs to know what documents are useful, but "useful" depends on what the reader can do with them. And the reader's performance depends on getting good documents. If either starts poorly, the system stays stuck.
Atlas solves this by training both components together, with the retriever receiving training signals derived from the reader's performance. The paper explores four retriever training objectives, each providing a different kind of feedback.
The idea behind Perplexity Distillation is elegant: for each retrieved document , ask the reader "how well can you predict the target if you only see this document?" The documents that lead to lower perplexity (better prediction) are the ones the retriever should rank higher. Formally, the retriever is trained to minimize the between its relevance scores and a distribution derived from the reader's per-document perplexity.
Fusion-in-Decoder: reading many documents efficiently
A naive approach to using retrieved documents would be to concatenate them all with the query and feed them to the model as one giant input. But has quadratic complexity, so doubling the input length quadruples the cost. With 20 retrieved passages, this becomes prohibitively expensive.
Fusion-in-Decoder (FiD) solves this elegantly. Each retrieved document is concatenated with the query and processed independently through the T5 encoder. The encoder produces a set of hidden representations for each document. These representations are then simply concatenated into one long sequence and passed to the decoder. The decoder uses cross-attention to attend over all document representations simultaneously.
The key insight: the encoder's self-attention — the expensive operation — processes each document separately (linear cost in number of documents). Only the decoder's cross-attention sees all documents together, and cross-attention is much cheaper than full self-attention over a concatenated input.
Pre-training: building a knowledge-seeking model from scratch
Atlas is not just fine-tuned for retrieval — it is pre-trained with retrieval from the start. The retriever is initialized from the unsupervised Contriever model, and the reader from T5-lm-adapt (a T5 variant trained only on unlabeled text). Both are then jointly pre-trained on a mixture of Wikipedia and Common Crawl data.
The pre-training task is : take a passage, mask some tokens, retrieve related documents, and predict the masked tokens using the retrieved context. This teaches the system to seek and use external knowledge from the very beginning of training.
A critical engineering detail: during pre-training, the document index must be periodically refreshed. As the retriever's weights change, the document embeddings become stale. Atlas re-encodes and re-indexes documents every 1,000 steps — a computationally expensive but necessary step that keeps the index aligned with the evolving retriever.
Few-shot results: small model, big performance
The paper's headline result: Atlas-11B, with only 64 training examples, reaches 42.4% accuracy on NaturalQuestions — beating PaLM-540B (39.6%), a model with 50× more parameters that required 50× more pre-training compute.
On TriviaQA, Atlas-11B reaches 84.7% with 64 examples. On MMLU (57 diverse tasks), it achieves strong results across domains, especially those that benefit from factual retrieval. On the KILT benchmark (8 knowledge-intensive tasks), Atlas sets new state-of-the-art with full and remains competitive even in the few-shot regime.
During fine-tuning, Atlas uses a practical optimization: query-side fine-tuning. Instead of re-encoding the entire document index (millions of documents) after each retriever update, only the is updated. The document encoder stays frozen and the index is re-ranked using the updated query embeddings. This makes fine-tuning feasible on a single machine.
Updateability: swapping the index, not the weights
One of Atlas's most practical advantages is updateability. Because knowledge lives in the document index rather than the model weights, you can update what the model "knows" by simply replacing the index — no retraining needed.
The authors demonstrated this with a temporal experiment on NaturalQuestions: they took questions whose answers changed over time (e.g., "Who is the president of X?") and showed that swapping from a 2017 Wikipedia index to a 2021 index updated the model's answers correctly, without any fine-tuning. A parameter-only model would need expensive retraining.
They also showed that the choice of index matters enormously. Using a Wikipedia-only index gives much better performance on knowledge-intensive tasks than a Common Crawl index, likely because Wikipedia is more curated and dense with factual information.
Ablations: what matters and what does not
The paper includes extensive ablation studies that reveal what makes Atlas work.
Retriever training objective: All four objectives (ADist, EMDR², PDist, LOOP) perform similarly, but PDist is the most stable and computationally efficient, making it the recommended default.
Pre-training task: Three pretext tasks were compared — prefix language modeling, masked language modeling, and title-to-section generation. Masked language modeling gives a slight edge and is adopted as the default.
Index content: Using Wikipedia as the index consistently outperforms Common Crawl for downstream tasks, even when Common Crawl was used during pre-training. This suggests that index quality matters more than index size.
Model scale: Atlas benefits from scale, but the gains from retrieval augmentation are larger than the gains from simply increasing parameters. Atlas-3B with retrieval outperforms closed-book T5-11B without retrieval on many tasks.
Code: how Atlas processes a query
Simplified to show the idea — not the real implementation.
# Atlas inference: retrieve → encode → fuse → generate
def atlas_inference(query: str, retriever, reader, index, k=20):
# Step 1: Encode query and retrieve top-k documents
query_emb = retriever.encode_query(query) # [1, d]
doc_scores = index.search(query_emb, k=k) # cosine similarity
top_docs = [index.get_doc(i) for i in doc_scores.top_k_ids]
# Step 2: Encode each (query + doc) pair independently
encoder_outputs = []
for doc in top_docs:
input_text = f"question: {query} context: {doc.text}"
enc_out = reader.encoder(input_text) # [seq_len, d_model]
encoder_outputs.append(enc_out)
# Step 3: Fuse — concatenate all encoder outputs
fused = torch.cat(encoder_outputs, dim=0) # [k * seq_len, d_model]
# Step 4: Decode — cross-attend over all documents
answer = reader.decoder.generate(
encoder_output=fused, # decoder sees ALL docs
max_length=64
)
return answerTimeline: from RAG to Atlas and beyond
2020
RAG — Retrieval-Augmented Generation
Lewis et al. introduced retrieval-augmented generation, combining DPR retrieval with a BART generator. Demonstrated that retrieval + generation outperforms pure generation on knowledge-intensive tasks.
2020
Fusion-in-Decoder (FiD)
Izacard & Grave proposed processing each retrieved document independently in the encoder and fusing in the decoder — the architecture that would become Atlas's reader.
2022
Contriever
Izacard et al. introduced unsupervised dense retrieval using contrastive learning — no labeled query-document pairs needed. This became Atlas's retriever backbone.
2023
Atlas
Joint pre-training of retriever + reader with four novel training objectives. 11B parameters beat 540B on knowledge-intensive few-shot tasks. Demonstrated updateability and interpretability.
2023
Self-RAG
Built on Atlas's insight that retrieval and generation should be tightly coupled, adding self-reflection tokens to decide when and what to retrieve.
Atlas's deepest insight is not about any particular architecture or training trick — it is about where knowledge should live. Large language models had been treating parameters as the universal storage medium: more knowledge meant more parameters. Atlas demonstrated that a well-designed retrieval system separates the skill of reasoning from the content being reasoned about, and that this separation leads to better performance at lower cost.
The principle endures: today's most capable RAG systems — from search-augmented chatbots to enterprise question-answering pipelines — trace their lineage through Atlas's demonstration that joint training, not just bolt-on retrieval, is what unlocks few-shot knowledge-intensive performance.
CitationIzacard, Lewis, Lomeli, Hosseini, Petroni, Schick, Dwivedi-Yu, Joulin, Riedel, Grave. Atlas: Few-shot Learning with Retrieval Augmented Language Models. JMLR, 2023.
Terms in this paper
- Retrieval-Augmented Generationالتوليد المعزَّز بالاسترجاع
- Few-Shot Learningالتعلّم بأمثلة قليلة
- Dense Retrievalالاسترجاع الكثيف
- Contrastive Learningالتعلم التبايُني
- Fusion-in-Decoderالدمج في المفكِّك
- Knowledge Distillationتقطير المعرفة
- Joint Trainingالتدريب المشترك
- Document Indexفهرس المستندات
- Query Encoderمُرمِّز الاستعلام
- Perplexity Distillationتقطير الحيرة
- Masked Language Modeling (MLM)نمذجة اللغة المُقنَّعة (MLM)
- In-Context Learningالتعلم في السياق
- Knowledge-Intensiveكثيف المعرفة
- Encoder-Decoderمرمِّز-فاكّ ترميز
- Information Retrievalاسترجاع المعلومات