Language Models2023intermediate12 min read
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Self-RAG: كيف يتعلّم النموذج أن يسترجع ويُولِّد وينتقد نفسه
Asai, A. · Wu, Z. · Wang, Y. · Sil, A. · Hajishirzi, H. — ICLR
The problem
Large language models produce fluent text but frequently hallucinate — stating facts that sound plausible but are wrong. (RAG) tries to fix this by always prepending retrieved documents to the prompt. But this one-size-fits-all approach has two problems: it retrieves even when retrieval is unnecessary (hurting versatility), and the model is never trained to verify whether its output actually matches the retrieved evidence (hurting ).
The contribution
Self-RAG trains a single to decide on-the-fly whether retrieval is needed, retrieve relevant passages, generate text segment by segment, and each segment using special (Retrieve, IsRelevant, IsSupported, IsUseful). These tokens are predicted as part of the normal , enabling controllable at time — adjusting retrieval frequency and quality trade-offs without retraining. A segment-level selects the best output based on weighted critique scores. Self-RAG 7B and 13B outperform ChatGPT and retrieval-augmented Llama2-chat across six diverse tasks.
The impact
Self-RAG demonstrated that retrieval does not have to be all-or-nothing. By making retrieval and critique learnable decisions within the generation loop, it opened a new design axis for RAG systems: controlled by the model itself. The reflection framework influenced subsequent work on self-evaluating LLMs and controllable generation, and its segment-level critique approach became a reference point for factuality-aware generation.
Imagine two students taking an open-book exam. The first student — standard RAG — photocopies five random pages from the textbook before every single question, staples them to the answer sheet, and writes an answer hoping the pages help. Even for "describe your summer vacation," they staple five pages.
The second student — Self-RAG — reads the question and first asks: "Do I need the book for this?" If the answer is a personal essay, they skip the book entirely. If it is a factual question, they look up specific passages, check each one for relevance, write a sentence, then re-check: "Does my sentence match what the passage says?" Only then do they move to the next sentence.
The first student is thorough but wasteful. The second is precise, self-aware, and produces answers you can actually verify.
The problem: standard RAG retrieves blindly
Retrieval-Augmented Generation (RAG) was a major step forward: instead of relying entirely on memorized knowledge, the model consults external documents. But standard RAG has a rigid pipeline — retrieve documents, prepend them, generate. Three problems arise:
First, retrieval happens indiscriminately. Whether the question needs factual grounding or is a creative writing prompt, the model always retrieves. This can inject irrelevant or distracting context.
Second, the model is never trained to check whether retrieved passages actually support its output. It may generate text that contradicts the evidence, producing a fluent but factually wrong answer with false citations.
Third, retrieval is one-shot: the model retrieves once at the beginning and never revisits the decision. If the first retrieval misses, there is no second chance.
The solution: Self-RAG — retrieve, generate, and critique in one model
Self-RAG rethinks the entire RAG pipeline. Instead of a fixed retrieve-then-generate pipeline, Self-RAG trains a single language model to perform three actions as an integrated loop, segment by segment:
Step 1 — Decide: Given the input and any preceding text, the model predicts a special Retrieve token. If retrieval is needed, it outputs Retrieve=Yes and calls a retriever. If not, it outputs Retrieve=No and generates freely from its own knowledge.
Step 2 — Generate with evidence: When retrieval is triggered, the retriever returns passages. The model processes all in parallel, generating a candidate continuation from each. For each passage, it first predicts an IsRelevant token — is this passage actually useful for the query?
Step 3 — Critique: After generating each candidate segment, the model predicts two more reflection tokens: IsSupported (is the output actually backed by the passage?) and IsUseful (is this a helpful response overall?). A segment-level beam search then selects the best continuation based on weighted scores from all three critique dimensions.
The key insight is that all of these decisions — retrieve, assess relevance, verify support, rate utility — are made by the same language model as part of its normal next-token prediction, using special tokens added to its vocabulary.
Reflection tokens: the language of self-critique
The central innovation of Self-RAG is the reflection token system. These are special tokens added to the model's vocabulary — not natural language feedback, but discrete control signals that the model learns to generate as part of its output. There are four types:
Retrieve decides whether the next segment needs retrieval. Its values are Yes, No, or Continue (keep using the current passage). Think of it as the model raising its hand to say "I need to look something up."
IsRelevant evaluates whether a retrieved passage provides useful information for the input. Values: Relevant or Irrelevant. This acts as a filter — if the passage is irrelevant, the model knows not to trust it.
IsSupported checks whether the generated output is actually backed by the passage. Values: Fully Supported, Partially Supported, or No Support / Contradictory. This is the factuality guard — it catches the model when it invents claims that go beyond the evidence.
IsUseful rates the overall quality of the response on a 1–5 scale, independently of whether it is factual. A response could be fully supported but still unhelpful if it misses the point of the question.
Training: building the critic, then the generator
Self-RAG's has two stages. The first stage trains a model that can predict reflection tokens. The second stage uses the critic to annotate a large training corpus, then trains the generator model on this annotated data.
Stage 1 — Training the critic: Manually labeling reflection tokens for every segment would be prohibitively expensive. Instead, the authors prompt GPT-4 with type-specific instructions (e.g., "Given this instruction and output, decide if retrieval would help") and collect 4k–20k labeled examples per token type. These labels are then used to fine-tune a smaller in-house critic model (Llama 2 7B) via standard language modeling. The resulting critic achieves over 90% agreement with GPT-4 predictions — meaning the authors distill GPT-4's judgment into a small, efficient model they fully control.
Stage 2 — Training the generator: Once the critic is trained, it is used offline to annotate the full training corpus. For each input-output pair, the critic decides whether retrieval is needed, evaluates retrieved passages for relevance, checks whether the output is supported, and rates overall utility. All of these reflection tokens are inserted directly into the text. The generator model is then trained on this enriched corpus using the standard next-token prediction objective:
Inference: segment-level beam search with critique scores
At inference time, Self-RAG generates text segment by segment (typically sentence by sentence). At each segment step, the process unfolds as follows:
If Retrieve=Yes, the retriever returns passages. The model generates candidate segments — one from each passage — in parallel. Each candidate is scored by a weighted sum of the critique token probabilities:
This design makes Self-RAG controllable without retraining. A practitioner deploying the model for a medical QA system can increase the IsSupported weight to demand strong evidence grounding. The same model deployed for creative writing can reduce retrieval frequency by raising the threshold for triggering Retrieve=Yes.
The retrieval threshold itself is adaptive. Instead of always retrieving when Retrieve=Yes has the highest probability, Self-RAG checks whether the normalized probability of Retrieve=Yes exceeds a threshold . A higher means less retrieval — useful for tasks where parametric knowledge suffices.
Results: outperforming much larger models
Self-RAG was evaluated on six diverse tasks: two closed-set tasks (PubHealth fact verification and ARC-Challenge reasoning), two open-domain QA tasks (PopQA and TriviaQA), and two long-form generation tasks (biography generation evaluated with FactScore, and ASQA evaluated for correctness, , and accuracy).
Self-RAG 7B and 13B consistently outperform baselines across all tasks. On PopQA, Self-RAG 13B achieves 55.8% accuracy, beating retrieval-augmented ChatGPT (50.8%) and Llama2-chat 13B (51.8%). On PubHealth, Self-RAG 13B reaches 74.5%, far surpassing ChatGPT (70.1%). Most strikingly, on biography generation, Self-RAG 7B achieves a FactScore of 81.2, outperforming the 65B-parameter CoVE model (71.2).
On ASQA citation metrics, Self-RAG 7B achieves 66.9% citation precision and 67.8% recall — approaching ChatGPT's citation recall (76.6%) while exceeding its precision (65.1%). This shows Self-RAG generates claims that are genuinely backed by cited evidence, not just plausible-sounding citations.
Why it works: ablation studies
The ablation studies reveal which components matter most. Removing the critic entirely (No Critic) — training on data always augmented with the top passage but without reflection tokens — causes major drops, especially on ASQA (from 32.1 to 18.1 str-em). This shows that blindly prepending retrieved text is insufficient; the model needs to learn when retrieval is useful and when it is not.
Removing the retriever (No Retriever) and training without any retrieved passages also hurts, confirming that retrieval does provide value. But the gap from removing the critic is larger — suggesting that the mechanism matters more than retrieval alone.
At inference time, using only the top-1 passage (like standard RAG) instead of processing multiple candidates in parallel causes notable drops on PopQA and ASQA. And removing the IsSupported score from beam search specifically hurts citation quality, confirming that each critique dimension provides independent value.
The big picture: from blind retrieval to self-aware generation
Self-RAG introduced three ideas that shifted how the community thinks about retrieval-augmented generation:
First, adaptive retrieval: the model decides when retrieval is helpful rather than always retrieving. This preserves the versatility of the base LM — creative tasks skip retrieval entirely, factual tasks retrieve aggressively.
Second, integrated self-critique: instead of relying on an external verifier or reward model at inference time, the model's own generation contains quality assessments. This makes verification a built-in capability rather than an add-on.
Third, inference-time controllability: by adjusting the weights on critique token scores, practitioners can tune the same model for different applications without retraining. High citation precision for legal or medical applications, high fluency for creative writing — all from the same checkpoint.
These ideas point toward a future where language models are not just generators but active agents that monitor their own reliability, consult external knowledge only when needed, and produce outputs that carry their own evidence trail.
Context: the evolution toward self-aware retrieval
2020
RAG (Lewis et al.)
Retrieval-Augmented Generation jointly trains a retriever and generator. Always retrieves, never critiques — the foundation that Self-RAG improves upon.
2022
InstructGPT / RLHF (Ouyang et al.)
Trained GPT-3 with human feedback via PPO. Showed that reward signals can steer generation quality — Self-RAG's reflection tokens draw inspiration from this but avoid the cost of online RL.
2023
Self-Refine (Madaan et al.)
Iteratively prompts a model to generate, critique in natural language, then refine. Unlike Self-RAG, the critique is unstructured text, making it slower and harder to use as a scoring signal.
2023
Self-RAG (this paper)
Unified retrieval, generation, and critique into one model with reflection tokens. Adaptive retrieval, segment-level beam search, inference-time controllability.
2023
Active RAG (Jiang et al.)
Adaptively retrieves passages during generation using a proprietary LLM. Self-RAG achieves similar adaptive behavior but through learned reflection tokens in an open model.
Self-RAG sits at a turning point in the retrieval-augmented generation landscape. Before it, retrieval was a fixed preprocessing step. After it, retrieval became a learnable, adaptive decision embedded within the generation process itself. The reflection token framework showed that self-assessment does not require expensive reinforcement learning — it can be learned through supervised distillation and exercised through standard next-token prediction.
CitationAsai, Wu, Wang, Sil, Hajishirzi. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. ICLR, 2024.
Terms in this paper
- Retrieval-Augmented Generationالتوليد المعزَّز بالاسترجاع
- Self-Reflectionالمراجعة الذاتية
- Reflection Tokensرموز التأمُّل
- Adaptive Retrievalالاسترجاع التكيُّفي
- Critiqueالنقد الذاتي للبنية
- Hallucinationالهلوسة الرقمية
- Fine-Tuningالضبط الدقيق
- Beam Searchبحث الحزمة
- Factualityالدقة الواقعية
- Inferenceالاستدلال
- Language Modelالنموذج اللغوي
- Knowledge Distillationتقطير المعرفة
- Reward Modelنموذج المكافأة
- Instruction Tuningالضبط التعليمي
- Passage Retrievalاسترجاع الفقرات