Language Models2023intermediate12 min read

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Self-RAG: كيف يتعلّم النموذج أن يسترجع ويُولِّد وينتقد نفسه

Asai, A. · Wu, Z. · Wang, Y. · Sil, A. · Hajishirzi, H. — ICLR

The problem

Large language models produce fluent text but frequently hallucinate — stating facts that sound plausible but are wrong. (RAG) tries to fix this by always prepending retrieved documents to the prompt. But this one-size-fits-all approach has two problems: it retrieves even when retrieval is unnecessary (hurting versatility), and the model is never trained to verify whether its output actually matches the retrieved evidence (hurting ).

The contribution

Self-RAG trains a single to decide on-the-fly whether retrieval is needed, retrieve relevant passages, generate text segment by segment, and each segment using special (Retrieve, IsRelevant, IsSupported, IsUseful). These tokens are predicted as part of the normal , enabling controllable at time — adjusting retrieval frequency and quality trade-offs without retraining. A segment-level selects the best output based on weighted critique scores. Self-RAG 7B and 13B outperform ChatGPT and retrieval-augmented Llama2-chat across six diverse tasks.

The impact

Self-RAG demonstrated that retrieval does not have to be all-or-nothing. By making retrieval and critique learnable decisions within the generation loop, it opened a new design axis for RAG systems: controlled by the model itself. The reflection framework influenced subsequent work on self-evaluating LLMs and controllable generation, and its segment-level critique approach became a reference point for factuality-aware generation.

Imagine two students taking an open-book exam. The first student — standard RAG — photocopies five random pages from the textbook before every single question, staples them to the answer sheet, and writes an answer hoping the pages help. Even for "describe your summer vacation," they staple five pages.

The second student — Self-RAG — reads the question and first asks: "Do I need the book for this?" If the answer is a personal essay, they skip the book entirely. If it is a factual question, they look up specific passages, check each one for relevance, write a sentence, then re-check: "Does my sentence match what the passage says?" Only then do they move to the next sentence.

The first student is thorough but wasteful. The second is precise, self-aware, and produces answers you can actually verify.

The problem: standard RAG retrieves blindly

Retrieval-Augmented Generation (RAG) was a major step forward: instead of relying entirely on memorized knowledge, the model consults external documents. But standard RAG has a rigid pipeline — retrieve KK documents, prepend them, generate. Three problems arise:

First, retrieval happens indiscriminately. Whether the question needs factual grounding or is a creative writing prompt, the model always retrieves. This can inject irrelevant or distracting context.

Second, the model is never trained to check whether retrieved passages actually support its output. It may generate text that contradicts the evidence, producing a fluent but factually wrong answer with false citations.

Third, retrieval is one-shot: the model retrieves once at the beginning and never revisits the decision. If the first retrieval misses, there is no second chance.

Open in Lab
Compare standard RAG (left) and Self-RAG (right). Notice how Self-RAG decides whether to retrieve at each step and critiques its own output.
The demo wakes as you arrive…

The solution: Self-RAG — retrieve, generate, and critique in one model

Self-RAG rethinks the entire RAG pipeline. Instead of a fixed retrieve-then-generate pipeline, Self-RAG trains a single language model to perform three actions as an integrated loop, segment by segment:

Step 1 — Decide: Given the input and any preceding text, the model predicts a special Retrieve token. If retrieval is needed, it outputs Retrieve=Yes and calls a retriever. If not, it outputs Retrieve=No and generates freely from its own knowledge.

Step 2 — Generate with evidence: When retrieval is triggered, the retriever returns KK passages. The model processes all KK in parallel, generating a candidate continuation from each. For each passage, it first predicts an IsRelevant token — is this passage actually useful for the query?

Step 3 — Critique: After generating each candidate segment, the model predicts two more reflection tokens: IsSupported (is the output actually backed by the passage?) and IsUseful (is this a helpful response overall?). A segment-level beam search then selects the best continuation based on weighted scores from all three critique dimensions.

The key insight is that all of these decisions — retrieve, assess relevance, verify support, rate utility — are made by the same language model as part of its normal next-token prediction, using special tokens added to its vocabulary.

Open in Lab
Explore the four reflection token types. Click each to see its role, possible values, and where it appears in the generation flow.
The demo wakes as you arrive…

Reflection tokens: the language of self-critique

The central innovation of Self-RAG is the reflection token system. These are special tokens added to the model's vocabulary — not natural language feedback, but discrete control signals that the model learns to generate as part of its output. There are four types:

Retrieve decides whether the next segment needs retrieval. Its values are Yes, No, or Continue (keep using the current passage). Think of it as the model raising its hand to say "I need to look something up."

IsRelevant evaluates whether a retrieved passage provides useful information for the input. Values: Relevant or Irrelevant. This acts as a filter — if the passage is irrelevant, the model knows not to trust it.

IsSupported checks whether the generated output is actually backed by the passage. Values: Fully Supported, Partially Supported, or No Support / Contradictory. This is the factuality guard — it catches the model when it invents claims that go beyond the evidence.

IsUseful rates the overall quality of the response on a 1–5 scale, independently of whether it is factual. A response could be fully supported but still unhelpful if it misses the point of the question.

Training: building the critic, then the generator

Self-RAG's has two stages. The first stage trains a model that can predict reflection tokens. The second stage uses the critic to annotate a large training corpus, then trains the generator model on this annotated data.

Stage 1 — Training the critic: Manually labeling reflection tokens for every segment would be prohibitively expensive. Instead, the authors prompt GPT-4 with type-specific instructions (e.g., "Given this instruction and output, decide if retrieval would help") and collect 4k–20k labeled examples per token type. These labels are then used to fine-tune a smaller in-house critic model (Llama 2 7B) via standard language modeling. The resulting critic achieves over 90% agreement with GPT-4 predictions — meaning the authors distill GPT-4's judgment into a small, efficient model they fully control.

max⁡C  E((x,y), r) ∼ Dcriticlog⁡ pC(r∣x,y)\max_{C} \; \mathbb{E}_{((x,y),\,r)\,\sim\,\mathcal{D}_{\text{critic}}} \log\, p_C(r \mid x, y)
Critic training objective — predicting reflection tokens given input and output — The critic model CC is trained to maximize the probability of the correct reflection token rr given the input xx and output yy. This is a standard conditional language modeling objective — the same way any language model learns to predict the next token, the critic learns to predict the appropriate critique.

Stage 2 — Training the generator: Once the critic is trained, it is used offline to annotate the full training corpus. For each input-output pair, the critic decides whether retrieval is needed, evaluates retrieved passages for relevance, checks whether the output is supported, and rates overall utility. All of these reflection tokens are inserted directly into the text. The generator model is then trained on this enriched corpus using the standard next-token prediction objective:

max⁡M  E(x,y,r) ∼ Dgenlog⁡ pM(y,r∣x)\max_{M} \; \mathbb{E}_{(x,y,r)\,\sim\,\mathcal{D}_{\text{gen}}} \log\, p_M(y, r \mid x)
Generator training objective — learning to produce text and reflection tokens together — The generator MM learns to predict both the output text yy and the reflection tokens rr given input xx. Critically, the retrieved passages themselves are masked out during loss computation — the model does not memorize specific passages but learns the pattern of when to retrieve, what is relevant, and what is supported.
Open in Lab
The two-stage training pipeline: GPT-4 labels are distilled into a critic model, which then annotates the training data for the generator.
The demo wakes as you arrive…

Inference: segment-level beam search with critique scores

At inference time, Self-RAG generates text segment by segment (typically sentence by sentence). At each segment step, the process unfolds as follows:

If Retrieve=Yes, the retriever returns KK passages. The model generates KK candidate segments — one from each passage — in parallel. Each candidate is scored by a weighted sum of the critique token probabilities:

f(yt,d)=p(yt∣x,d,y<t)+∑G∈{IsRel, IsSup, IsUse}wG⋅stGf(y_t, d) = p(y_t \mid x, d, y_{<t}) + \sum_{G \in \{\text{IsRel},\,\text{IsSup},\,\text{IsUse}\}} w^G \cdot s^G_t
Segment scoring function — combining generation likelihood with critique scores — Each candidate segment yty_t generated from passage dd is scored by combining its generation probability with a weighted sum of critique scores. stGs^G_t is the normalized probability of the most desirable token for critique type GG (e.g., stIsRels^{\text{IsRel}}_t is the probability of *Relevant* normalized over {Relevant, Irrelevant}). The weights wGw^G are hyperparameters adjustable at inference time — raising wIsSupw^{\text{IsSup}} produces more citation-precise output, while lowering it yields more fluent, expansive text.
Open in Lab
Watch Self-RAG process multiple passages in parallel and select the best segment using weighted critique scores.
The demo wakes as you arrive…

This design makes Self-RAG controllable without retraining. A practitioner deploying the model for a medical QA system can increase the IsSupported weight to demand strong evidence grounding. The same model deployed for creative writing can reduce retrieval frequency by raising the threshold δ\delta for triggering Retrieve=Yes.

The retrieval threshold itself is adaptive. Instead of always retrieving when Retrieve=Yes has the highest probability, Self-RAG checks whether the normalized probability of Retrieve=Yes exceeds a threshold δ\delta. A higher δ\delta means less retrieval — useful for tasks where parametric knowledge suffices.

Results: outperforming much larger models

Self-RAG was evaluated on six diverse tasks: two closed-set tasks (PubHealth fact verification and ARC-Challenge reasoning), two open-domain QA tasks (PopQA and TriviaQA), and two long-form generation tasks (biography generation evaluated with FactScore, and ASQA evaluated for correctness, , and accuracy).

Self-RAG 7B and 13B consistently outperform baselines across all tasks. On PopQA, Self-RAG 13B achieves 55.8% accuracy, beating retrieval-augmented ChatGPT (50.8%) and Llama2-chat 13B (51.8%). On PubHealth, Self-RAG 13B reaches 74.5%, far surpassing ChatGPT (70.1%). Most strikingly, on biography generation, Self-RAG 7B achieves a FactScore of 81.2, outperforming the 65B-parameter CoVE model (71.2).

On ASQA citation metrics, Self-RAG 7B achieves 66.9% citation precision and 67.8% recall — approaching ChatGPT's citation recall (76.6%) while exceeding its precision (65.1%). This shows Self-RAG generates claims that are genuinely backed by cited evidence, not just plausible-sounding citations.

Why it works: ablation studies

The ablation studies reveal which components matter most. Removing the critic entirely (No Critic) — training on data always augmented with the top passage but without reflection tokens — causes major drops, especially on ASQA (from 32.1 to 18.1 str-em). This shows that blindly prepending retrieved text is insufficient; the model needs to learn when retrieval is useful and when it is not.

Removing the retriever (No Retriever) and training without any retrieved passages also hurts, confirming that retrieval does provide value. But the gap from removing the critic is larger — suggesting that the mechanism matters more than retrieval alone.

At inference time, using only the top-1 passage (like standard RAG) instead of processing multiple candidates in parallel causes notable drops on PopQA and ASQA. And removing the IsSupported score from beam search specifically hurts citation quality, confirming that each critique dimension provides independent value.

Open in Lab
Step through Self-RAG's decision process for a sample query. See where retrieval is triggered, how passages are filtered, and which segment is selected.
The demo wakes as you arrive…

The big picture: from blind retrieval to self-aware generation

Self-RAG introduced three ideas that shifted how the community thinks about retrieval-augmented generation:

First, adaptive retrieval: the model decides when retrieval is helpful rather than always retrieving. This preserves the versatility of the base LM — creative tasks skip retrieval entirely, factual tasks retrieve aggressively.

Second, integrated self-critique: instead of relying on an external verifier or reward model at inference time, the model's own generation contains quality assessments. This makes verification a built-in capability rather than an add-on.

Third, inference-time controllability: by adjusting the weights on critique token scores, practitioners can tune the same model for different applications without retraining. High citation precision for legal or medical applications, high fluency for creative writing — all from the same checkpoint.

These ideas point toward a future where language models are not just generators but active agents that monitor their own reliability, consult external knowledge only when needed, and produce outputs that carry their own evidence trail.

Context: the evolution toward self-aware retrieval

  1. 2020

    RAG (Lewis et al.)

    Retrieval-Augmented Generation jointly trains a retriever and generator. Always retrieves, never critiques — the foundation that Self-RAG improves upon.

  2. 2022

    InstructGPT / RLHF (Ouyang et al.)

    Trained GPT-3 with human feedback via PPO. Showed that reward signals can steer generation quality — Self-RAG's reflection tokens draw inspiration from this but avoid the cost of online RL.

  3. 2023

    Self-Refine (Madaan et al.)

    Iteratively prompts a model to generate, critique in natural language, then refine. Unlike Self-RAG, the critique is unstructured text, making it slower and harder to use as a scoring signal.

  4. 2023

    Self-RAG (this paper)

    Unified retrieval, generation, and critique into one model with reflection tokens. Adaptive retrieval, segment-level beam search, inference-time controllability.

  5. 2023

    Active RAG (Jiang et al.)

    Adaptively retrieves passages during generation using a proprietary LLM. Self-RAG achieves similar adaptive behavior but through learned reflection tokens in an open model.

Self-RAG sits at a turning point in the retrieval-augmented generation landscape. Before it, retrieval was a fixed preprocessing step. After it, retrieval became a learnable, adaptive decision embedded within the generation process itself. The reflection token framework showed that self-assessment does not require expensive reinforcement learning — it can be learned through supervised distillation and exercised through standard next-token prediction.

CitationAsai, Wu, Wang, Sil, Hajishirzi. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. ICLR, 2024.

Terms in this paper