Reinforcement Learning2023intermediate11 min read

Reflexion: Language Agents with Verbal Reinforcement Learning

Reflexion: وكلاء لغويون بالتعزيز اللفظي

Shinn, N. · Cassano, F. · Berman, E. · Gopinath, A. · Narasimhan, K. · Yao, S. — NeurIPS

The problem

By 2023, large language models were being used as goal-driven agents that interact with environments through actions — playing games, writing code, calling APIs. But when these agents failed, they had no good way to learn from mistakes. Traditional requires updating model weights, which is extremely expensive for billion-parameter LLMs. And simply retrying without reflection doesn't help — the tends to repeat the same errors. The field needed a way for language agents to learn from trial-and-error without any gradient updates.

The contribution

Reflexion: a framework that replaces gradient-based weight updates with verbal stored in an buffer. After each failed trial, a self-reflection model generates a natural-language analysis of what went wrong, which is appended to the agent's memory and used as context in subsequent attempts. The agent's is parameterized as (LLM + memory) rather than just LLM weights. Reflexion achieves 91% pass@1 on HumanEval (vs GPT-4's 80%), improves AlfWorld decision-making by 22%, and boosts HotPotQA reasoning by 20% — all without any .

The impact

Reflexion established verbal self-reflection as a legitimate reinforcement signal for language agents. It showed that LLMs can "learn" within a context window — no weight updates needed — by maintaining a structured memory of past mistakes. This insight shaped the design of agentic systems like Voyager, Self-RAG, and SWE-Bench agents. The idea that natural language can serve as both the and the channel became a foundational principle for the agentic AI wave of 2023–2024.

Imagine a chess player who loses a game, but instead of just replaying the board, they write in a notebook: "I lost my queen because I ignored the bishop's diagonal. Next time, check all diagonals before moving." Before the next game, they re-read their notes.

They don't study new openings. They don't hire a coach. They just talk to themselves — on paper — about what went wrong.

That notebook is Reflexion. The player's skill improves not because their brain changes, but because their memory carries forward hard-won lessons. This paper showed that language models can do the same thing: reflect on failures in words, store those reflections, and use them to avoid the same mistakes next time.

The problem: agents that repeat their mistakes

By 2023, LLM-based agents like ReAct could interact with external environments — searching the web, navigating text-based games, writing and executing code. But when they failed a task, they had only two options, both bad.

The first option is traditional reinforcement learning: compute gradients and update the model's weights. For a model with billions of parameters, this is prohibitively expensive. You'd need large datasets, specialized hardware, and hours of compute — just to teach the agent one lesson.

The second option is to simply retry. But retrying without memory is like re-taking an exam without studying: you tend to make the same mistakes. The agent's behavior is determined entirely by its fixed weights and the current . With no record of past failures, it has no reason to try a different approach.

Reflexion proposes a third path: let the agent talk to itself about what went wrong, save that self-talk, and read it before the next attempt.

The framework: three models in a loop

Reflexion has three components that work together in an iterative loop:

The Actor (MaM_a) is the LLM agent that generates actions. It can be any action-generating strategy — Chain of Thought for reasoning, or ReAct for multi-step decision-making. The Actor receives the current task plus any stored reflections from memory, and produces a of actions and observations.

The Evaluator (MeM_e) scores the Actor's output. This could be as simple as a binary pass/fail from a compiler, an exact-match check against a ground-truth answer, or even an LLM-based judge. It produces a reward signal rtr_t for trial tt.

The Self-Reflection model (MsrM_{sr}) is the heart of the framework. Given the failed trajectory τt\tau_t and the reward rtr_t, it generates a verbal analysis: what went wrong, which specific action caused the failure, and what to do differently. This text is appended to the agent's episodic memory.

Open in Lab
The Reflexion loop: the Actor attempts a task, the Evaluator scores it, the Self-Reflection model analyzes the failure, and the reflection is stored in memory for the next trial. Click each component to learn more.
The demo wakes as you arrive…

The critical design choice is how the policy is parameterized. In standard RL, the policy πθ\pi_\theta is the model weights θ\theta. In Reflexion, the policy is πθ\pi_\theta where θ={Ma,mem}\theta = \{M_a, \text{mem}\} — the frozen LLM plus the episodic memory buffer. Learning happens by updating the memory, not the weights:

πθ(at∣st),θ={Ma,  mem}\pi_\theta(a_t \mid s_t), \quad \theta = \{M_a, \;\text{mem}\}
Reflexion policy — the LLM weights are frozen, memory is the learnable parameter — The agent selects action ata_t given state sts_t by conditioning on both the frozen LLM MaM_a and the episodic memory buffer mem. The memory accumulates self-reflections from past failures. Learning is simply appending new reflections to mem — no gradients, no backpropagation.

Self-reflection: turning failure into lessons

The self-reflection step is where Reflexion converts a sparse reward signal into rich, actionable feedback. Consider a coding agent that submits code which fails 3 out of 5 unit tests. A traditional RL agent gets the reward r=0.4r = 0.4 and somehow has to figure out what to fix. The Reflexion self-reflection model instead produces text like:

"My implementation of the binary search used <= instead of < in the while condition, causing an infinite loop on edge cases. I also forgot to handle the empty array case. Next time, I should add boundary checks before the main loop and use < for the comparison."

This is vastly more informative than a scalar reward. It performs credit assignment — identifying which specific action caused the failure — and provides a concrete plan for the next attempt. The agent doesn't need to discover these insights through random exploration; the self-reflection model tells it directly.

Open in Lab
Watch the memory buffer grow across trials. Each failed attempt adds a self-reflection. The agent reads all past reflections before its next attempt.
The demo wakes as you arrive…

Evaluation signals: how the agent knows it failed

Reflexion is designed to work with different types and sources of feedback signals. The framework is flexible — what matters is that there is some signal the agent can reflect on.

For decision-making tasks like AlfWorld, the evaluator uses hand-written heuristics. If the agent repeats the same action three times or exceeds 30 steps, it's likely stuck — time to reflect. The environment itself signals whether the task is complete.

For reasoning tasks like HotPotQA, the evaluator uses exact-match grading: does the agent's answer match the ground truth? A simple binary signal — but the self-reflection model turns it into rich verbal feedback about why the answer was wrong.

For programming tasks, the agent generates its own unit tests, then runs the code against them. The compiler output — error messages, stack traces, failed assertions — becomes the raw material for self-reflection. This is particularly powerful because the feedback is already detailed and specific.

Open in Lab
Explore the three domains and their feedback types. Each domain converts its native signal into verbal reflection.
The demo wakes as you arrive…

Results: dramatic gains across three domains

Reflexion was tested on three fundamentally different types of tasks, each requiring different skills. Across all three, it substantially outperformed strong baselines — without any model fine-tuning.

Decision-making (AlfWorld): The agent navigates text-based household environments — find a spatula in a drawer, chill a tomato in the fridge. ReAct alone solved 75% of 134 tasks. ReAct + Reflexion solved 97% (130/134) across 12 iterative trials. The improvement came in two phases: a sharp initial spike (the agent quickly corrected obvious errors) and then a steady climb (it gradually learned to explore more systematically by using its memory to track which locations it had already checked).

Reasoning (HotPotQA): On 100 multi-hop questions requiring Wikipedia search and synthesis, Reflexion + ReAct improved accuracy by 20% over ReAct alone. Notably, the baseline agents showed zero improvement across retries (at temperature 0.7), while Reflexion agents steadily improved — proving that the self-reflection, not randomness, drove the gains.

Programming (HumanEval): Reflexion achieved 91% pass@1 accuracy on the HumanEval Python , surpassing GPT-4's 80%. The agent generates code, runs it against self-generated unit tests, reflects on failures, and iterates. This is where the framework shines brightest: compiler errors provide detailed, structured feedback that is ideal for self-reflection.

Open in Lab
Compare Reflexion against baselines across AlfWorld, HotPotQA, and HumanEval. Toggle between domains to see the improvement curves.
The demo wakes as you arrive…

Ablation: what makes self-reflection work?

The authors conducted ablation studies to isolate the contribution of each component. The most revealing experiment compared three approaches on HotPotQA reasoning with ground-truth context:

  1. Baseline retry — simply retry the task with no memory. Performance stays flat.

  2. Episodic memory (EPM) — include the most recent trajectory as context, but no verbal reflection. This helps moderately — the agent can see what it tried before.

  3. Reflexion — add the self-reflection step that distills the failure into a verbal lesson. This adds an 8% absolute improvement over EPM alone.

The gap between EPM and Reflexion demonstrates that raw experience replay is not enough. The distillation into verbal lessons is what makes the difference. The agent doesn't just see what it did — it understands why it failed and what to change.

For programming, removing self-generated tests cut accuracy from 91% to 82%. Without tests, the agent has no way to check its own work before submission. And removing self-reflection while keeping the tests dropped accuracy further — the tests alone provide raw error signals, but the self-reflection is needed to convert those signals into actionable plans.

Reflexion in code: the iterative loop

Reflexion loop pseudocodepython

Simplified to show the idea — not the real implementation.

# Reflexion reinforcement loop
def reflexion_loop(task, actor, evaluator, reflector, max_trials=5):
    memory = []  # episodic memory buffer
    for trial in range(max_trials):
        # Actor generates trajectory conditioned on task + memory
        trajectory = actor.generate(task, memory)

        # Evaluator scores the trajectory
        reward = evaluator.score(trajectory)
        if reward == 1.0:  # task solved
            return trajectory

        # Self-reflection: analyze why it failed
        reflection = reflector.reflect(trajectory, reward)
        # e.g. "I searched for the wrong entity. Next time,
        #        search for the director, not the movie."

        memory.append(reflection)  # store lesson
        # Keep only last 3 reflections (context limit)
        memory = memory[-3:]

    return None  # failed after max_trials

Limitations and the road ahead

Reflexion has clear limitations that the authors acknowledge. First, it relies on the LLM's ability to accurately identify its own mistakes — a form of self-evaluation that is not always reliable. If the model generates an incorrect self-reflection ("I failed because the API was slow" when the actual issue was a logic error), the bad lesson gets stored in memory and may hurt future attempts.

Second, the context window is a hard constraint. With memory capped at 1–3 reflections, the agent cannot accumulate unbounded experience. Older lessons are dropped, which means it may re-encounter and re-solve problems it had already learned from.

Third, there is no formal convergence guarantee. Traditional RL methods can prove convergence under certain conditions; Reflexion offers no such theoretical backing. Its effectiveness is empirical — it works well in practice, but we cannot prove it will always improve.

Despite these limitations, the core idea — that language can serve as a reinforcement signal — has proven remarkably durable and has influenced the design of nearly every subsequent agentic system.

Legacy: the self-improving agent paradigm

  1. 2022

    ReAct — reasoning meets acting

    Yao et al. interleaved reasoning traces with actions, letting LLMs think step-by-step while interacting with environments. This became the Actor backbone for Reflexion.

  2. 2023

    Reflexion (this paper)

    Introduced verbal reinforcement learning: reflect on failures in words, store reflections in episodic memory, and use them to improve on the next trial. 91% on HumanEval without fine-tuning.

  3. 2023

    Self-RAG — retrieval with self-reflection

    Extended the self-reflection idea to retrieval-augmented generation. The model decides when to retrieve, what to retrieve, and whether retrieved content is relevant — all through self-generated reflection tokens.

  4. 2023

    Voyager — lifelong learning agent in Minecraft

    Applied the reflect-and-store paradigm to open-ended exploration. Voyager builds a growing skill library where each skill is verified by the environment — a close cousin of Reflexion's memory buffer.

  5. 2024

    SWE-Bench agents — self-reflection in real codebases

    Agents tackling real GitHub issues adopted Reflexion-style iteration: attempt a fix, run tests, reflect on failures, and retry. The pattern proved essential for complex, multi-file debugging tasks.

CitationShinn, Cassano, Berman, Gopinath, Narasimhan, Yao. Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS, 2023.

Terms in this paper