Language Models2023intermediate11 min read

ReAct: Synergizing Reasoning and Acting in Language Models

ReAct: دمج الاستدلال والفعل في النماذج اللغوية

Yao, S. · Zhao, J. · Yu, D. · Du, N. · Shafran, I. · Narasimhan, K. · Cao, Y. — ICLR

The problem

By 2022, large language models had impressive but separate capabilities: chain-of-thought prompting let them reason step-by-step, and action planning let them interact with tools. But reasoning alone hallucinates — the model invents plausible-sounding facts from its own internal knowledge. Acting alone is blind — the model takes actions without understanding why or what to do with the results. Neither approach alone could solve tasks requiring both knowledge retrieval and multi-step reasoning.

The contribution

ReAct: a prompting paradigm that interleaves reasoning traces and task-specific actions in a single language model. The model generates thoughts (to decompose goals, track progress, handle exceptions) and actions (to interact with external tools like Wikipedia) in alternation. On question answering (HotpotQA) and (FEVER), ReAct eliminates by reasoning in retrieved facts. On interactive decision-making (ALFWorld, WebShop), ReAct outperforms imitation and reinforcement learning methods by 34% and 10% in success rate, using only 1–2 in-context examples.

The impact

ReAct is widely regarded as the founding paradigm of agentic AI. It proved that a single language model can serve as both the "brain" (reasoner) and the "hands" (actor) of an autonomous . LangChain, AutoGPT, and virtually every modern AI agent framework is built on the ReAct loop. The paper's key insight — that grounding reasoning in external observations eliminates hallucination while reasoning makes actions purposeful — became the blueprint for tool-using language models.

A student preparing for an exam can study two ways. The think-only student sits with eyes closed, trying to recall everything from memory — and inevitably invents facts they never learned. The act-only student flips through textbooks randomly, highlighting passages without knowing why they matter.

The ReAct student does what any effective learner does: reads a question, thinks about what they need to find, opens the right textbook, reads the relevant passage, thinks about what it means, then decides whether to look up more or write the answer. Thinking guides the search; the search grounds the thinking.

The problem: reasoning and acting lived in separate worlds

By 2022, two powerful capabilities had emerged in large language models, but they were studied in isolation:

  • Chain-of-thought (CoT) reasoning let models "think out loud" — decomposing problems into steps, performing arithmetic, and drawing conclusions. But CoT is a closed loop: the model reasons entirely from its own internal knowledge, with no way to check facts or gather new information. When its knowledge is wrong or incomplete, it hallucinates — confidently generating plausible-sounding but fabricated facts.

  • Action generation let models interact with external tools — searching the web, querying databases, or navigating environments. But without reasoning, the model takes actions blindly. It cannot decompose goals, track progress, handle unexpected results, or synthesize a final answer from what it finds.

Humans don't separate thinking from acting. When cooking, you think "I need salt," reach for the salt, discover you're out, think "soy sauce could work instead," and adapt. This tight loop between reasoning and acting is what makes human problem-solving robust. ReAct brings this loop to language models.

Open in Lab
Compare the same question solved by CoT (reason only), Act (act only), and ReAct (reason + act).
The demo wakes as you arrive…

The idea: augment the action space with thought

The ReAct idea is elegant in its simplicity. Consider an agent interacting with an . At each time step tt, the agent receives an oto_t and takes an action ata_t. Normally the AA contains only environment actions — search, click, navigate.

ReAct extends this action space to A^=A∪L\hat{A} = A \cup L, where LL is the space of natural language. An action in LL — a thought or — does not affect the external environment. It produces no observation. Instead, it updates the agent's internal context, helping it plan, reason, and decide.

This means thoughts are "free" — they cost no environment step, generate no side effects, but enrich the context that the model uses to choose its next real action. The model alternates between thinking (in language space LL) and acting (in environment space AA), creating a thought-action-observation loop.

A^=A∪L\hat{A} = A \cup L
The augmented action space — the core of ReAct — A = environment actions (search, lookup, finish) · L = language thoughts (plan, reason, extract) · Thoughts update internal context without affecting the environment
Open in Lab
Click any action to understand what it does. Notice how thoughts enrich context without costing an environment step.
The demo wakes as you arrive…

The loop in action: a HotpotQA walkthrough

ReAct uses : the model is given a small number of human-annotated trajectories as examples, then generates its own trajectories for new problems. Each trajectory is a sequence of thought-action-observation steps.

For knowledge-intensive tasks like HotpotQA (multi-hop question answering) and FEVER (fact verification), the model alternates thoughts and actions densely — every action is preceded by a thought explaining what to do and why. The action space is a simple Wikipedia API with three operations: search[entity] returns the first 5 sentences from a Wikipedia page, lookup[string] simulates Ctrl+F, and finish[answer] submits the final answer.

The walkthrough below shows the complete trajectory for a real HotpotQA question. Notice how each thought serves a purpose: decomposing the question, extracting relevant facts from observations, reformulating failed searches, and synthesizing the final answer.

Open in Lab
Step through a complete ReAct trajectory. Each thought guides the next action; each observation informs the next thought.
The demo wakes as you arrive…

Beyond QA: ReAct in interactive decision-making

ReAct is not limited to question answering. The paper tests it on two interactive decision-making environments where an agent must take many actions over long horizons:

  • ALFWorld — a text-based game where the agent navigates a simulated household to complete tasks like "put a clean knife on the countertop." The agent must plan subgoals (find knife → clean it → place it), reason about where household items are likely to be found (a knife is more likely on a countertop or in a drawer), and track which subgoals are complete.

  • WebShop — an online shopping environment with 1.18 million real products. The agent must find and purchase a product matching a user's instruction by searching, comparing products, and selecting the right options.

In these environments, thoughts appear sparsely — only at key decision points, not before every action. The model itself decides when a thought is needed. This flexible placement is one of ReAct's strengths: it adapts the density of reasoning to the task.

Results across all benchmarks

The paper evaluates ReAct on four diverse benchmarks using PaLM-540B with few-shot prompting. The results reveal a nuanced picture:

On FEVER (fact verification), ReAct clearly outperforms CoT (60.9 vs. 56.3 accuracy), because verifying claims often requires retrieving precise facts that the model's internal knowledge may get slightly wrong.

On HotpotQA (multi-hop QA), ReAct slightly lags behind CoT (27.4 vs. 29.4 exact match), because CoT's free-form reasoning is more flexible than ReAct's structured thought-action-observation format. However, the best overall method is combining both: ReAct→CoT-SC achieves 35.1, and CoT-SC→ReAct achieves 34.2 — significantly outperforming either alone.

On both ALFWorld and WebShop, ReAct dramatically outperforms all baselines, demonstrating that reasoning is most valuable in interactive environments where the agent must plan over many steps.

Open in Lab
Select a benchmark to compare ReAct against baselines. Notice how ReAct shines most in interactive tasks.
The demo wakes as you arrive…

Why ReAct is more trustworthy: a failure analysis

The authors performed a detailed human analysis of 200 trajectories from ReAct and CoT on HotpotQA. The findings are striking:

Among successful trajectories, 14% of CoT answers were correct by accident — the reasoning trace contained hallucinated facts that happened to lead to the right answer. For ReAct, only 6% had this problem. Among failed trajectories, the picture is even clearer: 56% of CoT failures were caused by hallucination — the model invented facts and reasoned from them confidently. For ReAct, hallucination caused 0% of failures. External grounding completely eliminates fabricated facts.

ReAct has its own failure mode: search result errors (23% of failures), where the search returns empty or unhelpful results that derail the reasoning. This is the trade-off between factuality and flexibility. But the authors show this can be mitigated by combining ReAct with CoT-SC — falling back to internal reasoning when external search fails, or vice versa.

Open in Lab
Switch between failure and success modes to see the dramatic difference in hallucination rates.
The demo wakes as you arrive…

Scaling up: from prompting to fine-tuning

An important finding is how ReAct scales with model size and training. With prompting alone, ReAct actually performs worst among all methods on small models (PaLM-8B, PaLM-62B) — learning both reasoning and acting from in-context examples is too complex for smaller models.

But when fine-tuned on just 3,000 ReAct trajectories with correct answers, the picture reverses dramatically. Fine-tuned ReAct becomes the best method across all model sizes. PaLM-8B fine-tuned with ReAct outperforms every PaLM-62B prompting method. PaLM-62B fine-tuned with ReAct outperforms every PaLM-540B prompting method. The intuition is clear: Standard or CoT teaches the model to memorize (potentially hallucinated) facts, while fine-tuning ReAct teaches the model how to retrieve and reason — a more generalizable skill.

Open in Lab
Toggle between prompting and fine-tuning to see how ReAct transforms from worst to best.
The demo wakes as you arrive…

Interpretability and human control

A distinctive advantage of ReAct is transparency. Because the model writes its reasoning in natural language before each action, humans can inspect the decision-making process in real time. This enables two powerful capabilities:

  • Diagnosability: when ReAct makes an error, the reasoning trace reveals exactly where the reasoning went wrong — whether the model misinterpreted an observation, hallucinated a fact, or chose the wrong search query. This is far more interpretable than a black-box model that simply outputs a wrong answer.

  • Human-in-the-loop editing: because thoughts are just text, a human can edit them mid-trajectory. The paper demonstrates that by simply removing a hallucinated sentence and adding a hint, a failing ReAct trajectory can be steered to success. This enables a new paradigm of human-AI collaboration where the human edits the model's thinking, not its actions.

What ReAct unlocked

  1. 2022

    ReAct — the paradigm is born

    Interleaving reasoning traces and actions in a single LLM. Proved that grounding eliminates hallucination and reasoning makes actions purposeful.

  2. 2023

    Toolformer — models learn to call APIs

    Extended the idea to let language models autonomously decide when and how to call external tools like calculators, translators, and search engines.

  3. 2023

    LangChain & AutoGPT — agent frameworks

    The ReAct loop became the backbone of agent frameworks. LangChain's zero-shot ReAct agent made tool-using AI accessible to every developer.

  4. 2024

    Claude, GPT-4, Gemini — tool use as standard

    Every major LLM now supports tool use natively. The think-act-observe pattern that ReAct pioneered is built into their inference pipelines.

  5. 2025

    Agentic AI — autonomous multi-step agents

    Computer-using agents, coding agents, and research agents all trace their lineage to ReAct's core insight: interleave reasoning with environment interaction.

ReAct's deepest contribution is not a model or an algorithm — it is a pattern. The think-act-observe loop is now so fundamental that it is hard to imagine building an AI agent without it. Every time Claude uses a tool, every time an AI agent searches the web, browses a page, or executes code, the ReAct pattern is at work: reason about what to do, do it, observe the result, then reason again.

The idea in code

The ReAct loop, simplifiedpython

Simplified to show the idea — not the real implementation.

def react(question, llm, env, max_steps=7):
    """The ReAct loop: think, act, observe, repeat."""
    context = f"Question: {question}\n"

    for step in range(1, max_steps + 1):
        # The LLM generates BOTH a thought and an action
        output = llm.generate(context)  # e.g. "Thought: I need to search X\nAction: search[X]"
        thought, action = parse(output)

        # Thought updates context (free — no env interaction)
        context += f"Thought {step}: {thought}\n"

        # Action interacts with environment → produces observation
        observation = env.execute(action)     # e.g. Wikipedia search result
        context += f"Action {step}: {action}\n"
        context += f"Observation {step}: {observation}\n"

        if action.startswith("finish"):
            return extract_answer(action)

    return "Could not find answer within step limit"

# That's it. Think, act, observe, repeat.
# This loop — or its descendants — runs inside every modern AI agent.

CitationYao, Zhao, Yu, Du, Shafran, Narasimhan, Cao. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR, 2023.

Terms in this paper