Reinforcement Learning2023intermediate12 min read

Voyager: An Open-Ended Embodied Agent with Large Language Models

Voyager: وكيل مُجسَّد مفتوح النهاية يعمل بالنماذج اللغوية الكبيرة

Wang, G. · Xie, Y. · Jiang, Y. · Mandlekar, A. · Xiao, C. · Zhu, Y. · Fan, L. · Anandkumar, A. — NeurIPS Workshop

The problem

Building generally capable embodied agents that continuously explore, plan, and develop new skills in open-ended worlds is a grand challenge. Classical agents operate on primitive low-level actions, making systematic difficult and generalization poor. Recent LLM-based agents can generate action plans, but they are not lifelong learners — they cannot progressively acquire, accumulate, and transfer knowledge over extended time spans. They forget old skills when learning new ones (), cannot compose simple skills into complex behaviors, and struggle to generalize to new environments.

The contribution

Voyager: the first LLM-powered embodied . It introduces three key components: (1) an driven by GPT-4 that proposes tasks matching the agent's skill level and world state, maximizing exploration; (2) a that stores mastered behaviors as executable code indexed by -based retrieval, enabling compositional reuse and avoiding catastrophic forgetting; and (3) an mechanism with , execution error correction, and . Voyager uses code as its action space, interacts with GPT-4 via black-box queries with no fine-tuning, obtains 3.3× more unique items, travels 2.3× longer distances, and unlocks tech tree milestones up to 15.3× faster than prior methods in Minecraft.

The impact

Voyager established the paradigm of LLM-as-agent with persistent skill memory. It showed that code-as-action is superior to low-level motor commands for embodied agents, that self-verifying program synthesis enables lifelong learning without catastrophic forgetting, and that LLM-powered curricula can drive . Voyager inspired a wave of agentic AI systems — from coding assistants to robotic planners — that use LLMs as reasoning engines with tool-augmented, memory-persistent architectures.

Imagine a new employee at a huge company. On day one, they're given no instructions — just dropped into the building. A traditional AI agent would wander the hallways randomly, sometimes re-entering rooms it already visited, and if it learned to use the coffee machine, it would forget how once it learned to use the printer.

Voyager is different. It has a personal mentor (the automatic curriculum) who looks at what it has learned so far and says: "You've mastered the basics — time to try the advanced lab." It carries a notebook (the skill library) where every mastered task is written as a clear recipe, so it never forgets and can combine old recipes into new ones. And it has a mirror (self-verification) — after each attempt, it checks: "Did I actually complete the task, or am I fooling myself?"

The challenge: lifelong learning in open-ended worlds

Minecraft is unlike most environments studied in AI. It has no predefined end goal, no fixed storyline — just an infinite procedurally generated world with endless possibilities. An effective agent in such a world needs three capabilities that existing approaches lack:

Propose suitable tasks based on its current skill level and world state. If it spawns in a desert, it should learn to harvest sand and cactus before attempting iron tools — not follow a rigid curriculum designed for a forest.

Refine and remember skills. When a program to fight zombies fails, the agent should debug it using environment feedback. Once mastered, the skill should be stored permanently and reusable for similar tasks (fighting spiders uses similar logic).

Explore continuously in a self-driven manner, seeking out new challenges without waiting for human instructions.

Classical reinforcement learning agents struggle with all three. They operate on low-level actions (move forward, turn left), making it hard to build temporally extended behaviors. Recent LLM-based agents like ReAct and Reflexion can reason and reflect, but they lack persistent memory — each episode starts from scratch.

Open in Lab
Voyager's three-component architecture: an automatic curriculum proposes tasks, a skill library stores mastered code, and iterative prompting refines programs using environment feedback.
The demo wakes as you arrive…

Component 1: the automatic curriculum

The automatic curriculum is Voyager's internal teacher. At each step, GPT-4 receives the agent's current state — its inventory, position, biome, health, time of day, completed tasks, and failed tasks — and generates the next task to attempt. The overarching directive is simple: "discover as many diverse things as possible."

This is not a hand-designed curriculum. GPT-4 uses to reason about what the agent should try next, given what it already knows and what surrounds it. If the agent is near a river and has never fished, the curriculum might propose "catch a fish." If it has wood tools but no stone tools, it might propose "mine cobblestone." The curriculum adapts to the agent's actual trajectory rather than following a fixed script.

Additional context comes from a GPT-3.5-based question-answering module backed by the Minecraft Wiki. When GPT-4 generates a task, GPT-3.5 retrieves relevant crafting recipes, mob behavior, or biome information, grounding the curriculum in factual game knowledge. This two-tier LLM design separates high-level reasoning (GPT-4) from factual retrieval (GPT-3.5), balancing cost and capability.

Open in Lab
See how the automatic curriculum adapts task proposals to the agent's inventory, biome, and exploration history.
The demo wakes as you arrive…

Component 2: the skill library

The skill library is Voyager's long-term memory. Every time the agent successfully completes a task, the corresponding program is stored as a reusable skill. Each skill consists of executable JavaScript code (using the Mineflayer API) and a natural language description generated by GPT-4.

Storage: Skills are indexed by the embedding of their descriptions. When GPT-4 writes a new program, it also generates a one-line description like "mine 3 cobblestone blocks" or "craft a wooden pickaxe." This description is embedded using a text embedding model, and the resulting vector is stored alongside the code in a vector database.

Retrieval: When the curriculum proposes a new task, Voyager queries the skill library with the new task description, retrieving the top-5 most similar skills. These retrieved programs are injected into the as examples, so GPT-4 doesn't start from scratch — it can reuse, adapt, or compose existing skills.

: This is where the skill library becomes powerful. A skill for "craft a wooden pickaxe" can call the skill for "get 3 planks," which in turn calls "chop a tree." Complex behaviors emerge from composing simpler ones — like building blocks. This compositional structure means Voyager's capabilities grow super-linearly: each new skill can leverage all previous skills.

No catastrophic forgetting: Because skills are stored as code in an external database rather than encoded in neural network weights, Voyager never forgets. A skill learned in the first hour is just as accessible in the hundredth hour. This is a fundamental advantage over agents that store knowledge in model parameters.

Open in Lab
Explore how skills are stored, retrieved, and composed. Click a task to see which existing skills are retrieved and how new skills build on old ones.
The demo wakes as you arrive…

Component 3: iterative prompting with self-verification

When GPT-4 generates code for a new task, the program rarely works on the first try. Voyager uses an iterative refinement loop with three feedback channels:

Environment feedback: After executing the code, the game environment returns the agent's new state — updated inventory, position, health, nearby blocks and entities. If the task was "mine 3 iron ore" and the inventory shows only 1 iron ore, GPT-4 sees the gap and revises the code.

Execution errors: If the code throws a runtime error (e.g., trying to craft an "acacia axe" that doesn't exist in Minecraft), the error message is fed back to GPT-4, which debugs the code and produces a corrected version.

Self-verification: This is a separate GPT-4 call that acts as a critic. Given the task description and the agent's current state after execution, the verifier decides whether the task was truly completed. If not, it provides a critique explaining what went wrong and how to fix it. Only after self-verification passes is the skill committed to the library.

This three-channel feedback loop runs for up to 4 iterations per task. If the task still fails after all retries, it's added to a "failed tasks" list that the curriculum uses to avoid proposing tasks that are currently too difficult.

Open in Lab
Step through an iterative prompting cycle: initial code → environment feedback → error correction → self-verification → skill committed.
The demo wakes as you arrive…

Code as action space: why programs beat motor commands

A critical design choice in Voyager is representing actions as executable code rather than low-level motor commands. In reinforcement learning, an agent typically acts through primitive operations — move forward, turn, click. Building complex behaviors from these primitives requires hundreds of sequential decisions, and the agent must re-learn the entire sequence for every new task.

Voyager writes JavaScript programs that call the Mineflayer bot API. A single "skill" program can contain loops, conditionals, function calls, and error handling — all expressed in a few lines of code. For example, "mine 3 iron ore" is a complete program that navigates to iron ore, mines it, checks the inventory, and repeats until 3 are collected.

This has three major advantages. Programs are temporally extended: one skill can represent minutes of gameplay. Programs are interpretable: a human can read the code and understand exactly what the agent is doing. And programs are compositional: complex skills naturally call simpler skills as subroutines.

Example Voyager skill: crafting a stone pickaxejavascript

Simplified to show the idea — not the real implementation.

// Skill: craftStonePickaxe
// Description: Craft a stone pickaxe using cobblestone and sticks
async function craftStonePickaxe(bot) {
  // Step 1: Check if we have enough materials
  const cobblestone = bot.inventory.count("cobblestone");
  const sticks = bot.inventory.count("stick");

  // Step 2: Gather missing materials by calling existing skills
  if (cobblestone < 3) {
    await mineCobblestone(bot, 3 - cobblestone);  // reuse existing skill
  }
  if (sticks < 2) {
    await craftSticks(bot, 2 - sticks);           // reuse existing skill
  }

  // Step 3: Find or place a crafting table
  const craftingTable = bot.findBlock("crafting_table", 32);
  if (!craftingTable) {
    await placeCraftingTable(bot);                 // reuse existing skill
  }

  // Step 4: Craft the pickaxe
  await bot.craft("stone_pickaxe", 1);
  bot.chat("Stone pickaxe crafted successfully!");
}

Results: Voyager vs. baselines

Voyager was evaluated against three baselines — ReAct, Reflexion, and AutoGPT — across four dimensions in Minecraft. The results demonstrate the compounding power of lifelong learning with a skill library:

Exploration: Voyager discovered 63 unique items in 160 prompting iterations, 3.3× more than the best baseline. While baselines plateau quickly (repeating known behaviors), Voyager keeps discovering because its curriculum pushes toward novelty and its skill library enables increasingly complex actions.

Tech tree mastery: The Minecraft tech tree requires crafting tools in a hierarchy — wooden → stone → iron → diamond. Voyager unlocked the wooden level 15.3× faster, the stone level 8.5× faster, the iron level 6.4× faster, and was the only agent to reach the diamond level. Baselines could not compose the multi-step skills needed.

Map coverage: Voyager traversed 2.3× longer distances than baselines, exploring diverse biomes (deserts, forests, mountains) while baselines remained confined to local areas.

Zero-shot generalization: When placed in a completely new world with an empty inventory, Voyager used its skill library to solve novel tasks (e.g., "obtain a diamond pickaxe") that it had never seen during training. Baselines could not solve any of these tasks within 50 iterations. Remarkably, even giving the Voyager skill library to AutoGPT boosted its performance, showing the library works as a plug-and-play asset.

Open in Lab
Compare Voyager against baselines across exploration, tech tree mastery, and zero-shot generalization.
The demo wakes as you arrive…

Ablation studies: every component matters

The ablation studies reveal that each of Voyager's three components contributes significantly to its performance:

Removing the automatic curriculum and replacing it with a fixed task list causes a sharp drop — the agent can't adapt to its environment and wastes iterations on tasks that don't match its current state.

Removing the skill library means the agent starts from scratch every time. Without the ability to retrieve and compose previously mastered programs, complex tasks become unreachable.

Removing self-verification leads to a library polluted with buggy, incomplete programs. The agent adds skills that don't actually work, and subsequent tasks that compose those skills fail in cascade.

Additionally, replacing GPT-4 with GPT-3.5 for significantly degrades performance, showing that the quality of the underlying language model is crucial for generating correct, complex programs.

The bigger picture: from Minecraft to the real world

  1. 2022

    ReAct — Reasoning + Acting

    LLMs interleave reasoning traces with actions in an environment. Showed that thinking out loud improves task performance. But no memory across episodes.

  2. 2023

    Reflexion — Learning from Mistakes

    Added verbal self-reflection to LLM agents: after failure, the agent writes a textual analysis of what went wrong. Improved one-episode performance but still no persistent skill storage.

  3. 2023

    Voyager — Lifelong Learning Agent

    Introduced persistent skill library, automatic curriculum, and self-verification. First LLM agent to demonstrate lifelong learning with zero-shot transfer to new worlds.

  4. 2023

    Generative Agents — Believable Human Behavior

    Extended LLM agents with memory streams, reflection, and planning for social simulation. Shared Voyager's insight that persistent memory transforms agent capabilities.

Voyager's deepest contribution is architectural, not domain-specific. The three-component pattern — a curriculum that proposes tasks, a library that stores solutions, and a verifier that ensures quality — is a general blueprint for lifelong learning agents. The same architecture could power a coding assistant that builds up a library of reusable code patterns, a robotic system that accumulates manipulation skills, or a research agent that composes investigative strategies.

The key insight is separating knowledge storage from model parameters. By externalizing skills as code in a retrievable database, Voyager achieves something neural networks alone cannot: it never forgets, it always compounds, and its knowledge is transferable across environments.

CitationWang, Xie, Jiang, Mandlekar, Xiao, Zhu, Fan, Anandkumar. Voyager: An Open-Ended Embodied Agent with Large Language Models. NeurIPS Workshop (IMOL), 2023.

Terms in this paper