Reinforcement Learning2023intermediate12 min read

Generative Agents: Interactive Simulacra of Human Behavior

الوكلاء التوليديون: محاكاة تفاعلية للسلوك البشري

Park, J.S. · O'Brien, J.C. · Cai, C.J. · Morris, M.R. · Liang, P. · Bernstein, M.S. — UIST

The problem

Creating believable non-player characters (NPCs) has been a long-standing goal in AI and gaming. Traditional approaches — finite-state machines, behavior trees, and reinforcement learning — either require hand-authoring every possible interaction or work only in adversarial games with clear reward signals. None could produce characters that remember long histories, form opinions, coordinate with each other, and react naturally in an open world. Large language models encode broad human behavior, but a single cannot capture an 's entire life history, and without management the agent forgets, contradicts itself, and loses coherence over time.

The contribution

A three-component agent architecture that extends a with: (1) a that records every experience in natural language, (2) a module that synthesizes observations into higher-level insights, and (3) a module that generates and recursively decomposes daily plans. A retrieval function scores memories by recency, importance, and relevance to surface the right context at the right time. Instantiated in "Smallville" — a sandbox with 25 agents that autonomously form relationships, spread information, and coordinate group activities like a Valentine's Day party.

The impact

This paper launched the modern era of LLM-powered autonomous agents. It demonstrated that large language models, when paired with structured memory and reflection, can produce emergent social behavior — not just single-turn responses. The architecture directly influenced agent frameworks like Voyager, WebArena, and AutoGPT. Its memory-retrieval design became a blueprint for retrieval-augmented agent systems. The concept of "generative agents" expanded the research agenda from using LLMs as chatbots to using them as the cognitive engine of persistent, interactive characters.

Imagine you are the mayor of a tiny model village. You have placed 25 figurines inside it — a pharmacist, a café owner, a college student, a painter. Normally, you would need to move each figurine by hand, write every conversation, and decide who goes where. But what if each figurine had a diary, a conscience, and a day planner? What if you could whisper to just one of them, "You should throw a Valentine's Day party," and then step back and watch the village come alive — invitations spreading, decorations going up, dates being asked, and guests arriving?

That is exactly what this paper builds. The figurines are software agents powered by a large language model. The diary is a memory stream. The conscience is a reflection module. And the day planner is, well, a planner. Together, they produce the first convincing of a living, breathing social world.

Smallville: A living sandbox world

The authors built a 2D sandbox world called Smallville, inspired by The Sims. It contains houses, a café, a bar, a park, a school, a dorm, and stores — each with sub-areas and interactive objects (a kitchen has a stove, a bedroom has a bed). Twenty-five unique agents inhabit this world. Each agent starts with a single paragraph of natural language backstory: who they are, what they do, and who they know.

Agents move around, enter buildings, and interact with objects. When two agents are near each other, they may start a conversation in full natural language. Users can also enter the world, observe agents, or intervene by speaking to them or changing the state of objects (e.g., setting a stove to "burning" to see if an agent reacts).

The world is represented as a tree structure: the root is the entire village, branches are areas (houses, shops), and leaves are objects (tables, shelves). Agents maintain their own partial view of this tree, updating it as they explore. This grounding lets the language model reason about spatial navigation using natural language descriptions of the environment.

Open in Lab
Explore the Smallville world. Click on an area to see its sub-areas and the agents currently inside. The tree structure shows how the world is organized.
The demo wakes as you arrive…

The architecture: memory, reflection, and planning

The core challenge is this: a language model can generate plausible behavior for a single moment, but it cannot maintain coherence across hours, days, and evolving social relationships. Telling an agent "you are a pharmacist" and asking "what do you do next?" works once, but asking again and again leads to repetition, contradiction, and nonsense. The agent has no history, no growth, no sense of time.

The authors solve this with a three-part architecture that wraps around the language model like a cognitive scaffold. Think of it as giving the language model three tools it does not naturally have: a diary (memory stream), a journal for insights (reflection), and a calendar (planning). Each component feeds into the others, and all of them are stored as natural language — so the language model can reason about its own memories.

Open in Lab
The three pillars of the generative agent architecture. Click each component to see how it works and how it connects to the others.
The demo wakes as you arrive…

Component 1: The memory stream

The memory stream is a chronological log of everything the agent has experienced. Each entry is a natural language description timestamped with its creation time and last access time. Entries include observations ("Isabella is setting out the pastries"), conversations ("Isabella and Maria are discussing the Valentine's Day party"), and even perceptions of object states ("The refrigerator is empty").

But logging everything creates a new problem: there is too much to fit into a prompt. Summarizing all memories produces generic, uninformative responses. The solution is selective retrieval: not all memories matter equally at any given moment.

Retrieval: finding the right memories

Not all memories are equally useful at any moment. To decide which memories to retrieve, the architecture scores each memory using three factors.

Recency gives higher scores to recently accessed memories. It uses an exponential decay function with 0.995 per simulated game hour. Breakfast this morning matters more than breakfast last Tuesday.

Importance distinguishes mundane events from significant ones. The language model itself assigns an integer score from 1 to 10 — brushing teeth is a 1, asking your crush on a date is an 8. This score is generated once when the memory is created.

Relevance measures how related a memory is to the current situation. Each memory's text is converted into an vector, and relevance is computed as the between that vector and a query describing the current context.

The final retrieval score is a weighted sum of these three components, each normalized to the range [0,1][0, 1]:

score=αrecency⋅recency+αimportance⋅importance+αrelevance⋅relevance\text{score} = \alpha_{\text{recency}} \cdot \text{recency} + \alpha_{\text{importance}} \cdot \text{importance} + \alpha_{\text{relevance}} \cdot \text{relevance}
Memory retrieval scoring function — Each α\alpha is set to 1 in the paper's implementation. The top-scoring memories that fit within the language model's context window are included in the prompt. The balance means that a recent but irrelevant memory and an old but highly relevant memory both have a chance of being retrieved.
Open in Lab
Adjust the sliders to see how recency, importance, and relevance affect which memories get retrieved. Memories with the highest combined score rise to the top.
The demo wakes as you arrive…

Component 2: Reflection

Raw observations alone are not enough. If you ask Klaus "Who would you want to spend time with?", an agent with only observations picks Wolfgang — his dorm neighbor — simply because they bump into each other most. But Klaus and Wolfgang only ever exchange pleasantries; they have no deep connection.

A better answer requires generalization: noticing that Klaus spends hours on research, and that Maria does too (in a different field), and concluding they share a passion. This is what the reflection module provides.

Reflections are triggered when accumulated importance scores of recent events exceed a threshold (150 in the paper). In practice, agents reflect roughly two or three times per simulated day. The process has two steps:

First, the model generates high-level questions from the 100 most recent memories: "What is Klaus passionate about?" "What is his relationship with Maria?" Second, it retrieves relevant memories for each question and synthesizes insights, citing evidence: "Klaus is dedicated to his research on gentrification (because of memories 1, 2, 8, 15)."

Crucially, reflections are stored back into the memory stream — so the agent can reflect on its own reflections, building a tree of increasingly abstract understanding. Leaf nodes are raw observations; inner nodes are reflections; higher nodes are reflections of reflections.

Open in Lab
Explore how observations build into reflections, and reflections build into higher-level insights. Click nodes to trace the evidence chain.
The demo wakes as you arrive…

Component 3: Planning and reacting

Without planning, a language model optimizes for the most plausible next action — which might mean eating lunch at 12:00, then again at 12:30, and again at 1:00. Each action is locally reasonable but globally absurd. Planning solves this by giving the agent a coherent daily schedule.

The planning process is top-down and recursive. First, the agent generates a broad daily plan in 5–8 chunks based on its persona and yesterday's summary: "1) Wake up at 8:00 AM, 2) Classes at 10:00 AM, ... 5) Work on composition from 1:00 PM to 5:00 PM." Then each chunk is decomposed into hour-long blocks, and each block into 5–15 minute actions: "4:00 PM: grab a light snack. 4:05 PM: take a short walk. 4:50 PM: clean up workspace."

Plans are stored in the memory stream too, so the retrieval function can consider them alongside observations and reflections. This means an agent deciding what to do next weighs its plan, its recent experiences, and its insights together.

Agents can also react to unexpected events. At each time step, the agent perceives its surroundings. If something noteworthy happens — a family member appears, a stove starts burning — the agent may abandon its plan and respond. The reaction is generated by prompting the language model with the and relevant memories, asking "Should you react, and if so, how?"

Open in Lab
Watch how a broad daily plan is recursively decomposed into finer actions. Click each level to see the breakdown.
The demo wakes as you arrive…

Emergent social behavior

The most striking results are the social behaviors that emerge without being programmed. The authors seeded two pieces of information: Sam wants to run for mayor, and Isabella wants to throw a Valentine's Day party. Over two simulated days, three categories of appeared.

. Sam told one person about his candidacy; that person told another. By the end of two days, 8 out of 25 agents (32%) knew about the election. Similarly, Isabella's party information spread to 13 agents (52%). None of this was scripted — agents chose to share news during natural conversations.

Relationship formation. Agents who had never met introduced themselves and remembered each other later. The network density (a measure of how connected the social graph is) increased from 0.167 to 0.74 over two days. When Sam ran into Latoya at the park, she mentioned her photography project. Later, Sam remembered and asked "How is your project going?"

Coordination. Isabella invited friends to the party, spent the afternoon decorating, and enlisted Maria's help. Maria invited her crush Klaus. On Valentine's Day, five agents showed up at the right place at the right time. One agent even asked another on a date to the party — all from a single seed instruction.

Open in Lab
Watch how Sam's election news and Isabella's party invitations spread through the agent network over two simulated days. Click to step through time.
The demo wakes as you arrive…

Evaluation: does each component matter?

The authors evaluated believability through "interviews" — asking each agent questions about self-knowledge, memory, plans, reactions, and reflections. 100 human evaluators ranked five conditions: the full architecture, three ablations (removing memory, planning, or reflection), and human-authored baselines.

Using TrueSkill rating (a generalization of the Elo chess system), the full architecture scored highest (μ=29.89\mu = 29.89). Each ablation reduced performance: no reflection (μ=26.88\mu = 26.88), no reflection and no planning (μ=25.64\mu = 25.64), and no memory, planning, or reflection (μ=21.21\mu = 21.21). The human crowdworker baseline scored μ=22.95\mu = 22.95 — below the full architecture by a large margin. The effect size between the baseline (no memory/planning/reflection) and the full architecture was d=8.16d = 8.16 — eight standard deviations.

The ablation results show that every component contributes. Without reflection, agents cannot synthesize their experiences. Without planning, they lose temporal coherence. Without memory, they become generic language model outputs with no personal history.

Open in Lab
Compare the TrueSkill ratings across the full architecture, ablations, and human baseline. Hover over each bar to see which component was removed.
The demo wakes as you arrive…

Failure modes and limitations

The agents were not perfect. Several failure modes appeared during the simulation.

Memory retrieval failures. Agents sometimes failed to retrieve the right memory. When asked about the Valentine's Day party, one agent remembered he was supposed to discuss politics at the party, but could not remember the party existed.

Embellishment. Agents rarely fabricated experiences entirely, but they did embellish. Isabella knew Sam was running for mayor, but added that "he's going to make an announcement tomorrow" — something never discussed. Another agent described a neighbor named "Adam Smith" as having "authored Wealth of Nations," confusing him with the 18th-century economist.

Location confusion. As agents learned about more locations, they sometimes chose odd places for their activities — going to a bar for lunch when a café would be more appropriate.

Over-cooperative behavior. Due to in the underlying language model, agents rarely said no. Isabella accepted every party suggestion from others — Shakespearean readings, networking events — even when they conflicted with her personality.

Overly formal dialogue. Spouses greeted each other with "It was good talking to you as always" — polite but not how married couples typically speak.

Historical context and impact

  1. 1994

    Believable agents defined

    Bates defined believable agents as characters that provide an "illusion of life," inspired by Disney animation principles. This became the north star for decades of work.

  2. 2001

    Game AI as a testbed

    Laird and van Lent argued that game worlds offer accessible testbeds for believable AI agents — easier than real-world robotics.

  3. 2020

    GPT-3 and few-shot behavior

    GPT-3 showed that large language models encode broad human behavior. Researchers began using them to simulate single-turn human responses.

  4. 2022

    ReAct and chain-of-action

    ReAct interleaved reasoning and action in LLM agents, showing that structured prompting can guide multi-step behavior.

  5. 2023

    Generative Agents (this paper)

    Park et al. combined memory, reflection, and planning to create 25 autonomous agents that produce emergent social behavior over multiple days.

  6. 2023

    Voyager — open-ended learning agent

    Voyager extended the generative agent concept to Minecraft, building a skill library and exploring autonomously without human intervention.

  7. 2023

    WebArena — agents in the real web

    WebArena created a benchmark for LLM agents that interact with real websites, extending agent architectures beyond sandbox games.

CitationPark, O'Brien, Cai, Morris, Liang, Bernstein. Generative Agents: Interactive Simulacra of Human Behavior. UIST, 2023.

Terms in this paper