Reinforcement Learning2020advanced13 min read

MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

MuZero: إتقان ألعاب Atari وGo والشطرنج وShogi بالتخطيط عبر نموذج مُتعلَّم

Schrittwieser, J. · Antonoglou, I. · Hubert, T. · Simonyan, K. · Sifre, L. · Schmitt, S. · Guez, A. · Lockhart, E. · Hassabis, D. · Graepel, T. · Lillicrap, T. · Silver, D. — Nature

The problem

algorithms like AlphaZero achieved superhuman performance in board games, but they required a perfect simulator — a complete, rule-based model of the that could predict the exact next state for any action. This works for chess and Go, where the rules are known and deterministic. But most real-world problems — video games, robotics, autonomous driving — have complex, unknown dynamics. Traditional tried to learn full environment models that reconstruct raw observations (pixel-level predictions), but these models were too inaccurate to support effective planning. The gap between "knows the rules" and "learns from raw experience" remained a fundamental barrier.

The contribution

MuZero bridges this gap with three learned neural networks: a function that encodes raw observations into a compact , a dynamics function that predicts the next hidden state and reward given an action, and a prediction function that outputs a policy and value estimate from any hidden state. Crucially, these networks are trained end-to-end to predict only what matters for planning — reward, value, and policy — not raw pixel reconstructions. Combined with , MuZero matched AlphaZero's superhuman performance in Go, chess, and shogi without knowing any game rules, and simultaneously achieved state-of-the-art on 57 Atari games.

The impact

MuZero unified model-free and model-based , showing that a single algorithm can master both perfect-information board games and visually complex video games. It proved that learning an abstract, planning-relevant model is more effective than learning a full environment simulator. This idea directly influenced AlphaTensor (discovering faster matrix multiplication algorithms), Dreamer-V3 (general world models), and the broader shift toward learned world models in AI. MuZero also demonstrated practical applications beyond games, including video compression optimization.

Imagine a chess grandmaster who has never read a rulebook. Instead of memorizing that "bishops move diagonally," she plays millions of games and develops an intuition: "from this position, moving that piece usually leads to winning." She doesn't simulate every possible board state in her head — she imagines just the aspects that matter for her next decision.

Now imagine she can do this not only for chess, but for any game you hand her — Go, shogi, or even a chaotic Atari game with flashing pixels. She learns a different mental model for each, but always the same kind of model: one that predicts what matters for planning, not every detail of the world.

MuZero is this grandmaster. It learns to plan in any environment by building an internal model that predicts only what's useful — rewards, good moves, and future prospects — without ever being told the rules.

The problem: why game rules are a luxury

Before MuZero, the most powerful planning algorithms fell into two camps that could not reach each other.

On one side stood AlphaZero, which combined deep neural networks with Monte Carlo Tree Search to achieve superhuman performance in Go, chess, and shogi. But AlphaZero had a critical dependency: it needed a perfect simulator — a function that, given any board state and any legal move, returns the exact next state. For board games with known rules, this is trivial. For an Atari game rendered as pixels, or a robot navigating a cluttered room, no such function exists.

On the other side stood model-free methods like DQN and R2D2, which learned directly from raw experience without any model of the environment. These worked on Atari and other complex domains, but they could not plan ahead — they made decisions reactively, one step at a time, without imagining future consequences.

Previous attempts at model-based RL tried to learn a full environment simulator — predicting every pixel of the next frame. But pixel-level prediction is enormously difficult, and small errors compound over multiple steps, making the learned model too unreliable for planning.

Open in Lab
Compare the two paradigms: AlphaZero uses a perfect simulator (left), while model-free methods react without planning (right). MuZero bridges the gap with a learned model.
The demo wakes as you arrive…

MuZero's three functions: representation, dynamics, and prediction

MuZero's architecture rests on three neural networks that work together like a team of specialists. Think of them as three departments in a planning agency:

The Representation Function hh is the translator. It takes the raw from the environment — a Go board position, an Atari screen of pixels — and compresses it into a compact hidden state s0s^0. This is like translating a photograph into a concise description that captures everything strategically relevant while discarding irrelevant visual noise. Formally: s0=h(o1,…,ot)s^0 = h(o_1, \ldots, o_t).

The Dynamics Function gg is the simulator. Given a hidden state and an action, it predicts the next hidden state and the immediate reward. This is MuZero's learned substitute for a game engine — it imagines what would happen next without actually taking the action in the real environment. Think of it as a mental rehearsal: "If I move my knight here, the position becomes..." Formally: rk,sk=g(sk−1,ak)r^k, s^k = g(s^{k-1}, a^k).

The Prediction Function ff is the evaluator. Given any hidden state, it estimates how good the position is (value) and which actions are most promising (policy). This is the intuition that guides the search: "This position looks strong, and the most promising move is..." Formally: pk,vk=f(sk)p^k, v^k = f(s^k).

Open in Lab
Click each function to see how observation flows through representation → dynamics → prediction. Toggle between board games and Atari to see the same architecture at work.
The demo wakes as you arrive…

The critical design choice is that the hidden state sks^k has no requirement to reconstruct the original observation. Unlike traditional world models that try to predict future frames pixel by pixel, MuZero's hidden state is free to represent whatever is most useful for predicting reward, value, and policy. The signal comes entirely from planning-relevant targets, not from reconstruction loss.

This is the equivalent of asking a chess player to describe a position. A novice might photograph the board; an expert says "White has a passed pawn on d6 with connected rooks." The expert's description is lossy — you cannot reconstruct the exact board from it — but it captures everything needed for planning.

Planning with MCTS: thinking ahead with imagination

At each decision point, MuZero does not simply pick the action with the highest predicted value. Instead, it runs Monte Carlo Tree Search (MCTS) — a systematic that builds a of possible futures using the learned model.

The process has three phases that repeat many times (typically hundreds of simulations per move):

Selection. Starting from the current state at the root, the algorithm walks down the tree, choosing actions that balance (trying less-visited actions) with (favoring actions with high estimated value). The selection criterion uses the PUCT formula, which combines the predicted value with a bonus for less-explored branches, weighted by the 's prior probability.

Expansion. When the walk reaches a leaf node — a position not yet explored — MuZero uses the dynamics function gg to imagine the next state and reward, then the prediction function ff to evaluate it. The new node is added to the tree with its policy and value estimates.

Backup. The value estimate is propagated back up the tree, updating the statistics of every node along the path. After many simulations, the visit counts at the root reflect how promising each action is. The final action is chosen proportionally to these visit counts.

Open in Lab
Watch MCTS build the search tree step by step. Click "Simulate" to run one selection → expansion → backup cycle. The node sizes reflect visit counts.
The demo wakes as you arrive…
ak=arg⁡max⁡a[Q(s,a)+c(s)⋅P(s,a)⋅N(s)1+N(s,a)]a^k = \arg\max_a \left[ Q(s,a) + c(s) \cdot P(s,a) \cdot \frac{\sqrt{N(s)}}{1 + N(s,a)} \right]
PUCT selection formula — balancing exploitation and exploration — Q(s,a)Q(s,a) is the mean value of action aa from state ss (exploitation). P(s,a)P(s,a) is the prior probability from the policy network. N(s)N(s) and N(s,a)N(s,a) are visit counts. The fraction gives a bonus to less-visited actions (exploration). The constant c(s)c(s) controls the exploration-exploitation trade-off, adapting to the scale of values.

Training: learning from self-play and imagination

MuZero's training loop has two interleaved processes: acting and learning.

During acting, the plays games (or Atari episodes) using MCTS to select actions. Each game generates a of observations, actions, rewards, MCTS policies, and MCTS value estimates. These trajectories are stored in a prioritized , much like in DQN.

During learning, the algorithm samples trajectories from the replay buffer and unrolls the learned model for KK hypothetical steps. Starting from a real observation oto_t, it applies the representation function to get s0s^0, then repeatedly applies the dynamics function with the actual actions taken in the trajectory to generate predicted states s1,s2,…,sKs^1, s^2, \ldots, s^K. At each step kk, three losses are computed.

ℓt(θ)=∑k=0K[lp(πt+k, ptk)⏟policy+lv(zt+k, vtk)⏟value+lr(ut+k, rtk)⏟reward]+c∥θ∥2\ell_t(\theta) = \sum_{k=0}^{K} \Big[ \underbrace{l^p(\pi_{t+k},\, p^k_t)}_{\text{policy}} + \underbrace{l^v(z_{t+k},\, v^k_t)}_{\text{value}} + \underbrace{l^r(u_{t+k},\, r^k_t)}_{\text{reward}} \Big] + c \|\theta\|^2
MuZero total loss — policy, value, and reward across K unrolled steps — At each unrolled step kk: the **policy loss** lpl^p matches the predicted policy pkp^k to the MCTS search policy πt+k\pi_{t+k} (cross-entropy). The **value loss** lvl^v matches the predicted value vkv^k to an nn-step bootstrapped target zt+kz_{t+k}. The **reward loss** lrl^r matches the predicted reward rkr^k to the actual observed reward ut+ku_{t+k}. All three losses are summed across KK steps, plus an L2 regularization term.
Open in Lab
Watch the training unroll: a trajectory from the replay buffer is unrolled K steps through the learned model, computing losses at each step.
The demo wakes as you arrive…

A critical detail: the value target uses nn-step . For board games, MuZero bootstraps all the way to the end of the game (the actual win/loss outcome). For Atari, it uses n=10n = 10 steps of observed rewards plus a discounted value estimate from the search:

zt=ut+1+γut+2+⋯+γn−1ut+n+γnvt+nz_t = u_{t+1} + \gamma u_{t+2} + \cdots + \gamma^{n-1} u_{t+n} + \gamma^n v_{t+n}

To maintain stable gradients across the KK unrolled steps, MuZero scales the at each step by 1/K1/K. This ensures that steps further from the root do not dominate the training signal.

Results: matching AlphaZero without rules, beating Atari with planning

MuZero was evaluated on two very different classes of domains, demonstrating its generality.

Board games (Go, chess, shogi). MuZero matched the performance of AlphaZero in all three games — despite never being given the game rules. In Go, it achieved an Elo rating within a few points of AlphaZero. It used the same amount of computation per move (800 simulations of MCTS) but replaced the perfect simulator with its learned dynamics function.

Atari (57 games). MuZero achieved a new state of the art on the Atari benchmark, outperforming the previous best model-free method (R2D2). This was particularly remarkable because model-based methods had historically struggled on Atari — the pixel-level observations and diverse game mechanics made it far harder than board games. MuZero showed that a learned abstract model, combined with planning, could outperform even the best reactive approaches.

Open in Lab
Compare MuZero's performance against AlphaZero (board games) and R2D2 (Atari). Toggle between domains to see the results.
The demo wakes as you arrive…

Ablation: what each piece contributes

The authors ran ablation studies to show which components matter most:

Search at action time. Removing MCTS and using only the raw policy network degraded performance significantly, especially on board games. Search provides a systematic way to improve beyond the current policy.

Learned model vs. perfect model. Replacing the learned dynamics function with a perfect simulator (where available) showed minimal improvement on board games — the learned model was already accurate enough. This confirmed that the abstract learned model captures the relevant dynamics.

Number of unroll steps (K). Training with K=5K = 5 unrolled steps worked well across domains. Too few steps gave insufficient learning signal; too many led to compounding errors in the dynamics model.

Code: MuZero pseudocode

MuZero MCTS — one simulation steppython

Simplified to show the idea — not the real implementation.

# One MCTS simulation: select → expand → backup
def run_simulation(root, model):
    node = root
    path = [node]

    # ── Selection: walk the tree using PUCT ──
    while node.is_expanded():
        action, child = select_child(node)  # PUCT formula
        node = child
        path.append(node)

    # ── Expansion: use learned model to imagine next state ──
    parent = path[-2]
    hidden_state = parent.hidden_state
    # Dynamics function: predict next state + reward
    next_state, reward = model.dynamics(hidden_state, action)
    # Prediction function: evaluate the new state
    policy, value = model.prediction(next_state)
    node.expand(next_state, reward, policy)

    # ── Backup: propagate value up the tree ──
    for node in reversed(path):
        node.visit_count += 1
        node.value_sum += value
        value = node.reward + discount * value  # bootstrap

The big picture: from game rules to learned world models

MuZero represents a fundamental shift in how we think about planning in AI. Previous systems drew a hard line: either you have a perfect model of the world (chess engines, AlphaZero), or you learn without one (DQN, policy gradient methods). MuZero showed that you can learn a model that is good enough for planning — and that this learned model does not need to reconstruct reality, only to predict the consequences of actions in a way that supports good decisions.

This insight extends far beyond games. Any domain where an agent must make sequential decisions — logistics, drug discovery, chip design, resource allocation — could benefit from the same approach: learn a compact predictive model, then plan with it using search. The model does not need to be photorealistic; it needs to be decision-relevant.

Timeline: from AlphaGo to learned world models

  1. 2016

    AlphaGo

    First program to defeat a human Go champion, using MCTS with neural network evaluation — but required training on human expert games and a perfect simulator.

  2. 2017

    AlphaGo Zero

    Learned Go entirely through self-play with no human data — tabula rasa. Still required the perfect simulator for MCTS.

  3. 2018

    AlphaZero

    Generalized self-play learning to chess, shogi, and Go with a single algorithm. Still required knowing the game rules for simulation.

  4. 2020

    MuZero (this paper)

    Replaced the perfect simulator with a learned model. Matched AlphaZero on board games and achieved state-of-the-art on Atari — all without knowing any rules.

  5. 2021

    MuZero Reanalyse

    Extended MuZero with offline reanalysis — reprocessing stored trajectories with the latest model to improve sample efficiency further.

  6. 2022

    AlphaTensor

    Applied MuZero-style planning to discover faster matrix multiplication algorithms, extending the framework beyond games into mathematical discovery.

  7. 2023

    Dreamer-V3

    General world model that learns across diverse domains. Shares MuZero's philosophy of learning abstract dynamics models, but uses imagination-based rollouts instead of MCTS.

MuZero's legacy is the demonstration that planning with a learned model can match or exceed planning with a perfect model. This resolved a decades-old debate in reinforcement learning: you do not need to choose between model-based and model-free approaches. You can learn the model and plan with it — as long as you train the model to predict what planning actually needs.

CitationSchrittwieser, Antonoglou, Hubert, Simonyan, Sifre, Schmitt, Guez, Lockhart, Hassabis, Graepel, Lillicrap, Silver. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model. Nature, 2020.

Terms in this paper