Reinforcement Learning2020advanced13 min read
MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
MuZero: إتقان ألعاب Atari وGo والشطرنج وShogi بالتخطيط عبر نموذج مُتعلَّم
Schrittwieser, J. · Antonoglou, I. · Hubert, T. · Simonyan, K. · Sifre, L. · Schmitt, S. · Guez, A. · Lockhart, E. · Hassabis, D. · Graepel, T. · Lillicrap, T. · Silver, D. — Nature
The problem
algorithms like AlphaZero achieved superhuman performance in board games, but they required a perfect simulator — a complete, rule-based model of the that could predict the exact next state for any action. This works for chess and Go, where the rules are known and deterministic. But most real-world problems — video games, robotics, autonomous driving — have complex, unknown dynamics. Traditional tried to learn full environment models that reconstruct raw observations (pixel-level predictions), but these models were too inaccurate to support effective planning. The gap between "knows the rules" and "learns from raw experience" remained a fundamental barrier.
The contribution
MuZero bridges this gap with three learned neural networks: a function that encodes raw observations into a compact , a dynamics function that predicts the next hidden state and reward given an action, and a prediction function that outputs a policy and value estimate from any hidden state. Crucially, these networks are trained end-to-end to predict only what matters for planning — reward, value, and policy — not raw pixel reconstructions. Combined with , MuZero matched AlphaZero's superhuman performance in Go, chess, and shogi without knowing any game rules, and simultaneously achieved state-of-the-art on 57 Atari games.
The impact
MuZero unified model-free and model-based , showing that a single algorithm can master both perfect-information board games and visually complex video games. It proved that learning an abstract, planning-relevant model is more effective than learning a full environment simulator. This idea directly influenced AlphaTensor (discovering faster matrix multiplication algorithms), Dreamer-V3 (general world models), and the broader shift toward learned world models in AI. MuZero also demonstrated practical applications beyond games, including video compression optimization.
Imagine a chess grandmaster who has never read a rulebook. Instead of memorizing that "bishops move diagonally," she plays millions of games and develops an intuition: "from this position, moving that piece usually leads to winning." She doesn't simulate every possible board state in her head — she imagines just the aspects that matter for her next decision.
Now imagine she can do this not only for chess, but for any game you hand her — Go, shogi, or even a chaotic Atari game with flashing pixels. She learns a different mental model for each, but always the same kind of model: one that predicts what matters for planning, not every detail of the world.
MuZero is this grandmaster. It learns to plan in any environment by building an internal model that predicts only what's useful — rewards, good moves, and future prospects — without ever being told the rules.
The problem: why game rules are a luxury
Before MuZero, the most powerful planning algorithms fell into two camps that could not reach each other.
On one side stood AlphaZero, which combined deep neural networks with Monte Carlo Tree Search to achieve superhuman performance in Go, chess, and shogi. But AlphaZero had a critical dependency: it needed a perfect simulator — a function that, given any board state and any legal move, returns the exact next state. For board games with known rules, this is trivial. For an Atari game rendered as pixels, or a robot navigating a cluttered room, no such function exists.
On the other side stood model-free methods like DQN and R2D2, which learned directly from raw experience without any model of the environment. These worked on Atari and other complex domains, but they could not plan ahead — they made decisions reactively, one step at a time, without imagining future consequences.
Previous attempts at model-based RL tried to learn a full environment simulator — predicting every pixel of the next frame. But pixel-level prediction is enormously difficult, and small errors compound over multiple steps, making the learned model too unreliable for planning.
MuZero's three functions: representation, dynamics, and prediction
MuZero's architecture rests on three neural networks that work together like a team of specialists. Think of them as three departments in a planning agency:
The Representation Function is the translator. It takes the raw from the environment — a Go board position, an Atari screen of pixels — and compresses it into a compact hidden state . This is like translating a photograph into a concise description that captures everything strategically relevant while discarding irrelevant visual noise. Formally: .
The Dynamics Function is the simulator. Given a hidden state and an action, it predicts the next hidden state and the immediate reward. This is MuZero's learned substitute for a game engine — it imagines what would happen next without actually taking the action in the real environment. Think of it as a mental rehearsal: "If I move my knight here, the position becomes..." Formally: .
The Prediction Function is the evaluator. Given any hidden state, it estimates how good the position is (value) and which actions are most promising (policy). This is the intuition that guides the search: "This position looks strong, and the most promising move is..." Formally: .
The critical design choice is that the hidden state has no requirement to reconstruct the original observation. Unlike traditional world models that try to predict future frames pixel by pixel, MuZero's hidden state is free to represent whatever is most useful for predicting reward, value, and policy. The signal comes entirely from planning-relevant targets, not from reconstruction loss.
This is the equivalent of asking a chess player to describe a position. A novice might photograph the board; an expert says "White has a passed pawn on d6 with connected rooks." The expert's description is lossy — you cannot reconstruct the exact board from it — but it captures everything needed for planning.
Planning with MCTS: thinking ahead with imagination
At each decision point, MuZero does not simply pick the action with the highest predicted value. Instead, it runs Monte Carlo Tree Search (MCTS) — a systematic that builds a of possible futures using the learned model.
The process has three phases that repeat many times (typically hundreds of simulations per move):
Selection. Starting from the current state at the root, the algorithm walks down the tree, choosing actions that balance (trying less-visited actions) with (favoring actions with high estimated value). The selection criterion uses the PUCT formula, which combines the predicted value with a bonus for less-explored branches, weighted by the 's prior probability.
Expansion. When the walk reaches a leaf node — a position not yet explored — MuZero uses the dynamics function to imagine the next state and reward, then the prediction function to evaluate it. The new node is added to the tree with its policy and value estimates.
Backup. The value estimate is propagated back up the tree, updating the statistics of every node along the path. After many simulations, the visit counts at the root reflect how promising each action is. The final action is chosen proportionally to these visit counts.
Training: learning from self-play and imagination
MuZero's training loop has two interleaved processes: acting and learning.
During acting, the plays games (or Atari episodes) using MCTS to select actions. Each game generates a of observations, actions, rewards, MCTS policies, and MCTS value estimates. These trajectories are stored in a prioritized , much like in DQN.
During learning, the algorithm samples trajectories from the replay buffer and unrolls the learned model for hypothetical steps. Starting from a real observation , it applies the representation function to get , then repeatedly applies the dynamics function with the actual actions taken in the trajectory to generate predicted states . At each step , three losses are computed.
A critical detail: the value target uses -step . For board games, MuZero bootstraps all the way to the end of the game (the actual win/loss outcome). For Atari, it uses steps of observed rewards plus a discounted value estimate from the search:
To maintain stable gradients across the unrolled steps, MuZero scales the at each step by . This ensures that steps further from the root do not dominate the training signal.
Results: matching AlphaZero without rules, beating Atari with planning
MuZero was evaluated on two very different classes of domains, demonstrating its generality.
Board games (Go, chess, shogi). MuZero matched the performance of AlphaZero in all three games — despite never being given the game rules. In Go, it achieved an Elo rating within a few points of AlphaZero. It used the same amount of computation per move (800 simulations of MCTS) but replaced the perfect simulator with its learned dynamics function.
Atari (57 games). MuZero achieved a new state of the art on the Atari benchmark, outperforming the previous best model-free method (R2D2). This was particularly remarkable because model-based methods had historically struggled on Atari — the pixel-level observations and diverse game mechanics made it far harder than board games. MuZero showed that a learned abstract model, combined with planning, could outperform even the best reactive approaches.
Ablation: what each piece contributes
The authors ran ablation studies to show which components matter most:
Search at action time. Removing MCTS and using only the raw policy network degraded performance significantly, especially on board games. Search provides a systematic way to improve beyond the current policy.
Learned model vs. perfect model. Replacing the learned dynamics function with a perfect simulator (where available) showed minimal improvement on board games — the learned model was already accurate enough. This confirmed that the abstract learned model captures the relevant dynamics.
Number of unroll steps (K). Training with unrolled steps worked well across domains. Too few steps gave insufficient learning signal; too many led to compounding errors in the dynamics model.
Code: MuZero pseudocode
Simplified to show the idea — not the real implementation.
# One MCTS simulation: select → expand → backup
def run_simulation(root, model):
node = root
path = [node]
# ── Selection: walk the tree using PUCT ──
while node.is_expanded():
action, child = select_child(node) # PUCT formula
node = child
path.append(node)
# ── Expansion: use learned model to imagine next state ──
parent = path[-2]
hidden_state = parent.hidden_state
# Dynamics function: predict next state + reward
next_state, reward = model.dynamics(hidden_state, action)
# Prediction function: evaluate the new state
policy, value = model.prediction(next_state)
node.expand(next_state, reward, policy)
# ── Backup: propagate value up the tree ──
for node in reversed(path):
node.visit_count += 1
node.value_sum += value
value = node.reward + discount * value # bootstrapThe big picture: from game rules to learned world models
MuZero represents a fundamental shift in how we think about planning in AI. Previous systems drew a hard line: either you have a perfect model of the world (chess engines, AlphaZero), or you learn without one (DQN, policy gradient methods). MuZero showed that you can learn a model that is good enough for planning — and that this learned model does not need to reconstruct reality, only to predict the consequences of actions in a way that supports good decisions.
This insight extends far beyond games. Any domain where an agent must make sequential decisions — logistics, drug discovery, chip design, resource allocation — could benefit from the same approach: learn a compact predictive model, then plan with it using search. The model does not need to be photorealistic; it needs to be decision-relevant.
Timeline: from AlphaGo to learned world models
2016
AlphaGo
First program to defeat a human Go champion, using MCTS with neural network evaluation — but required training on human expert games and a perfect simulator.
2017
AlphaGo Zero
Learned Go entirely through self-play with no human data — tabula rasa. Still required the perfect simulator for MCTS.
2018
AlphaZero
Generalized self-play learning to chess, shogi, and Go with a single algorithm. Still required knowing the game rules for simulation.
2020
MuZero (this paper)
Replaced the perfect simulator with a learned model. Matched AlphaZero on board games and achieved state-of-the-art on Atari — all without knowing any rules.
2021
MuZero Reanalyse
Extended MuZero with offline reanalysis — reprocessing stored trajectories with the latest model to improve sample efficiency further.
2022
AlphaTensor
Applied MuZero-style planning to discover faster matrix multiplication algorithms, extending the framework beyond games into mathematical discovery.
2023
Dreamer-V3
General world model that learns across diverse domains. Shares MuZero's philosophy of learning abstract dynamics models, but uses imagination-based rollouts instead of MCTS.
MuZero's legacy is the demonstration that planning with a learned model can match or exceed planning with a perfect model. This resolved a decades-old debate in reinforcement learning: you do not need to choose between model-based and model-free approaches. You can learn the model and plan with it — as long as you train the model to predict what planning actually needs.
CitationSchrittwieser, Antonoglou, Hubert, Simonyan, Sifre, Schmitt, Guez, Lockhart, Hassabis, Graepel, Lillicrap, Silver. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model. Nature, 2020.
Terms in this paper
- Monte Carlo Tree Searchخوارزمية بحث شجرة مونت كارلو
- Model-Based RLالتعلم بالتعزيز المعتمد على بناء بيئة
- World Modelنموذج العالَم
- Self-Playاللعب الذاتي
- Hidden Stateالحالة المخفية
- Value Functionدالة تقييم العوائد
- Value Networkشبكة القيمة
- Policy Networkشبكة السياسة
- Representationالتمثيل الرقمي
- Planningالتخطيط
- Rolloutالتوليد التجريبي
- Replay Bufferذاكرة التجارب
- Discount Factorمُعامل الخصم
- Action Spaceفضاء الأفعال
- Observationملاحظة