Reinforcement Learning2018intermediate11 min read

A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play

خوارزمية تعلّم مُعزَّز عامة تتقن الشطرنج والشوغي والغو عبر اللعب الذاتي

Silver, D. · Hubert, T. · Schrittwieser, J. · Antonoglou, I. · Lai, M. · Guez, A. · Lanctot, M. · Sifre, L. · Kumaran, D. · Graepel, T. · Lillicrap, T. · Simonyan, K. · Hassabis, D. — Science

The problem

The strongest game-playing programs — Deep Blue in chess, Elmo in shogi, expert Go engines — relied on decades of hand-engineered evaluation functions, domain-specific search tricks, and massive opening databases crafted by human grandmasters. Each program was a specialist: tuned to one game, brittle outside it, and incapable of transferring its intelligence to another domain. Building a new champion for a new game meant starting from scratch.

The contribution

AlphaZero: a single, general-purpose algorithm that masters chess, shogi, and Go from scratch. It combines a deep residual with . The network has two heads — one outputs a (move probabilities) and the other a value (win probability). The only input is the game rules. Through millions of games and , AlphaZero defeated the world-champion programs in all three games: Stockfish in chess, Elmo in shogi, and AlphaGo Zero in Go — all within 24 hours of .

The impact

AlphaZero proved that a single learning algorithm, given only the rules, can surpass decades of human-engineered expertise across multiple domains. It catalyzed a family of successors: MuZero (which learns even the rules), AlphaStar (StarCraft II), and AlphaTensor (discovering faster matrix multiplication algorithms). The idea that self-play plus search plus a neural network is a general recipe for mastery reshaped how the field thinks about artificial general intelligence.

Imagine three different board games — chess, shogi, and Go — each with its own guild of craftsmen who have spent decades carving specialized tools: custom-made chisels, jigs, and measuring sticks, each useless outside its own workshop.

AlphaZero walks in with one universal toolkit — a notebook, a pencil, and a rule that says "play yourself, write down what works." Within a single day, it builds furniture that outclasses every guild's best work. Same toolkit, three workshops, zero prior carpentry knowledge.

The notebook is the neural network. The pencil is self-play. And the rule for deciding which shelf to inspect next is Monte Carlo tree search.

The problem: domain-specific engineering does not generalize

Before AlphaZero, the world's strongest chess program — Stockfish — was the product of decades of collective human effort: hand-tuned evaluation functions that assigned precise numerical scores to piece positions, pawn structures, king safety, and thousands of other chess-specific patterns. Sophisticated alpha-beta search with domain-specific pruning and extension heuristics explored millions of positions per second.

For shogi (Japanese chess), a completely separate ecosystem of engine developers built Elmo, using similar but distinct domain-specific tricks — different piece values, drop rules, promotion patterns. For Go, the story was the same: expert systems with hand-crafted patterns.

The core limitation was clear: nothing transferred. A chess engine knew nothing about shogi. A Go engine knew nothing about chess. Intelligence was locked inside domain-specific engineering, and each new game demanded a new army of experts starting from scratch.

Open in Lab
Compare the traditional approach (years of hand-crafted features) vs. AlphaZero's approach (a single algorithm, no domain knowledge).
The demo wakes as you arrive…

The idea: one algorithm, three games, zero human knowledge

AlphaZero's recipe has three ingredients that together form a self-improving loop:

1. A neural network with two heads. Given a board position, the network outputs two things simultaneously: a policy p\mathbf{p} — a probability distribution over all legal moves suggesting which ones look promising — and a value vv — a single number estimating the probability of winning from this position. The network is built from residual blocks with — the same building blocks proven in ResNet.

2. Monte Carlo tree search (MCTS) guided by the network. Instead of searching exhaustively (like alpha-beta), MCTS uses the neural network as a guide. The policy head says "look here first" and the value head says "this position is worth exploring." MCTS runs hundreds of simulated games forward from the current position, building a search tree focused on the most promising variations.

3. Self-play reinforcement learning. The system plays millions of games against itself. After each game, the actual result (win/draw/loss) becomes the training signal. The network learns to predict which moves MCTS would recommend (policy target) and who actually won (value target). As the network improves, MCTS makes better decisions, which generates better training data, which improves the network further — a virtuous cycle.

The neural network: residual tower with dual heads

The network architecture is elegantly simple. The board state is encoded as a stack of planes — for chess, this includes piece positions for the last 8 moves, castling rights, move count, and side to move (119 input planes per 8×8 board). This stack passes through:

  • A with 256 filters and batch normalization.
  • 19 residual blocks, each containing two convolutional layers with batch normalization and activations, plus a (identical to ResNet blocks). These blocks extract hierarchical features from the position — patterns within patterns.
  • A policy head that outputs a probability distribution over all legal moves via a convolution followed by a fully-connected layer and .
  • A value head that outputs a single scalar in [-1, +1] representing the expected outcome (win/loss/draw) via a convolution, fully-connected layers, and activation.

The same architecture is used for all three games — only the input encoding and output size change to reflect different board sizes and move spaces. The network structure itself is completely game-agnostic.

Open in Lab
Click each layer to see its role. The same architecture serves chess, shogi, and Go.
The demo wakes as you arrive…

Monte Carlo tree search: thinking ahead with a learned guide

Traditional chess engines use alpha-beta search: they consider every move, every reply, every counter-reply, pruning only branches that are provably worse. This is exhaustive but wasteful — most positions don't deserve deep analysis.

AlphaZero uses MCTS with a (Predictor + Upper Confidence bound for Trees) selection rule. At each node in the search tree, the algorithm picks the action that maximizes a balance of (choosing moves that have worked well so far) and (trying moves the network thinks are promising but haven't been tested much).

The PUCT formula selects the action that maximizes a sum of two terms. The first term captures exploitation — how well this action has performed in past simulations. The second term captures exploration — it is higher for actions the neural network's policy head rates highly but that have been visited relatively few times. This ensures the search both trusts the network's intuition and verifies it through actual simulation.

a∗=arg⁡max⁡a[Q(s,a)+cpuct⋅P(s,a)⋅∑bN(s,b)1+N(s,a)]a^* = \arg\max_a \left[ Q(s,a) + c_{\text{puct}} \cdot P(s,a) \cdot \frac{\sqrt{\sum_b N(s,b)}}{1 + N(s,a)} \right]
PUCT Selection Rule — Q(s,a) is the mean value of action a from state s (exploitation). P(s,a) is the neural network's prior probability for that move. N(s,a) counts how many times this action has been explored. The second term shrinks as an action is visited more, naturally shifting from exploration to exploitation.

To ensure the search doesn't become too narrow, is added to the neural network's prior probabilities at the root node. This is like occasionally throwing a dart at the board randomly — it forces the system to explore moves it might otherwise overlook. The noise parameter α\alpha is scaled inversely to the number of legal moves: α=0.3\alpha = 0.3 for chess (~30 legal moves), α=0.15\alpha = 0.15 for shogi (~80 legal moves), and α=0.03\alpha = 0.03 for Go (~250 legal moves).

Open in Lab
Step through MCTS — watch how the search tree grows, guided by the neural network.
The demo wakes as you arrive…

The training loop: playing yourself into mastery

The training process is a continuous cycle. Start with a randomly initialized neural network — it plays moves essentially at random. Use this network to guide MCTS as it plays games against itself. After each game, store every position along with two labels:

  • The MCTS policy π\boldsymbol{\pi} — the distribution of visit counts at the root node, which represents the search's recommendation after careful deliberation.
  • The game outcome z∈{−1,0,+1}z \in \{-1, 0, +1\} — who actually won (loss, draw, or win from the current player's perspective).

Then train the neural network to predict both of these: make its policy head match π\boldsymbol{\pi} and its value head match zz. The trained network becomes a better MCTS guide, the better MCTS generates higher-quality games, and the cycle repeats. AlphaZero uses 800 MCTS simulations per move during training and generates millions of games per training run.

ℓ=(z−v)2⏟value loss−π⊤log⁡p⏟policy loss+c∥θ∥2⏟regularization\ell = \underbrace{(z - v)^2}_{\text{value loss}} - \underbrace{\boldsymbol{\pi}^\top \log \mathbf{p}}_{\text{policy loss}} + \underbrace{c \|\theta\|^2}_{\text{regularization}}
AlphaZero Loss Function — The loss combines mean squared error between the predicted value v and actual outcome z, cross-entropy between the predicted policy p and the MCTS policy π, and L2 regularization. The value loss teaches the network to evaluate positions; the policy loss teaches it which moves to recommend; and regularization prevents overfitting.
Open in Lab
Watch the self-play training loop — the same cycle that teaches all three games.
The demo wakes as you arrive…

Results: superhuman in 24 hours

AlphaZero's climbed rapidly during training. In chess, it surpassed Stockfish after just 4 hours (300,000 training steps). In shogi, it surpassed Elmo after 2 hours (110,000 steps). In Go, it surpassed the version of AlphaGo that defeated Lee Sedol after 30 hours.

In head-to-head matches under tournament conditions (3 hours per game, 15 seconds per move increment):

  • Chess vs Stockfish: AlphaZero won 155 games, drew 839, lost only 6 out of 1,000 games.
  • Shogi vs Elmo: AlphaZero won 91.2% of games.
  • Go vs AlphaGo Zero: AlphaZero won 61% of games.

Perhaps most impressively, AlphaZero's MCTS examined only about 80,000 positions per second in chess — compared to Stockfish's 70 million. It searched 1,000× fewer positions but searched far more intelligently, focusing computation on the positions that mattered most.

Open in Lab
Training progress of AlphaZero across all three games. Notice how it surpasses the world champion programs within hours.
The demo wakes as you arrive…
Open in Lab
Compare board size, branching factor, and game length across all three domains.
The demo wakes as you arrive…

Why it works: the virtuous cycle of search and learning

The power of AlphaZero comes from the synergy between its components. The neural network alone cannot play perfectly — it makes approximate evaluations. MCTS alone would be impractically slow without a guide. Together, they amplify each other:

The network provides MCTS with a strong prior over which moves to explore and an informed evaluation of leaf positions. MCTS, in turn, produces better move recommendations than the network alone — it looks multiple steps ahead and corrects the network's errors through actual simulation. These corrected recommendations become the training signal for the next version of the network.

This creates a self-reinforcing loop: better network → better search → better training data → even better network. The network is the fast, intuitive "gut feeling" (System 1), and MCTS is the slow, deliberate "analytical reasoning" (System 2). AlphaZero's genius is using each to improve the other.

Impact: from board games to scientific discovery

AlphaZero's significance extends far beyond games. It demonstrated that a single general-purpose algorithm could match or surpass highly-optimized, domain-specific systems that took decades to build. This shifted the AI community's perspective: perhaps generality is not the enemy of performance, but a path to it.

The same core ideas — self-play, neural network evaluation, and tree search — were extended into a family of algorithms that tackled progressively harder problems beyond perfect-information board games.

  1. 2016

    AlphaGo defeats Lee Sedol

    DeepMind's Go program beat a world champion 4-1, using deep learning + MCTS + human game data. A milestone, but still relied on human expertise for initial training.

  2. 2017

    AlphaGo Zero — no human data at all

    Learned Go from scratch via pure self-play, surpassing all previous versions. Proved human knowledge was not just unnecessary but a bottleneck.

  3. 2018

    AlphaZero — generality across three games

    Same algorithm masters chess, shogi, and Go. Defeated Stockfish, Elmo, and AlphaGo Zero. Demonstrated that the self-play + MCTS + neural network recipe generalizes.

  4. 2019

    MuZero — learning even the rules

    Extended AlphaZero to learn the environment dynamics (the rules) from experience. Mastered Go, chess, shogi, and Atari games without being told the rules.

  5. 2019

    AlphaStar — real-time strategy

    Applied self-play to StarCraft II — a game with imperfect information, continuous actions, and enormous complexity. Reached Grandmaster level.

  6. 2022

    AlphaTensor — discovering algorithms

    Used AlphaZero's framework to discover faster matrix multiplication algorithms — the first time an AI system found novel algorithms for a fundamental math problem.

The legacy of AlphaZero is not just that it won games — it is that it won them by learning, not by being told. The same three ingredients — a neural network that evaluates and suggests, a search procedure that deliberates, and a self-play loop that generates its own curriculum — have become the foundation for AI systems tackling problems far beyond games.

CitationSilver, Hubert, Schrittwieser, Antonoglou, Lai, Guez, Lanctot, Sifre, Kumaran, Graepel, Lillicrap, Simonyan, Hassabis. A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play. Science, 2018.

Terms in this paper