Reinforcement Learning2018intermediate11 min read
A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play
خوارزمية تعلّم مُعزَّز عامة تتقن الشطرنج والشوغي والغو عبر اللعب الذاتي
Silver, D. · Hubert, T. · Schrittwieser, J. · Antonoglou, I. · Lai, M. · Guez, A. · Lanctot, M. · Sifre, L. · Kumaran, D. · Graepel, T. · Lillicrap, T. · Simonyan, K. · Hassabis, D. — Science
The problem
The strongest game-playing programs — Deep Blue in chess, Elmo in shogi, expert Go engines — relied on decades of hand-engineered evaluation functions, domain-specific search tricks, and massive opening databases crafted by human grandmasters. Each program was a specialist: tuned to one game, brittle outside it, and incapable of transferring its intelligence to another domain. Building a new champion for a new game meant starting from scratch.
The contribution
AlphaZero: a single, general-purpose algorithm that masters chess, shogi, and Go from scratch. It combines a deep residual with . The network has two heads — one outputs a (move probabilities) and the other a value (win probability). The only input is the game rules. Through millions of games and , AlphaZero defeated the world-champion programs in all three games: Stockfish in chess, Elmo in shogi, and AlphaGo Zero in Go — all within 24 hours of .
The impact
AlphaZero proved that a single learning algorithm, given only the rules, can surpass decades of human-engineered expertise across multiple domains. It catalyzed a family of successors: MuZero (which learns even the rules), AlphaStar (StarCraft II), and AlphaTensor (discovering faster matrix multiplication algorithms). The idea that self-play plus search plus a neural network is a general recipe for mastery reshaped how the field thinks about artificial general intelligence.
Imagine three different board games — chess, shogi, and Go — each with its own guild of craftsmen who have spent decades carving specialized tools: custom-made chisels, jigs, and measuring sticks, each useless outside its own workshop.
AlphaZero walks in with one universal toolkit — a notebook, a pencil, and a rule that says "play yourself, write down what works." Within a single day, it builds furniture that outclasses every guild's best work. Same toolkit, three workshops, zero prior carpentry knowledge.
The notebook is the neural network. The pencil is self-play. And the rule for deciding which shelf to inspect next is Monte Carlo tree search.
The problem: domain-specific engineering does not generalize
Before AlphaZero, the world's strongest chess program — Stockfish — was the product of decades of collective human effort: hand-tuned evaluation functions that assigned precise numerical scores to piece positions, pawn structures, king safety, and thousands of other chess-specific patterns. Sophisticated alpha-beta search with domain-specific pruning and extension heuristics explored millions of positions per second.
For shogi (Japanese chess), a completely separate ecosystem of engine developers built Elmo, using similar but distinct domain-specific tricks — different piece values, drop rules, promotion patterns. For Go, the story was the same: expert systems with hand-crafted patterns.
The core limitation was clear: nothing transferred. A chess engine knew nothing about shogi. A Go engine knew nothing about chess. Intelligence was locked inside domain-specific engineering, and each new game demanded a new army of experts starting from scratch.
The idea: one algorithm, three games, zero human knowledge
AlphaZero's recipe has three ingredients that together form a self-improving loop:
1. A neural network with two heads. Given a board position, the network outputs two things simultaneously: a policy — a probability distribution over all legal moves suggesting which ones look promising — and a value — a single number estimating the probability of winning from this position. The network is built from residual blocks with — the same building blocks proven in ResNet.
2. Monte Carlo tree search (MCTS) guided by the network. Instead of searching exhaustively (like alpha-beta), MCTS uses the neural network as a guide. The policy head says "look here first" and the value head says "this position is worth exploring." MCTS runs hundreds of simulated games forward from the current position, building a search tree focused on the most promising variations.
3. Self-play reinforcement learning. The system plays millions of games against itself. After each game, the actual result (win/draw/loss) becomes the training signal. The network learns to predict which moves MCTS would recommend (policy target) and who actually won (value target). As the network improves, MCTS makes better decisions, which generates better training data, which improves the network further — a virtuous cycle.
The neural network: residual tower with dual heads
The network architecture is elegantly simple. The board state is encoded as a stack of planes — for chess, this includes piece positions for the last 8 moves, castling rights, move count, and side to move (119 input planes per 8×8 board). This stack passes through:
- A with 256 filters and batch normalization.
- 19 residual blocks, each containing two convolutional layers with batch normalization and activations, plus a (identical to ResNet blocks). These blocks extract hierarchical features from the position — patterns within patterns.
- A policy head that outputs a probability distribution over all legal moves via a convolution followed by a fully-connected layer and .
- A value head that outputs a single scalar in [-1, +1] representing the expected outcome (win/loss/draw) via a convolution, fully-connected layers, and activation.
The same architecture is used for all three games — only the input encoding and output size change to reflect different board sizes and move spaces. The network structure itself is completely game-agnostic.
Monte Carlo tree search: thinking ahead with a learned guide
Traditional chess engines use alpha-beta search: they consider every move, every reply, every counter-reply, pruning only branches that are provably worse. This is exhaustive but wasteful — most positions don't deserve deep analysis.
AlphaZero uses MCTS with a (Predictor + Upper Confidence bound for Trees) selection rule. At each node in the search tree, the algorithm picks the action that maximizes a balance of (choosing moves that have worked well so far) and (trying moves the network thinks are promising but haven't been tested much).
The PUCT formula selects the action that maximizes a sum of two terms. The first term captures exploitation — how well this action has performed in past simulations. The second term captures exploration — it is higher for actions the neural network's policy head rates highly but that have been visited relatively few times. This ensures the search both trusts the network's intuition and verifies it through actual simulation.
To ensure the search doesn't become too narrow, is added to the neural network's prior probabilities at the root node. This is like occasionally throwing a dart at the board randomly — it forces the system to explore moves it might otherwise overlook. The noise parameter is scaled inversely to the number of legal moves: for chess (~30 legal moves), for shogi (~80 legal moves), and for Go (~250 legal moves).
The training loop: playing yourself into mastery
The training process is a continuous cycle. Start with a randomly initialized neural network — it plays moves essentially at random. Use this network to guide MCTS as it plays games against itself. After each game, store every position along with two labels:
- The MCTS policy — the distribution of visit counts at the root node, which represents the search's recommendation after careful deliberation.
- The game outcome — who actually won (loss, draw, or win from the current player's perspective).
Then train the neural network to predict both of these: make its policy head match and its value head match . The trained network becomes a better MCTS guide, the better MCTS generates higher-quality games, and the cycle repeats. AlphaZero uses 800 MCTS simulations per move during training and generates millions of games per training run.
Results: superhuman in 24 hours
AlphaZero's climbed rapidly during training. In chess, it surpassed Stockfish after just 4 hours (300,000 training steps). In shogi, it surpassed Elmo after 2 hours (110,000 steps). In Go, it surpassed the version of AlphaGo that defeated Lee Sedol after 30 hours.
In head-to-head matches under tournament conditions (3 hours per game, 15 seconds per move increment):
- Chess vs Stockfish: AlphaZero won 155 games, drew 839, lost only 6 out of 1,000 games.
- Shogi vs Elmo: AlphaZero won 91.2% of games.
- Go vs AlphaGo Zero: AlphaZero won 61% of games.
Perhaps most impressively, AlphaZero's MCTS examined only about 80,000 positions per second in chess — compared to Stockfish's 70 million. It searched 1,000× fewer positions but searched far more intelligently, focusing computation on the positions that mattered most.
Why it works: the virtuous cycle of search and learning
The power of AlphaZero comes from the synergy between its components. The neural network alone cannot play perfectly — it makes approximate evaluations. MCTS alone would be impractically slow without a guide. Together, they amplify each other:
The network provides MCTS with a strong prior over which moves to explore and an informed evaluation of leaf positions. MCTS, in turn, produces better move recommendations than the network alone — it looks multiple steps ahead and corrects the network's errors through actual simulation. These corrected recommendations become the training signal for the next version of the network.
This creates a self-reinforcing loop: better network → better search → better training data → even better network. The network is the fast, intuitive "gut feeling" (System 1), and MCTS is the slow, deliberate "analytical reasoning" (System 2). AlphaZero's genius is using each to improve the other.
Impact: from board games to scientific discovery
AlphaZero's significance extends far beyond games. It demonstrated that a single general-purpose algorithm could match or surpass highly-optimized, domain-specific systems that took decades to build. This shifted the AI community's perspective: perhaps generality is not the enemy of performance, but a path to it.
The same core ideas — self-play, neural network evaluation, and tree search — were extended into a family of algorithms that tackled progressively harder problems beyond perfect-information board games.
2016
AlphaGo defeats Lee Sedol
DeepMind's Go program beat a world champion 4-1, using deep learning + MCTS + human game data. A milestone, but still relied on human expertise for initial training.
2017
AlphaGo Zero — no human data at all
Learned Go from scratch via pure self-play, surpassing all previous versions. Proved human knowledge was not just unnecessary but a bottleneck.
2018
AlphaZero — generality across three games
Same algorithm masters chess, shogi, and Go. Defeated Stockfish, Elmo, and AlphaGo Zero. Demonstrated that the self-play + MCTS + neural network recipe generalizes.
2019
MuZero — learning even the rules
Extended AlphaZero to learn the environment dynamics (the rules) from experience. Mastered Go, chess, shogi, and Atari games without being told the rules.
2019
AlphaStar — real-time strategy
Applied self-play to StarCraft II — a game with imperfect information, continuous actions, and enormous complexity. Reached Grandmaster level.
2022
AlphaTensor — discovering algorithms
Used AlphaZero's framework to discover faster matrix multiplication algorithms — the first time an AI system found novel algorithms for a fundamental math problem.
The legacy of AlphaZero is not just that it won games — it is that it won them by learning, not by being told. The same three ingredients — a neural network that evaluates and suggests, a search procedure that deliberates, and a self-play loop that generates its own curriculum — have become the foundation for AI systems tackling problems far beyond games.
CitationSilver, Hubert, Schrittwieser, Antonoglou, Lai, Guez, Lanctot, Sifre, Kumaran, Graepel, Lillicrap, Simonyan, Hassabis. A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go Through Self-Play. Science, 2018.
Terms in this paper
- Self-Playاللعب الذاتي
- Monte Carlo Tree Searchخوارزمية بحث شجرة مونت كارلو
- PUCTحدّ الثقة الأعلى التنبّؤي للأشجار
- Tabula Rasaالصفحة البيضاء
- Reinforcement Learningالتعلم المعزز
- Policyالسياسة
- Value Functionدالة تقييم العوائد
- Explorationالاستكشاف (تجربة أفعال جديدة)
- Exploitationالاستغلال (اعتماد الأفعال الناجحة)
- Dirichlet Noiseتشويش ديريكليه
- Residual Blockالكتلة المتبقّية
- Elo Rating (for model comparison)تصنيف إيلو القياسي للنماذج