Reinforcement Learning2016advanced11 min read

Mastering the Game of Go with Deep Neural Networks and Tree Search

إتقان لعبة غو باستخدام الشبكات العصبية العميقة والبحث الشجري

Silver, D. · Huang, A. · Maddison, C. J. · Guez, A. · Sifre, L. · van den Driessche, G. · Schrittwieser, J. · Antonoglou, I. · Panneershelvam, V. · Lanctot, M. · Dieleman, S. · Grewe, D. · Nham, J. · Kalchbrenner, N. · Sutskever, I. · Lillicrap, T. · Leach, M. · Kavukcuoglu, K. · Graepel, T. · Hassabis, D. — Nature

The problem

The game of Go has ~10¹⁷⁰ possible board positions — more than atoms in the observable universe — making brute-force search impossible. Traditional game-playing AI (like Deep Blue for chess) relied on handcrafted evaluation functions and exhaustive search, but Go's subtle positional nuances defied such simplification. Prior programs reached only amateur levels. Go was considered AI's "grand challenge" and experts predicted it would take another decade to solve.

The contribution

AlphaGo combines deep convolutional neural networks with Tree Search through a four-stage pipeline: (1) a trained on 30 million expert moves to predict human plays, (2) a fast policy for rapid game simulations, (3) a policy network refined by , and (4) a that evaluates board positions. During play, MCTS uses the policy network to guide which moves to explore and the value network (mixed with rollouts) to evaluate positions — replacing handcrafted evaluation with learned intuition.

The impact

AlphaGo defeated the European Go champion Fan Hui 5–0 in October 2015, and world champion Lee Sedol 4–1 in March 2016 — a feat experts had predicted was at least a decade away. It proved that deep reinforcement learning can master domains once thought to require human intuition, launching a lineage of systems (AlphaGo Zero, AlphaZero, MuZero) that removed human data entirely, and inspiring applications from protein folding to chip design to mathematical theorem discovery.

Imagine a chess grandmaster who can see ten moves ahead. Now imagine a Go board where the number of possible games exceeds the number of atoms in the universe. No grandmaster — and no computer — can see even a fraction of that tree.

AlphaGo's trick: train a scout (the policy network) that has watched millions of expert games and knows which moves look promising, and a judge (the value network) that can glance at a board and say "this side is winning." Then send the scout and judge into a controlled expedition (MCTS) where they explore only the most promising paths and evaluate positions along the way — trading exhaustive search for intelligent search.

The wall: Go's search space defies brute force

Chess has roughly 10⁴⁷ legal positions; Go has approximately 10¹⁷⁰. In chess, the average is ~35 moves per turn; in Go it is ~250. Deep Blue conquered chess in 1997 by evaluating 200 million positions per second with handcrafted evaluation functions. That strategy is mathematically impossible for Go — even if you evaluated a trillion positions per second for the age of the universe, you would not scratch the surface.

Before AlphaGo, the best computer Go programs used Monte Carlo Tree Search with shallow pattern features and reached only amateur-level play (~2 dan amateur). The gap between amateur and professional was enormous — roughly equivalent to the gap between a club chess player and a world champion. Bridging that gap required replacing handcrafted features with learned representations that capture the deep positional intuition human experts develop over decades.

Open in Lab
Compare the search space of Chess vs. Go — notice the exponential gap.
The demo wakes as you arrive…

The training pipeline: four stages from imitation to mastery

AlphaGo's training pipeline is a four-stage curriculum — think of it as an apprenticeship: first watch the master, then practice on your own, then develop judgment about positions.

Stage 1 — Supervised Learning (SL) Policy Network: A 13- deep is trained on 30 million board positions from expert human games (KGS Go Server). The network takes the 19×19 board as input (with 48 planes encoding stone positions, liberties, capture history, etc.) and outputs a probability over all 361 intersections — predicting the move a human expert would play. This network achieved 57% accuracy in predicting expert moves, a significant leap over the prior state of the art of 44%.

Stage 2 — Fast Rollout Policy: A simpler, lightweight policy trained on the same data but using only local pattern features. It sacrifices accuracy (24.2%) for speed — selecting a move in 2 microseconds instead of 3 milliseconds. This is used later inside MCTS to simulate thousands of games rapidly.

Stage 3 — Reinforcement Learning (RL) Policy Network: The SL policy is cloned and improved through self-play. The network plays games against randomly selected earlier versions of itself, and updates its weights using methods () to maximize the probability of winning. After training, the RL policy wins 80% of games against the SL policy — it has learned moves that win games, not just moves that look human.

Stage 4 — Value Network: A network with the same architecture as the policy network but outputting a single scalar — the probability of winning from a given position. Trained on 30 million positions generated by RL self-play (using different games for each position to avoid ), it learns to evaluate board positions without playing to the end.

Open in Lab
Click through the four training stages — watch AlphaGo evolve from imitating experts to surpassing them.
The demo wakes as you arrive…

Two neural networks, two questions

The genius of AlphaGo lies in factoring the game-playing problem into two complementary questions answered by two separate networks:

The policy network answers: "What move should I consider next?" Given a board position, it outputs a probability distribution over all legal moves. High-probability moves are explored first in the , dramatically reducing the effective branching factor from ~250 to a handful of promising candidates. This is like a chess player's intuition — an instant sense of which moves "feel right" without calculating deeply.

The value network answers: "Who is winning in this position?" Given a board position, it outputs a single number between 0 and 1 — the estimated probability of winning. This replaces the need to play a game all the way to the end to know the outcome. A single through the value network is equivalent to thousands of random simulations, compressing judgment into learned intuition.

Both networks share the same deep convolutional architecture (13 layers), but differ in their final output: a distribution (policy) vs. a scalar (value).

Open in Lab
Compare the two networks side by side — the policy network narrows the search, the value network evaluates depth.
The demo wakes as you arrive…

The math: policy gradient and value regression

The SL policy network learns by maximizing the log-likelihood of expert moves — the same cross-entropy objective used in . Given a board state ss and the expert move aa, the update pushes the network to assign higher probability to the correct move:

Δσ∝∂log⁡pσ(a∣s)∂σ\Delta\sigma \propto \frac{\partial \log p_\sigma(a \mid s)}{\partial \sigma}
SL Policy Gradient — Update the weights σ in the direction that increases the probability of the expert action a in state s. This is standard supervised classification applied to move prediction.

The RL policy network is then trained with REINFORCE: play a complete game, observe the outcome zz (+1 for win, −1 for loss), and update all moves in the game to make winning moves more likely:

Δρ∝∂log⁡pρ(at∣st)∂ρ⋅zt\Delta\rho \propto \frac{\partial \log p_\rho(a_t \mid s_t)}{\partial \rho} \cdot z_t
REINFORCE Policy Gradient — This learning rule adjusts the policy based on the outcome of an episode. Actions that contribute to successful outcomes become more likely to be selected in the future, while actions associated with poor outcomes become less likely. Through repeated self-play and feedback from results, the policy gradually learns strategies that maximize long-term success rather than simply imitating existing human behavior.

Finally, the value network is trained by regression — minimizing the between its prediction and the actual game outcome:

L(θ)=1N∑i=1N(zi−vθ(si))2L(\theta) = \frac{1}{N}\sum_{i=1}^{N}(z_i - v_\theta(s_i))^2
Value Network Loss — The value network is trained to predict the eventual outcome of a game from a given position. During training, its predictions are compared with the actual results observed after the game ends, and the network is updated to reduce the difference. Once trained, the value network can estimate the quality of a position with a single forward pass, providing a fast alternative to running large numbers of simulation-based evaluations.

Monte Carlo Tree Search: the decision engine

During actual gameplay, AlphaGo uses MCTS — a four-step cycle repeated thousands of times per move:

Selection: Starting from the root (current board position), traverse the tree by picking at each node the action that maximizes Q(s,a)+u(s,a)Q(s,a) + u(s,a), where QQ is the current estimated value and uu is a bonus that encourages exploring moves the policy network considers promising but that haven't been tried much yet. This balances the classic tension between (trying new paths) and (following paths known to be good).

Expansion: When the traversal reaches a leaf node (a position not yet in the tree), expand it by adding the new position and using the SL policy network to initialize the prior probabilities of its children.

Evaluation: Evaluate the leaf position using a weighted combination of: (a) the value network's prediction, and (b) the outcome of a fast rollout simulation using the lightweight rollout policy. The mixing parameter λ balances the neural network's deep understanding against the rollout's concrete game completion.

Backup: Propagate the evaluation value back up the tree, updating the QQ values of every node along the traversal path. After many iterations, the root's children have reliable value estimates, and AlphaGo selects the most-visited child as its move.

V(sL)=(1−λ) vθ(sL)+λ zLV(s_L) = (1 - \lambda)\, v_\theta(s_L) + \lambda\, z_L
Leaf Evaluation (mixing formula) — Rather than relying on a single source of information, AlphaGo combines two complementary estimates when evaluating a search leaf. One estimate comes from the value network, which provides a learned assessment of the position, while the other comes from simulated game continuations that play the game out to completion. A mixing factor controls how much each source contributes to the final evaluation, balancing strategic judgment with direct empirical evidence.
Open in Lab
Step through the four MCTS phases — see how the policy network guides exploration and the value network evaluates positions.
The demo wakes as you arrive…

Self-play: learning beyond human knowledge

The critical insight that elevates AlphaGo beyond imitation is self-play. Once the SL policy network learned to mimic expert moves, it was cloned and set to play against itself millions of times. Through policy gradient reinforcement learning, it discovered moves that no human expert had played — moves that don't look "natural" but lead to victory.

Self-play solves two key problems simultaneously. First, it generates an unlimited supply of training data — no longer dependent on the finite corpus of human games. Second, it creates a curriculum of increasing difficulty: as the network improves, its opponent (an earlier version of itself) also improves, creating a natural arms race that pushes performance beyond what human data alone could teach.

The RL policy network after self-play training won 80% of games against the SL policy network. More importantly, it discovered novel strategies: moves that professional commentators initially called "mistakes" but later recognized as deeply creative and correct. This foreshadowed the even more dramatic discoveries in AlphaGo Zero, which learned entirely from self-play without any human data at all.

Open in Lab
Watch the RL policy improve through generations of self-play — each generation plays its predecessor and learns from the outcome.
The demo wakes as you arrive…

Results: from amateur to superhuman

AlphaGo achieved a 99.8% win rate against all other Go programs and became the first program to defeat a human professional player on the full 19×19 board. In October 2015, it beat Fan Hui (European champion, 2 dan professional) 5–0 in a formal match. In March 2016, it defeated Lee Sedol (9 dan professional, widely considered one of the greatest players in the game's history) 4–1 in a globally televised match watched by over 200 million people.

The distributed version of AlphaGo used 1920 CPUs and 280 GPUs. However, even the single-machine version (48 CPUs + 8 GPUs) was strong enough to defeat Fan Hui — demonstrating that the algorithmic innovations, not just raw compute, were responsible for the breakthrough.

An revealed the relative contribution of each component: the value network alone was stronger than rollouts alone; combining both was stronger than either; and using the policy network to guide MCTS was essential — without it, the search explored too many irrelevant moves.

Open in Lab
Compare the Elo ratings of AlphaGo versions against top Go programs and human champions.
The demo wakes as you arrive…

What AlphaGo unlocked

  1. 2016

    AlphaGo (this paper)

    Combined deep CNNs with MCTS and self-play RL to defeat a professional Go player for the first time. Used 30 million human games as a starting point.

  2. 2017

    AlphaGo Zero — no human data at all

    Learned Go entirely from self-play starting with random weights. Surpassed the original AlphaGo in 40 hours. Proved that human knowledge was a helpful starting point but not a ceiling.

  3. 2018

    AlphaZero — one algorithm, three games

    Generalized AlphaGo Zero to chess and shogi with no game-specific tuning. Defeated Stockfish (chess) and Elmo (shogi) after hours of self-play training.

  4. 2019

    MuZero — learning without knowing the rules

    Extended the approach to environments where the rules are unknown. Learned a model of the environment alongside the policy and value networks. Mastered Go, chess, shogi, and Atari games.

  5. 2023

    Tree of Thoughts — MCTS meets language models

    Applied tree search ideas to LLM reasoning — branching and evaluating chains of thought like moves in a game. AlphaGo's core insight (search + learned evaluation) scaled beyond games.

AlphaGo's most lasting contribution is not defeating a Go champion — it is the proof that combining neural network intuition with tree search planning creates systems far stronger than either approach alone. This principle — learned evaluation + intelligent search — has become a blueprint for AI systems, from AlphaGo Zero's pure self-play to modern chain-of-thought and tree-of-thought reasoning in large language models.

CitationSilver, Huang, Maddison, Guez, Sifre, van den Driessche, Schrittwieser, Antonoglou, Panneershelvam, Lanctot, Dieleman, Grewe, Nham, Kalchbrenner, Sutskever, Lillicrap, Leach, Kavukcuoglu, Graepel, Hassabis. Mastering the Game of Go with Deep Neural Networks and Tree Search. Nature, 2016.

Terms in this paper