Reinforcement Learning2017advanced12 min read
Mastering the Game of Go Without Human Knowledge
إتقان لعبة غو دون أي معرفة بشرية
Silver, D. · Schrittwieser, J. · Simonyan, K. · Antonoglou, I. · Huang, A. · Guez, A. · Hubert, T. · Baker, L. · Lai, M. · Bolton, A. · Chen, Y. · Lillicrap, T. · Hui, F. · Sifre, L. · van den Driessche, G. · Graepel, T. · Hassabis, D. — Nature
The problem
The original AlphaGo defeated world champions, but it depended on two crutches: supervised learning from millions of human expert games to bootstrap its , and hand-crafted features encoding Go heuristics. This created a ceiling — the system could only be as creative as the human data it learned from — and made the approach impossible to transfer to domains where expert data doesn't exist.
The contribution
AlphaGo Zero starts from random play — — with zero human data. A single dual-headed residual predicts both move probabilities () and win probability (value). uses this network to plan ahead, producing stronger move policies that the network then learns from. This loop — play, search, learn, repeat — is the entire algorithm. After 72 hours on 4 TPUs it defeated the version that beat Lee Sedol, 100–0.
The impact
AlphaGo Zero proved that superhuman performance doesn't require human knowledge — self-play from scratch can surpass centuries of accumulated expertise. This insight led directly to AlphaZero, which mastered chess, shogi, and Go with one algorithm, and inspired a wave of tabula-rasa approaches across science and engineering.
Imagine two students preparing for an exam. Student A memorizes every answer from last year's test papers — fast to start, but limited to patterns humans already discovered.
Student B never sees a single past exam. Instead, they invent practice questions, solve them, grade themselves, and invent harder ones. After enough cycles, Student B not only masters every known technique, but discovers tricks no textbook ever contained.
AlphaGo Zero is Student B. The "exam" is Go — a game with more legal positions than atoms in the universe — and the "practice questions" are games it plays against itself.
The problem: human knowledge as a ceiling
The original AlphaGo was a landmark — the first program to defeat a professional Go player. But its pipeline had three dependencies on human expertise:
- Supervised pre- from 30 million positions in expert games to initialize the policy network.
- Hand-crafted features — liberty counts, ladder patterns, stone age — engineered by Go experts.
- Separate networks — a policy network for move selection and a for position evaluation, trained independently.
These dependencies meant AlphaGo could refine human strategies but never fundamentally transcend them. Worse, the approach was Go-specific: there was no path to transfer it to chess, protein folding, or any domain lacking millions of expert demonstrations.
The core idea: self-play as the only teacher
AlphaGo Zero's premise is radical: throw away all human data. Start from random weights. Let the system play against itself, and use the game outcomes to improve. The entire training loop has just three steps:
- Play — the current neural network plays games against itself, using Monte Carlo Tree Search to select moves.
- Learn — the network is trained to predict the MCTS-improved move probabilities and the eventual game winner.
- Evaluate — the new network plays the current best; if it wins more than 55% of games, it replaces the best and becomes the new self-play generator.
This is a form of generalized policy iteration: MCTS acts as a policy improvement operator (it finds stronger moves than the raw network), and self-play acts as a policy evaluation operator (game outcomes evaluate how good the policy really is). The network distills the search's wisdom, and the improved network makes the next search even better — a virtuous cycle.
Monte Carlo Tree Search: thinking ahead with a neural compass
At each turn, AlphaGo Zero doesn't just use the network's raw move predictions. Instead, it runs 1,600 simulations of MCTS — each simulation a journey from the current board down a . Think of it as mentally rehearsing possible futures before choosing a move.
Each simulation has four stages:
Select — starting from the root (current position), walk down the tree, at each node picking the action that maximizes Q(s,a) + U(s,a). Q is the average value from past simulations; U is an bonus based on the network's prior probability and how rarely this action has been tried. This is the formula — it balances (moves that have worked well) with exploration (moves the network thinks are promising but haven't been tested enough).
Expand — when you reach a leaf node (a position never seen before), expand it: feed the board to the neural network, which returns a policy p (how promising each move is) and a value v (who's likely to win from here).
Backup — carry the value v back up the path, updating Q for every edge you traversed. After many simulations, Q converges to a reliable estimate of each move's worth.
Play — after all 1,600 simulations, choose the move proportional to how often each action was visited at the root. , not raw value, is the selection criterion — a subtle but critical detail that ensures robust move selection.
The PUCT formula: balancing exploration and exploitation
At each node in the search tree, MCTS must decide which branch to explore next. It selects the action that maximizes the sum of two terms: the estimated value Q(s,a) and an exploration bonus U(s,a). The idea is intuitive — follow paths that either have proven high value (exploit) or are promising but under-explored (explore).
The neural network: one brain, two outputs
Unlike the original AlphaGo, which used two separate networks, AlphaGo Zero uses a single dual-headed residual network. The input is a 19×19×17 : the current board plus the last 8 history positions for each player, plus a color indicator.
The network has three parts:
Residual tower — 19 (or 39) residual blocks, each containing two 3×3 convolutional layers with and . The residual connections (skip connections) allow gradients to flow through the deep stack without vanishing — the same trick that made deep image networks possible.
— a 1×1 followed by a outputting a probability over all 19×19 + 1 = 362 possible moves (including pass). This answers: "what should I play next?"
Value head — a 1×1 convolution, a fully connected with 256 units, and a tanh output giving a single scalar in [-1, +1]. This answers: "who is winning from this position?"
Sharing a residual tower between both heads means the network learns features useful for both tasks simultaneously — a position pattern that helps predict good moves also helps predict who will win.
The training objective: learning from your own search
After each self-play game produces a dataset of (state, MCTS policy, game outcome) triples, the network is trained by minimizing a combined :
Read this loss as two teaching signals and a safety net. The value loss says: "get better at predicting who wins." The policy loss says: "learn to mimic the search's move choices — they're smarter than your raw guesses." The says: "don't memorize specific positions — generalize."
As training progresses, the network predicts the search's output more accurately, which makes the search itself more powerful (because it starts from a better network), which produces even better training data — the virtuous cycle again.
Dirichlet noise: ensuring the search doesn't get stuck
A danger of self-play is converging to a narrow strategy — mastering one opening and never discovering that a different opening is even better. AlphaGo Zero guards against this by adding to the root node's prior probabilities:
With and for Go's 19×19 board. This means 25% of the root prior comes from random noise rather than the network — enough to ensure the search occasionally tries surprising moves. The noise only applies at the root (not deeper in the tree), so the search remains focused once a direction is chosen.
Think of it as a creative constraint: the network proposes moves, but a random nudge forces it to occasionally consider moves it would otherwise dismiss. Some of these surprises turn out to be genuinely good — and that's how AlphaGo Zero discovers strategies no human ever played.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
class Node:
"""One position in the MCTS search tree."""
def __init__(self, prior):
self.prior = prior # P(s,a) from the network
self.visit_count = 0 # N(s,a)
self.value_sum = 0.0 # total value accumulated
self.children = {} # action -> Node
def value(self):
if self.visit_count == 0:
return 0
return self.value_sum / self.visit_count # Q(s,a)
def puct_score(parent, child, c_puct=1.5):
"""PUCT: balance exploitation (Q) with exploration (U)."""
exploration = c_puct * child.prior * np.sqrt(parent.visit_count) / (1 + child.visit_count)
return child.value() + exploration
def mcts_simulate(root, network, game, n_simulations=1600):
"""Run n_simulations of select → expand → backup."""
for _ in range(n_simulations):
node = root
path = [node]
state = game.clone()
# SELECT — walk down tree by PUCT
while node.children:
action = max(node.children, key=lambda a: puct_score(node, node.children[a]))
node = node.children[action]
state.apply(action)
path.append(node)
# EXPAND — use network to evaluate leaf
policy, value = network.predict(state) # (p, v) = f_θ(s)
for action, prob in enumerate(policy):
if state.is_legal(action):
node.children[action] = Node(prior=prob)
# BACKUP — propagate value up the path
for ancestor in reversed(path):
ancestor.value_sum += value
ancestor.visit_count += 1
value = -value # flip perspective each level
# Return visit-count distribution as the improved policy
visits = {a: c.visit_count for a, c in root.children.items()}
total = sum(visits.values())
return {a: n / total for a, n in visits.items()}Results: from random to superhuman in 72 hours
AlphaGo Zero's training curve is one of the most dramatic in history. Starting from random play — literally choosing moves uniformly at random — the system:
- After 3 hours: plays basic legal Go, capturing stones and securing territory.
- After 19 hours: rediscovers standard human openings (star-point, 3-4 point).
- After 36 hours: surpasses AlphaGo Lee (the version that defeated Lee Sedol 4-1).
- After 72 hours: defeats AlphaGo Lee 100–0, using only 4 TPUs vs. 48 TPUs.
- After 40 days: achieves an of 5,185, surpassing AlphaGo Master (4,858) which had defeated top professionals 60–0 online.
The final system used only a single machine with 4 TPUs. The original AlphaGo Lee was distributed across many machines with 48 TPUs. Superior performance with 1/12th the compute — that's the power of starting from scratch.
Rediscovering — and transcending — human knowledge
Perhaps the most fascinating aspect of AlphaGo Zero is what it learned, not just how well it played. During training, the system independently rediscovered Go concepts that took humans thousands of years to develop: joseki (corner patterns), life-and-death patterns, influence, thickness, and territory. It went through recognizable phases of Go knowledge: early game focused on local fights, then gradually developing whole-board strategy.
But it didn't stop there. AlphaGo Zero also discovered novel strategies that challenged conventional wisdom — moves that human experts initially dismissed as mistakes but later recognized as genuinely superior. It preferred some openings that humans rarely play and discarded others that humans consider standard.
This is the deepest lesson of the paper: a system with no human bias can discover knowledge that human bias actively prevents us from seeing. Centuries of accumulated Go wisdom, while mostly correct, also contained blind spots — strategies dismissed by tradition that were actually optimal.
Why these design choices matter: ablation study
The paper carefully ablates each design choice to show it matters:
- Dual network vs. separate networks: The single (dual-res) significantly outperforms using separate policy and value networks (sep-res), even though the separate value network predicts expert moves slightly better. Why? Because shared features create a synergy — learning what moves are good also helps learn who's winning.
- Residual vs. plain convolutional: Residual blocks dramatically improve both prediction accuracy and playing strength, especially as depth increases. Without skip connections, deep networks can't effectively learn.
- No human data vs. human data: The tabula rasa version (AlphaGo Zero) ultimately surpasses a version trained with human data (AlphaGo Master). Human data gives a head start but introduces bias that the system must later unlearn.
Why it changed everything
2015
AlphaGo Fan
Defeated European champion Fan Hui. Used supervised learning from human games plus Monte Carlo rollouts. The first program to beat a professional Go player.
2016
AlphaGo Lee
Defeated world champion Lee Sedol 4–1 in a historic match. Still dependent on human expert games for bootstrapping. Used 48 TPUs distributed across many machines.
2017
AlphaGo Zero
Tabula rasa — zero human data. Single dual-headed residual network, no rollouts. Defeated AlphaGo Lee 100–0 using only 4 TPUs on a single machine.
2017
AlphaGo Master
Same architecture as Zero but trained with human data. Defeated top professionals 60–0 online. Eventually surpassed by the tabula rasa Zero version.
2018
AlphaZero
Generalized the approach to chess, shogi, and Go with a single algorithm. Proved that tabula rasa self-play works across domains, not just in Go.
AlphaGo Zero showed that the self-play + MCTS + neural network recipe is domain-general. AlphaZero proved it by mastering three different games with the same code. The principle continues to influence , robotics, and scientific discovery today.
CitationSilver, Schrittwieser, Simonyan, Antonoglou, Huang, Guez, Hubert, Baker, Lai, Bolton, Chen, Lillicrap, Hui, Sifre, van den Driessche, Graepel, Hassabis. Mastering the Game of Go Without Human Knowledge. Nature, 2017.
Terms in this paper
- Self-Playاللعب الذاتي
- Tabula Rasaالصفحة البيضاء
- Monte Carlo Tree Searchخوارزمية بحث شجرة مونت كارلو
- Policy Networkشبكة السياسة
- Value Networkشبكة القيمة
- PUCTحدّ الثقة الأعلى التنبّؤي للأشجار
- Residual Connectionالوصلة التجاوزية
- Reinforcement Learningالتعلم المعزز
- Dual-Headed Networkالشبكة ذات الرأسين
- Dirichlet Noiseتشويش ديريكليه