Reinforcement Learning2017advanced12 min read

Mastering the Game of Go Without Human Knowledge

إتقان لعبة غو دون أي معرفة بشرية

Silver, D. · Schrittwieser, J. · Simonyan, K. · Antonoglou, I. · Huang, A. · Guez, A. · Hubert, T. · Baker, L. · Lai, M. · Bolton, A. · Chen, Y. · Lillicrap, T. · Hui, F. · Sifre, L. · van den Driessche, G. · Graepel, T. · Hassabis, D. — Nature

The problem

The original AlphaGo defeated world champions, but it depended on two crutches: supervised learning from millions of human expert games to bootstrap its , and hand-crafted features encoding Go heuristics. This created a ceiling — the system could only be as creative as the human data it learned from — and made the approach impossible to transfer to domains where expert data doesn't exist.

The contribution

AlphaGo Zero starts from random play — — with zero human data. A single dual-headed residual predicts both move probabilities () and win probability (value). uses this network to plan ahead, producing stronger move policies that the network then learns from. This loop — play, search, learn, repeat — is the entire algorithm. After 72 hours on 4 TPUs it defeated the version that beat Lee Sedol, 100–0.

The impact

AlphaGo Zero proved that superhuman performance doesn't require human knowledge — self-play from scratch can surpass centuries of accumulated expertise. This insight led directly to AlphaZero, which mastered chess, shogi, and Go with one algorithm, and inspired a wave of tabula-rasa approaches across science and engineering.

Imagine two students preparing for an exam. Student A memorizes every answer from last year's test papers — fast to start, but limited to patterns humans already discovered.

Student B never sees a single past exam. Instead, they invent practice questions, solve them, grade themselves, and invent harder ones. After enough cycles, Student B not only masters every known technique, but discovers tricks no textbook ever contained.

AlphaGo Zero is Student B. The "exam" is Go — a game with more legal positions than atoms in the universe — and the "practice questions" are games it plays against itself.

The problem: human knowledge as a ceiling

The original AlphaGo was a landmark — the first program to defeat a professional Go player. But its pipeline had three dependencies on human expertise:

  • Supervised pre- from 30 million positions in expert games to initialize the policy network.
  • Hand-crafted features — liberty counts, ladder patterns, stone age — engineered by Go experts.
  • Separate networks — a policy network for move selection and a for position evaluation, trained independently.

These dependencies meant AlphaGo could refine human strategies but never fundamentally transcend them. Worse, the approach was Go-specific: there was no path to transfer it to chess, protein folding, or any domain lacking millions of expert demonstrations.

Open in Lab
Compare the training pipelines: AlphaGo needs human expert games, AlphaGo Zero starts from scratch.
The demo wakes as you arrive…

The core idea: self-play as the only teacher

AlphaGo Zero's premise is radical: throw away all human data. Start from random weights. Let the system play against itself, and use the game outcomes to improve. The entire training loop has just three steps:

  1. Play — the current neural network plays games against itself, using Monte Carlo Tree Search to select moves.
  2. Learn — the network is trained to predict the MCTS-improved move probabilities and the eventual game winner.
  3. Evaluate — the new network plays the current best; if it wins more than 55% of games, it replaces the best and becomes the new self-play generator.

This is a form of generalized policy iteration: MCTS acts as a policy improvement operator (it finds stronger moves than the raw network), and self-play acts as a policy evaluation operator (game outcomes evaluate how good the policy really is). The network distills the search's wisdom, and the improved network makes the next search even better — a virtuous cycle.

Open in Lab
Step through the self-play training loop — watch the cycle of play, learn, evaluate.
The demo wakes as you arrive…

Monte Carlo Tree Search: thinking ahead with a neural compass

At each turn, AlphaGo Zero doesn't just use the network's raw move predictions. Instead, it runs 1,600 simulations of MCTS — each simulation a journey from the current board down a . Think of it as mentally rehearsing possible futures before choosing a move.

Each simulation has four stages:

Select — starting from the root (current position), walk down the tree, at each node picking the action that maximizes Q(s,a) + U(s,a). Q is the average value from past simulations; U is an bonus based on the network's prior probability and how rarely this action has been tried. This is the formula — it balances (moves that have worked well) with exploration (moves the network thinks are promising but haven't been tested enough).

Expand — when you reach a leaf node (a position never seen before), expand it: feed the board to the neural network, which returns a policy p (how promising each move is) and a value v (who's likely to win from here).

Backup — carry the value v back up the path, updating Q for every edge you traversed. After many simulations, Q converges to a reliable estimate of each move's worth.

Play — after all 1,600 simulations, choose the move proportional to how often each action was visited at the root. , not raw value, is the selection criterion — a subtle but critical detail that ensures robust move selection.

Open in Lab
Watch MCTS build a search tree. Click "Simulate" to run one cycle of select → expand → backup.
The demo wakes as you arrive…

The PUCT formula: balancing exploration and exploitation

At each node in the search tree, MCTS must decide which branch to explore next. It selects the action that maximizes the sum of two terms: the estimated value Q(s,a) and an exploration bonus U(s,a). The idea is intuitive — follow paths that either have proven high value (exploit) or are promising but under-explored (explore).

at=arg⁡max⁡a[Q(s,a)+cpuct⋅P(s,a)⋅∑bN(s,b)1+N(s,a)]a_t = \arg\max_a \bigl[ Q(s,a) + c_{\text{puct}} \cdot P(s,a) \cdot \frac{\sqrt{\sum_b N(s,b)}}{1 + N(s,a)} \bigr]
PUCT selection rule — the heart of AlphaGo Zero's search — Q(s,a) = average value from past simulations · P(s,a) = neural network's prior probability for action a · N(s,a) = visit count for action a · c_puct = exploration constant. The exploration term shrinks as N(s,a) grows — well-tested moves are exploited, untested moves are explored.
Open in Lab
Drag the exploration constant c_puct and watch how the balance shifts between exploitation and exploration.
The demo wakes as you arrive…

The neural network: one brain, two outputs

Unlike the original AlphaGo, which used two separate networks, AlphaGo Zero uses a single dual-headed residual network. The input is a 19×19×17 : the current board plus the last 8 history positions for each player, plus a color indicator.

The network has three parts:

Residual tower — 19 (or 39) residual blocks, each containing two 3×3 convolutional layers with and . The residual connections (skip connections) allow gradients to flow through the deep stack without vanishing — the same trick that made deep image networks possible.

— a 1×1 followed by a outputting a probability over all 19×19 + 1 = 362 possible moves (including pass). This answers: "what should I play next?"

Value head — a 1×1 convolution, a fully connected with 256 units, and a tanh output giving a single scalar in [-1, +1]. This answers: "who is winning from this position?"

Sharing a residual tower between both heads means the network learns features useful for both tasks simultaneously — a position pattern that helps predict good moves also helps predict who will win.

Open in Lab
Click on any layer to see its role. The shared tower feeds both the policy head and the value head.
The demo wakes as you arrive…

After each self-play game produces a dataset of (state, MCTS policy, game outcome) triples, the network is trained by minimizing a combined :

L=(z−v)2−πTlog⁡p+c∥θ∥2L = (z - v)^2 - \boldsymbol{\pi}^T \log \mathbf{p} + c \|\theta\|^2
AlphaGo Zero's combined loss function — (z - v)² = mean squared error between the predicted value v and the actual game outcome z · -π^T log p = cross-entropy between the MCTS policy π and the network's policy p · c||θ||² = L2 regularization to prevent overfitting

Read this loss as two teaching signals and a safety net. The value loss says: "get better at predicting who wins." The policy loss says: "learn to mimic the search's move choices — they're smarter than your raw guesses." The says: "don't memorize specific positions — generalize."

As training progresses, the network predicts the search's output more accurately, which makes the search itself more powerful (because it starts from a better network), which produces even better training data — the virtuous cycle again.

Dirichlet noise: ensuring the search doesn't get stuck

A danger of self-play is converging to a narrow strategy — mastering one opening and never discovering that a different opening is even better. AlphaGo Zero guards against this by adding to the root node's prior probabilities:

P(s,a)=(1−ϵ)⋅pa+ϵ⋅ηa,η∼Dir(α)P(s, a) = (1 - \epsilon) \cdot p_a + \epsilon \cdot \eta_a, \quad \eta \sim \text{Dir}(\alpha)

With ϵ=0.25\epsilon = 0.25 and α=0.03\alpha = 0.03 for Go's 19×19 board. This means 25% of the root prior comes from random noise rather than the network — enough to ensure the search occasionally tries surprising moves. The noise only applies at the root (not deeper in the tree), so the search remains focused once a direction is chosen.

Think of it as a creative constraint: the network proposes moves, but a random nudge forces it to occasionally consider moves it would otherwise dismiss. Some of these surprises turn out to be genuinely good — and that's how AlphaGo Zero discovers strategies no human ever played.

The idea in code

Simplified AlphaGo Zero self-play and MCTSpython

Simplified to show the idea — not the real implementation.

import numpy as np

class Node:
    """One position in the MCTS search tree."""
    def __init__(self, prior):
        self.prior = prior        # P(s,a) from the network
        self.visit_count = 0      # N(s,a)
        self.value_sum = 0.0      # total value accumulated
        self.children = {}        # action -> Node

    def value(self):
        if self.visit_count == 0:
            return 0
        return self.value_sum / self.visit_count  # Q(s,a)

def puct_score(parent, child, c_puct=1.5):
    """PUCT: balance exploitation (Q) with exploration (U)."""
    exploration = c_puct * child.prior * np.sqrt(parent.visit_count) / (1 + child.visit_count)
    return child.value() + exploration

def mcts_simulate(root, network, game, n_simulations=1600):
    """Run n_simulations of select → expand → backup."""
    for _ in range(n_simulations):
        node = root
        path = [node]
        state = game.clone()

        # SELECT — walk down tree by PUCT
        while node.children:
            action = max(node.children, key=lambda a: puct_score(node, node.children[a]))
            node = node.children[action]
            state.apply(action)
            path.append(node)

        # EXPAND — use network to evaluate leaf
        policy, value = network.predict(state)  # (p, v) = f_θ(s)
        for action, prob in enumerate(policy):
            if state.is_legal(action):
                node.children[action] = Node(prior=prob)

        # BACKUP — propagate value up the path
        for ancestor in reversed(path):
            ancestor.value_sum += value
            ancestor.visit_count += 1
            value = -value  # flip perspective each level

    # Return visit-count distribution as the improved policy
    visits = {a: c.visit_count for a, c in root.children.items()}
    total = sum(visits.values())
    return {a: n / total for a, n in visits.items()}

Results: from random to superhuman in 72 hours

AlphaGo Zero's training curve is one of the most dramatic in history. Starting from random play — literally choosing moves uniformly at random — the system:

  • After 3 hours: plays basic legal Go, capturing stones and securing territory.
  • After 19 hours: rediscovers standard human openings (star-point, 3-4 point).
  • After 36 hours: surpasses AlphaGo Lee (the version that defeated Lee Sedol 4-1).
  • After 72 hours: defeats AlphaGo Lee 100–0, using only 4 TPUs vs. 48 TPUs.
  • After 40 days: achieves an of 5,185, surpassing AlphaGo Master (4,858) which had defeated top professionals 60–0 online.

The final system used only a single machine with 4 TPUs. The original AlphaGo Lee was distributed across many machines with 48 TPUs. Superior performance with 1/12th the compute — that's the power of starting from scratch.

Open in Lab
Watch AlphaGo Zero's Elo rating climb over 40 days, surpassing each previous version.
The demo wakes as you arrive…

Rediscovering — and transcending — human knowledge

Perhaps the most fascinating aspect of AlphaGo Zero is what it learned, not just how well it played. During training, the system independently rediscovered Go concepts that took humans thousands of years to develop: joseki (corner patterns), life-and-death patterns, influence, thickness, and territory. It went through recognizable phases of Go knowledge: early game focused on local fights, then gradually developing whole-board strategy.

But it didn't stop there. AlphaGo Zero also discovered novel strategies that challenged conventional wisdom — moves that human experts initially dismissed as mistakes but later recognized as genuinely superior. It preferred some openings that humans rarely play and discarded others that humans consider standard.

This is the deepest lesson of the paper: a system with no human bias can discover knowledge that human bias actively prevents us from seeing. Centuries of accumulated Go wisdom, while mostly correct, also contained blind spots — strategies dismissed by tradition that were actually optimal.

Why these design choices matter: ablation study

The paper carefully ablates each design choice to show it matters:

  • Dual network vs. separate networks: The single (dual-res) significantly outperforms using separate policy and value networks (sep-res), even though the separate value network predicts expert moves slightly better. Why? Because shared features create a synergy — learning what moves are good also helps learn who's winning.
  • Residual vs. plain convolutional: Residual blocks dramatically improve both prediction accuracy and playing strength, especially as depth increases. Without skip connections, deep networks can't effectively learn.
  • No human data vs. human data: The tabula rasa version (AlphaGo Zero) ultimately surpasses a version trained with human data (AlphaGo Master). Human data gives a head start but introduces bias that the system must later unlearn.
Open in Lab
Compare the four architecture variants — dual-res consistently wins.
The demo wakes as you arrive…

Why it changed everything

  1. 2015

    AlphaGo Fan

    Defeated European champion Fan Hui. Used supervised learning from human games plus Monte Carlo rollouts. The first program to beat a professional Go player.

  2. 2016

    AlphaGo Lee

    Defeated world champion Lee Sedol 4–1 in a historic match. Still dependent on human expert games for bootstrapping. Used 48 TPUs distributed across many machines.

  3. 2017

    AlphaGo Zero

    Tabula rasa — zero human data. Single dual-headed residual network, no rollouts. Defeated AlphaGo Lee 100–0 using only 4 TPUs on a single machine.

  4. 2017

    AlphaGo Master

    Same architecture as Zero but trained with human data. Defeated top professionals 60–0 online. Eventually surpassed by the tabula rasa Zero version.

  5. 2018

    AlphaZero

    Generalized the approach to chess, shogi, and Go with a single algorithm. Proved that tabula rasa self-play works across domains, not just in Go.

AlphaGo Zero showed that the self-play + MCTS + neural network recipe is domain-general. AlphaZero proved it by mastering three different games with the same code. The principle continues to influence , robotics, and scientific discovery today.

CitationSilver, Schrittwieser, Simonyan, Antonoglou, Huang, Guez, Hubert, Baker, Lai, Bolton, Chen, Lillicrap, Hui, Sifre, van den Driessche, Graepel, Hassabis. Mastering the Game of Go Without Human Knowledge. Nature, 2017.

Terms in this paper