Reinforcement Learning2021intermediate11 min read
First Return, Then Explore
عُد أولاً، ثم استكشِف
Ecoffet, A. · Huizinga, J. · Lehman, J. · Stanley, K.O. · Clune, J. — Nature
The problem
algorithms struggle with environments where rewards are sparse or deceptive. In games like Montezuma's Revenge and Pitfall, an might need to perform hundreds of correct actions before receiving any at all. Existing methods — including and curiosity-driven approaches — fail because of two fundamental problems: they forget how to reach previously discovered promising states (detachment), and their exploration noise prevents them from reliably returning to those states in the first place (derailment).
The contribution
Go-Explore: a family of algorithms that directly combats detachment and derailment. It maintains an of all visited states, explicitly returns to promising ones before exploring, and optionally robustifies the discovered trajectories into stochastic-robust policies via . Go-Explore solved all previously unsolved Atari games and achieved orders-of-magnitude improvements on Montezuma's Revenge and Pitfall. It also solved a challenging sparse-reward robotics pick-and-place task.
The impact
Go-Explore demonstrated that the simplest exploration principles — remember states, go back to them, then explore — can dramatically outperform sophisticated intrinsic motivation methods. It was published in Nature in 2021 and reshaped how the RL community thinks about hard-exploration problems. Its archive-based paradigm influenced subsequent work on quality-diversity algorithms and goal-conditioned exploration, showing that structured exploration is more important than clever .
Imagine you are exploring a vast, dark cave system with a notebook, a teleporter, and a flashlight. Every time you find an interesting junction, you jot down its location. When you want to explore more, you don't wander from the entrance again — you teleport back to the most promising junction and shine your flashlight into the unexplored tunnels from there.
Most RL agents are like cavers without a notebook or teleporter. They stumble from the entrance every time, hoping to randomly walk further than before. Go-Explore is the caver with the full toolkit: remember, return, explore.
The problem: why existing exploration fails
Reinforcement learning has achieved superhuman performance in Go, StarCraft II, and Dota II — but all these successes relied on dense, well-shaped reward functions that provide frequent feedback. In many practical problems, rewards are sparse: a household robot might receive a reward only when it successfully places a cup in a cupboard, requiring hundreds of precise motor actions beforehand. Rewards can also be deceptive: following the shortest-distance reward might lead the robot into a dead end.
The standard approach to handling sparse rewards is intrinsic motivation — giving the agent internal rewards for visiting novel states, essentially rewarding curiosity. Methods like ICM (Intrinsic Curiosity Module) and RND (Random Network Distillation) have shown improvements, but they still fail on the hardest benchmarks.
Go-Explore identifies why they fail through two precise failure modes: detachment and derailment.
The solution: remember, return, explore
Go-Explore is built on three remarkably simple principles:
1. Remember — maintain an archive of all interesting states the agent has visited. States are grouped into cells by mapping them to a low-dimensional representation. The archive stores the best trajectory to reach each cell.
2. First return — at the start of each exploration round, select a promising cell from the archive and return to it without exploration noise. In the simplest variant, this is done by restoring the simulator state directly. In the -based variant, a navigates back.
3. Then explore — once the agent is back at the selected state, switch to pure exploration (random actions or a learned exploration policy). Any new states encountered are added to the archive.
By separating returning from exploring, Go-Explore avoids both detachment (the archive never forgets) and derailment (no exploration noise during the return trip).
Cell representations: grouping states wisely
Real environments have far too many unique states to store individually. Go-Explore solves this by mapping similar states to the same cell — a coarsened representation that groups states that look alike from an exploration standpoint.
For Atari games without domain knowledge, the cell representation is a downscaled version of the game frame: the RGB image is converted to grayscale, reduced in resolution, and quantized to fewer pixel depths. The downscaling parameters are optimized dynamically to strike a balance between too many cells (computationally expensive, no aggregation benefit) and too few cells (exploration stalls because distinct states are merged).
For Montezuma's Revenge with domain knowledge, the cell includes the room number, the agent's discretized x-y position, the current level, and which keys the agent holds. This compact representation captures exactly the features relevant for exploration progress.
The choice of cell representation is critical. A good representation groups states that are interchangeable for exploration purposes, while distinguishing states that represent genuinely different exploration frontiers.
Two phases: explore, then robustify
The simplest version of Go-Explore operates in two distinct phases:
Phase 1 — Exploration: The algorithm iteratively selects cells from the archive, restores the simulator to that state, and takes random actions to explore. It can exploit the deterministic nature of simulators by saving and restoring internal state snapshots. The exploration phase does not produce a policy — it produces a collection of high-scoring trajectories (sequences of actions).
Phase 2 — : The trajectories from Phase 1 serve as expert demonstrations for Learning from Demonstrations (LfD). Using a modified "backward algorithm," a policy is trained to replicate and improve upon these trajectories in a stochastic (with sticky actions that add randomness). This phase converts brittle, deterministic trajectories into robust, generalizable policies.
The separation is powerful: Phase 1 can use any means to explore (including determinism and simulator restoring), because Phase 2 will ensure the final policy handles real-world stochasticity.
Choosing where to explore: cell selection
Not all cells in the archive are equally worth returning to. Go-Explore selects cells probabilistically based on how promising they are for further exploration. The key intuition is to prefer cells that have been visited fewer times during the explore step, because they likely have unexplored territory nearby.
The selection weight for a cell is inversely proportional to the square root of the number of times it was chosen for exploration:
For Montezuma's Revenge with domain knowledge, Go-Explore adds spatial heuristics to the selection. Cells with fewer horizontal neighbours already in the archive get a bonus, because they sit at the edge of explored territory. A key bonus also prioritizes cells that hold the most keys at a given location. These domain-aware weights dramatically accelerate exploration in structured environments.
Policy-based Go-Explore: no simulator needed
Restoring simulator states is efficient, but not always possible — especially in the real world. Policy-based Go-Explore replaces the simulator restore with a learned goal-conditioned policy that navigates back to the selected cell by following the archived trajectory as a series of sub-goals.
This variant brings two key advantages. First, having a policy available during the explore step dramatically improves exploration: instead of random actions, the agent can be given new goals to move toward — including cells not yet in the archive. Experiments showed that policy-based exploration discovers roughly four times more cells than random actions on both Montezuma's Revenge and Pitfall.
Second, because the policy trains in a stochastic environment from the start, there is no need for a separate robustification phase. The policy naturally learns to handle environmental randomness during .
The policy is trained with PPO, enhanced by to extract maximum value from the rare successful trajectories discovered early in training. An injection mechanism increases exploration only when the agent gets stuck, avoiding the global derailment that comes from a fixed .
Results: solving the unsolvable
Go-Explore's results are dramatic. On the Atari :
On Montezuma's Revenge, Go-Explore achieved a score of 43,791 after robustification with the downscaled representation — quadrupling the previous state of the art of 11,618 and far exceeding average human performance of 4,753. With domain knowledge, the score reached 1,731,645, surpassing even the human world record of 1.2 million.
On Pitfall, Go-Explore scored 6,954 — the first algorithm to score any points on this game under proper stochastic evaluation, while the previous state of the art was zero.
The exploration phase alone (before robustification) finds superhuman trajectories for all 55 Atari games, including beating state-of-the-art scores in 83.6% of games.
On a challenging robotics pick-and-place task — where a robot arm must grasp an object and place it in one of four shelves (two behind latched doors) — Go-Explore achieved a 99% success rate. Standard PPO found zero rewards after a billion frames. Count-based intrinsic motivation with the same domain knowledge also found zero rewards, failing even to learn reliable grasping due to derailment.
Why this matters: simplicity beats complexity
The most striking aspect of Go-Explore is the contrast between the simplicity of its mechanisms and the magnitude of its results. Years of research on sophisticated intrinsic motivation methods — prediction error, curiosity modules, count-based bonuses — produced incremental gains. Go-Explore, with its straightforward "remember, return, explore" loop, produced orders-of-magnitude improvements.
This suggests that the bottleneck in hard-exploration problems was not a lack of clever reward signals, but a structural failure: algorithms were not systematically returning to promising states before exploring. The exploration community was optimizing the wrong thing — how to encourage exploration at each step, rather than how to structure the overall exploration process.
The paper's insight extends beyond Go-Explore itself. The decomposition of exploration into returning and exploring is a general principle that can be combined with any RL algorithm. It represents a shift from "explore everywhere a little" to "return somewhere specific, then explore a lot" — a paradigm that more closely mirrors how humans and animals explore their environments.
Core loop in code
Simplified to show the idea — not the real implementation.
archive = {initial_cell: (initial_state, empty_trajectory)}
while not done:
# 1. SELECT — pick a promising cell
cell = select_cell(archive, weights)
# 2. GO — return to that state (no exploration noise!)
restore_simulator_state(archive[cell].state)
# 3. EXPLORE — take random actions from the restored state
for step in range(explore_steps):
action = random_action()
new_state, reward = env.step(action)
new_cell = map_to_cell(new_state)
# Update archive if this cell is new or trajectory is better
if new_cell not in archive or is_better(trajectory, archive[new_cell]):
archive[new_cell] = (new_state, current_trajectory)Context: the exploration timeline
2015
DQN on Atari (Mnih et al.)
Deep Q-Networks achieved human-level performance on many Atari games but failed on hard-exploration games like Montezuma's Revenge, scoring near zero.
2017
ICM — Intrinsic Curiosity Module
Pathak et al. introduced curiosity-driven exploration using prediction error as intrinsic reward. Improved on some hard games but still far from human on Montezuma's Revenge.
2018
RND — Random Network Distillation
Burda et al. used prediction error on random features as intrinsic reward. Reached ~10,000 on Montezuma's Revenge — best at the time, but still below human.
2019
Go-Explore preprint (Ecoffet et al.)
First preprint introducing Go-Explore with simulator-state restore. Demonstrated massive improvements but the community initially debated reliance on determinism.
2020
Agent57 (Badia et al.)
DeepMind's Agent57 also achieved superhuman on all Atari games concurrently, but under easier evaluation conditions. A different approach validating that all games are solvable.
2021
Go-Explore published in Nature
The full paper, including policy-based Go-Explore and robotics results, was published in Nature. Established "first return, then explore" as a fundamental exploration principle.
Go-Explore's legacy is not just its scores, but its diagnosis: detachment and derailment are named, understood, and solvable. Any future exploration algorithm that does not address these failure modes is likely leaving performance on the table. The paper showed that structured exploration — not smarter rewards — is the key to unlocking hard-exploration problems.
CitationEcoffet, Huizinga, Lehman, Stanley, Clune. First return, then explore. Nature, 2021.
Terms in this paper
- Explorationالاستكشاف (تجربة أفعال جديدة)
- Sparse Rewardالمكافأة الشحيحة
- Intrinsic Motivationالدافعية الجوهرية
- Reinforcement Learningالتعلم المعزز
- Archiveالأرشيف
- Policy Gradientتدرج السياسة التشغيلية
- Imitation Learningالتعلّم بالتقليد
- Deep Reinforcement Learningالتعلم العميق بالتعزيز
- Reward Shapingصياغة وهندسة دالة المكافأة
- Goal-Conditioned Policyالسياسة المشروطة بالأهداف
- Deterministic Policyالسياسة الحتمية
- Experience Replayإعادة تشغيل التجارب
- Environmentالبيئة التفاعلية
- Episodeجولة تفاعلية كاملة
- Action Spaceفضاء الأفعال