Reinforcement Learning2017intermediate10 min read

Curiosity-Driven Exploration by Self-Supervised Prediction

الاستكشاف المدفوع بالفضول عبر التنبؤ ذاتي الإشراف

Pathak, D. · Agrawal, P. · Efros, A. A. · Darrell, T. — ICML

The problem

In many real-world RL tasks, extrinsic rewards are extremely sparse or absent altogether. An playing Super Mario Bros gets no score until it reaches a distant goal. Random (epsilon-greedy) almost never stumbles upon that goal. Previous curiosity methods predicted raw pixels, but pixel is hard, wasteful, and easily fooled by irrelevant visual — like leaves blowing in the wind or static on a TV screen — trapping the agent in an endless loop of artificial surprise.

The contribution

The Intrinsic Curiosity Module (ICM): a self-supervised architecture with two neural networks working together. An learns a that captures only agent-relevant aspects of the , filtering out visual noise. A predicts the next in that learned space, and its prediction error becomes the intrinsic . The agent explores by seeking states where its forward model fails — genuine novelty, not pixel noise. Combined with A3C, the agent explores VizDoom mazes and plays Super Mario Bros without any extrinsic reward.

The impact

ICM established learned-feature-space curiosity as the standard paradigm for in deep RL. It directly inspired RND (Random Network Distillation), Go-Explore, and dozens of follow-up methods. The key insight — predict in a task-relevant feature space, not raw pixels — became a design principle across exploration, world models, and learning. The paper also pioneered evaluating RL by testing on unseen levels, a practice now standard in the field.

Imagine you just moved to a new city. A tourist with a guidebook (extrinsic reward) only visits the landmarks circled on the map — if nothing is circled, she stays at the hotel.

A curious explorer is different: she wanders toward streets she can't predict what's around the corner. Once she's seen a neighborhood, her surprise fades and she moves on to the next unknown block. Crucially, she has learned to ignore the flickering neon signs and swaying trees — they're unpredictable but irrelevant. She only tracks the things she can interact with: doors she can open, paths she can take.

ICM gives an RL agent exactly this explorer's instinct: a forward model that predicts what happens next, and an inverse model that teaches the agent to only pay attention to things its actions can actually affect.

The problem: when there is no treasure map

Standard RL agents learn by trial and error, guided by rewards from the environment. But what happens when rewards are extremely sparse — or absent altogether?

Consider Super Mario Bros: the agent gets a score only for reaching a distant flag. With random exploration (epsilon-greedy), the probability of stumbling upon that flag through thousands of random button presses is essentially zero. The agent never gets a learning signal, so it never improves.

Humans solve this effortlessly — a child explores a playground not because anyone promised a reward, but out of sheer curiosity. This paper formalizes that intuition into a computable intrinsic reward.

Open in Lab
Compare how much of the grid an agent explores under dense, sparse, and no-reward settings. Curiosity-driven agents explore far more than random ones.
The demo wakes as you arrive…

The pixel trap: why raw prediction fails

A natural first idea for curiosity: predict what the next frame of pixels will look like after taking an action, and use the prediction error as reward. But this has two fatal problems.

First, predicting thousands of pixel values is extremely hard — the model wastes capacity on irrelevant details like exact textures and lighting.

Second, and more fundamentally, some parts of the environment are inherently unpredictable but irrelevant to the agent. Imagine a TV screen showing static noise in the corner of a game level. A pixel-prediction model can never predict that noise, so the agent remains permanently curious about the TV — trapped staring at randomness while ignoring the actual game.

The paper identifies three categories of environmental changes: (1) things the agent can control — like opening a door; (2) things it cannot control but that affect it — like another agent's vehicle; and (3) things it cannot control and that don't affect it — like blowing leaves or TV static. A good curiosity signal should capture (1) and (2) but ignore (3).

Open in Lab
Toggle between pixel-space and feature-space prediction to see how the agent handles irrelevant visual noise.
The demo wakes as you arrive…

The solution: Intrinsic Curiosity Module (ICM)

ICM solves the pixel trap with an elegant two-part design. Instead of predicting raw pixels, it first transforms observations into a compact feature space, then makes predictions there. The trick is how that feature space is learned.

The module has three components working together: a feature ϕ\phi that maps raw states to compact representations, an inverse dynamics model that learns what features matter, and a forward dynamics model whose prediction error becomes the curiosity reward.

Open in Lab
The ICM architecture: click on each component to see its role. The inverse model trains the encoder, and the forward model generates curiosity reward.
The demo wakes as you arrive…

Step 1: the inverse model learns what matters

The inverse dynamics model takes the feature representations of two consecutive states ϕ(st)\phi(s_t) and ϕ(st+1)\phi(s_{t+1}), and predicts what action the agent took to move between them.

Why does this help? If the feature encoder includes information about blowing leaves or TV static, that information is useless for predicting the agent's action — leaves blow regardless of what the agent does. So the inverse model's pressure forces the encoder to keep only features that are relevant to the agent's actions. This is the key insight that makes ICM robust to visual distractors.

a^t=g(ϕ(st),  ϕ(st+1);  θI)\hat{a}_t = g\bigl(\phi(s_t),\; \phi(s_{t+1});\; \theta_I\bigr)
Inverse dynamics model — predicting the agent's action from consecutive feature states — Given features of two consecutive states, predict what action the agent took. This forces the encoder to represent only action-relevant information. Trained with cross-entropy loss for discrete actions.

Step 2: the forward model generates curiosity

The forward dynamics model takes the current state features ϕ(st)\phi(s_t) and the action ata_t, and predicts what the features of the next state ϕ(st+1)\phi(s_{t+1}) will look like.

The prediction is made in the learned feature space, not in pixel space. Since that feature space already filters out irrelevant noise (thanks to the inverse model), the forward model's prediction errors reflect genuine novelty — places the agent hasn't explored yet — rather than uncontrollable randomness.

ϕ^(st+1)=f(ϕ(st),  at;  θF)\hat{\phi}(s_{t+1}) = f\bigl(\phi(s_t),\; a_t;\; \theta_F\bigr)
Forward dynamics model — predicting the next feature state — Given current features and action, predict the next state's features. The squared error between the prediction and reality becomes the curiosity reward.
rti=η2 ∥ϕ^(st+1)−ϕ(st+1)∥22r^i_t = \frac{\eta}{2}\,\bigl\|\hat{\phi}(s_{t+1}) - \phi(s_{t+1})\bigr\|_2^2
Intrinsic curiosity reward — The intrinsic reward is the scaled prediction error in feature space. High error means the agent encountered something genuinely new. The scaling factor η controls the reward magnitude.
Open in Lab
Watch how the forward model's prediction error decreases as the agent revisits familiar states and spikes when encountering new ones.
The demo wakes as you arrive…

The full picture: joint optimization

The agent combines extrinsic reward (if any) with intrinsic curiosity reward at every time step: rt=rti+rter_t = r^i_t + r^e_t. The is trained with A3C to maximize the expected sum of this combined reward.

All three components — encoder, inverse model, and forward model — are trained jointly with the policy. The overall function balances three objectives: policy learning, inverse model accuracy, and forward model accuracy.

min⁡θP, θI, θF, θE  [−λ Eπ(st;θP)[∑trt]  +  (1−β) LI  +  β LF]\min_{\theta_P,\,\theta_I,\,\theta_F,\,\theta_E}\; \Bigl[-\lambda\,\mathbb{E}_{\pi(s_t;\theta_P)}\bigl[\textstyle\sum_t r_t\bigr] \;+\;(1-\beta)\,L_I \;+\; \beta\,L_F\Bigr]
Joint optimization objective — λ weighs the policy gradient loss against the prediction losses. β (between 0 and 1) balances the inverse model loss LIL_I against the forward model loss LFL_F. The policy gradient is NOT backpropagated through the forward model to prevent the agent from learning to reward itself.

Experiments: VizDoom and Super Mario Bros

The paper evaluates ICM in three settings, each testing a different claim about curiosity.

Sparse extrinsic reward: In VizDoom's 3D navigation maze, the agent must reach a goal from spawn points at varying distances. With dense spawning (close to goal), vanilla A3C sometimes succeeds. With sparse spawning (270+ steps away), only ICM+A3C reliably finds the goal — scoring 100% vs 0% for baseline A3C.

No extrinsic reward: With zero environmental reward, ICM-driven Mario crosses over 30% of Level-1, learning to jump over pipes, dodge enemies, and avoid pits — all discovered purely through curiosity. In VizDoom, the curious agent explores 5+ rooms while a random agent barely leaves the starting room.

Generalization: A Mario agent pre-trained with curiosity on Level-1 explores Level-2 faster than an agent trained from scratch on Level-2. The skills it learned (running, jumping, dodging) transfer to new levels. In VizDoom, a curiosity-pretrained agent fine-tuned on a new map with new textures outperforms an agent trained from scratch.

Open in Lab
Watch a curiosity-driven agent explore a grid world. Cells turn green when visited. The agent seeks high prediction error (bright cells) and avoids already-explored areas.
The demo wakes as you arrive…

Robustness: surviving the noisy TV

To test robustness against irrelevant stochasticity, the authors replaced 40% of the agent's visual observation with white noise — simulating the "noisy TV" problem. Results were striking: ICM with learned features succeeded just as well as without the noise. The pixel-prediction baseline (ICM-pixels) collapsed entirely.

This confirms the core thesis: the inverse-dynamics-trained feature space filters out uncontrollable factors. The agent literally cannot see the noise because its encoder learned that noise carries zero information about actions.

Legacy: from ICM to modern exploration

ICM established two principles that shaped everything that followed in exploration research.

First, predict in a learned feature space, not raw observations. This principle appears in RND (which uses a random network as the feature encoder), BYOL-Explore, and virtually every modern intrinsic motivation method.

Second, evaluate generalization. Before ICM, RL papers almost never tested on unseen environments. By showing that curiosity-learned skills transfer to new levels and maps, the paper pushed the community toward generalization as a first-class evaluation criterion.

The paper's influence extends beyond exploration: the idea that an inverse dynamics model provides a good self-supervised training signal for visual representations has been adopted in robotics, video understanding, and world models.

  1. 1991

    Schmidhuber's curiosity and compression

    Proposed using prediction error and compression progress as intrinsic motivation — the intellectual ancestor of all curiosity-driven exploration methods.

  2. 2016

    VIME: Variational Information Maximizing Exploration

    Used information gain about environment dynamics as intrinsic reward. Effective but computationally expensive and struggled with high-dimensional observations.

  3. 2017

    ICM (this paper)

    Introduced feature-space curiosity with the inverse-forward model architecture. First demonstration of learning to play games with zero extrinsic reward.

  4. 2018

    RND: Random Network Distillation

    Simplified ICM by replacing the inverse model with a fixed random network as the feature target. Achieved first superhuman results on Montezuma's Revenge.

  5. 2019

    Go-Explore: archive-based exploration

    Combined curiosity with an explicit archive of promising states, addressing the "boredom" limitation of pure curiosity methods like ICM.

  6. 2019

    Large-Scale Study of Curiosity-Driven Learning

    Burda et al. scaled ICM to 54 Atari games and found that curiosity alone — with no extrinsic reward — learns meaningful behaviors across diverse environments.

Today, intrinsic curiosity is a standard component in the exploration toolkit. Whether through learned features (ICM), random networks (RND), or state counting, the core idea remains: give the agent a reason to seek the unknown. ICM was the paper that made this practical at scale with raw visual observations.

CitationPathak, Agrawal, Efros, Darrell. Curiosity-driven Exploration by Self-supervised Prediction. ICML, 2017.

Terms in this paper