Reinforcement Learning2018intermediate11 min read

World Models

نماذج العالَم

Ha, D. · Schmidhuber, J. — NeurIPS

The problem

By 2018, model-free agents could solve complex tasks but required millions of interactions — expensive, slow, and impractical for real-world robotics. These agents treated the environment as a black box: observe, act, receive , repeat. They had no internal understanding of how the world works, no ability to imagine consequences before acting, and no way to practice offline. Humans, by contrast, build rich mental models of the world and rehearse actions mentally before executing them.

The contribution

A three-component architecture that gives RL agents an internal . The Vision model (V) is a that compresses high-dimensional observations into a compact latent code. The Memory model (M) is an MDN-RNN that learns temporal dynamics in — predicting the probability distribution of future states. The Controller (C) is a tiny linear (867 parameters) trained with evolutionary strategies. The key breakthrough: agents can be trained entirely inside hallucinated "dreams" generated by the world model and then transferred to the real environment.

The impact

World Models revived model-based reinforcement learning and established the template that modern systems still follow: learn a compressed representation, predict dynamics in latent space, and plan using imagined rollouts. It directly inspired PlaNet, the Dreamer family (v1–v3), and MuZero. The idea that agents can "dream" to learn — in imagination rather than reality — became a foundational concept in sample-efficient RL and is now central to robotics, game AI, and autonomous driving research.

Imagine learning to ride a bicycle. Before you even touch the pedals, you can picture yourself pedaling, leaning into a turn, and recovering from a wobble. Your brain runs a mental simulation — a private rehearsal that costs no scraped knees.

Now imagine a robot that can do the same. It watches the world through a camera, compresses each frame into a quick sketch (Vision), uses past sketches to imagine what comes next (Memory), and then a tiny coach decides which way to steer (Controller).

The revolutionary idea in this paper: the robot can train entirely inside its own imagination — pedaling through dreamed roads — and then ride competently in the real world.

The architecture: Vision, Memory, and Controller

The core insight of World Models is to decompose an RL into three specialized modules, each solving a different subproblem. This decomposition mirrors how cognitive scientists describe human perception: we compress sensory input, maintain a predictive model of the world, and make decisions using both.

Vision (V) — a Variational — compresses each raw observation (a 64×64 RGB image) into a compact latent vector ztz_t of just 32 dimensions. This is like reducing a photograph to a quick sketch that captures the essential layout while discarding pixel-level noise.

Memory (M) — an MDN-RNN (a recurrent network with a output) — receives the current latent ztz_t and action ata_t, then predicts the distribution of the next latent zt+1z_{t+1}. It maintains a hth_t that accumulates temporal context. Think of it as an internal crystal ball that says "given what I've seen and what I just did, here is what might happen next."

Controller (C) — a single linear layer — takes the concatenation of ztz_t (what I see now) and hth_t (what I remember and predict) and outputs the action ata_t. With only 867 parameters, it is deliberately tiny so that the world model carries the complexity.

Open in Lab
The full V-M-C pipeline: raw pixels flow through the Vision encoder to become a latent sketch, the Memory model predicts the future, and the Controller picks an action.
The demo wakes as you arrive…

Vision (V): compressing the world into a sketch

The Vision model is a convolutional Variational Autoencoder. The takes a 64×64 RGB frame and maps it to the parameters of a Gaussian distribution in a 32-dimensional latent space: a mean μ\mu and a log-variance log⁡σ2\log \sigma^2. A sample ztz_t is drawn from this distribution using the reparameterization trick. The decoder reconstructs the frame from ztz_t.

Why a VAE and not a plain autoencoder? Because the VAE learns a smooth, continuous latent space. Nearby points in this space correspond to visually similar scenes. This smoothness is critical for the Memory model — it needs to predict trajectories through latent space, and a jumpy, fragmented space would make prediction nearly impossible. The VAE's KL divergence penalty acts like a gentle force that prevents the latent space from collapsing into disconnected islands.

The result: a 32-number "sketch" that captures the essential structure of each frame — the position of the car, the shape of the road, the presence of obstacles — while discarding irrelevant pixel details.

LVAE=Eq(z∣x)[log⁡p(x∣z)]⏟reconstruction−DKL(q(z∣x)∥p(z))⏟smoothness\mathcal{L}_{\text{VAE}} = \underbrace{\mathbb{E}_{q(z|x)} [\log p(x|z)]}_{\text{reconstruction}} - \underbrace{D_{KL}(q(z|x) \| p(z))}_{\text{smoothness}}
VAE loss — balancing faithful reconstruction with smooth latent structure — The first term rewards accurate reconstruction: the decoder should recover the original image from the latent code. The second term penalizes deviation from a standard Gaussian prior, enforcing a smooth, regular latent space. Together they ensure that similar scenes map to nearby points and that the space has no "holes" that would confuse the Memory model.
Open in Lab
Drag in the 2D latent space to see how the VAE decoder reconstructs different scenes. Nearby points produce similar frames — evidence of a smooth latent structure.
The demo wakes as you arrive…

Memory (M): predicting the future in latent space

With a compact representation of the present (ztz_t), the next challenge is predicting what happens next. The Memory model is an whose output is not a single prediction but a probability distribution over possible futures.

Specifically, the MDN-RNN outputs the parameters of a Mixture of Gaussians: for each of the 32 latent dimensions, it produces KK Gaussian components (means μk\mu_k, standard deviations σk\sigma_k, and mixing weights πk\pi_k). This means the model can express uncertainty and multimodality — if the car is approaching a fork, the model can assign probability to both branches rather than blurring between them.

At each step, the LSTM receives the concatenation of ztz_t and ata_t and updates its hidden state hth_t. The hidden state accumulates temporal context: it remembers the trajectory so far, the velocity, the rhythm of the track. The MDN head then maps hth_t to the mixture parameters.

P(zt+1∣at,zt,ht)=∑k=1Kπk  N(zt+1∣μk,σk2)P(z_{t+1} \mid a_t, z_t, h_t) = \sum_{k=1}^{K} \pi_k \; \mathcal{N}(z_{t+1} \mid \mu_k, \sigma_k^2)
MDN-RNN prediction — a mixture of Gaussians over future states — The model does not predict a single future state but a distribution: KK possible outcomes weighted by πk\pi_k. Each Gaussian component captures a different plausible next state. During dream training, samples are drawn from this distribution to generate imagined rollouts.
Open in Lab
See how the MDN-RNN predicts a distribution of possible future states. Each Gaussian component represents a different possible outcome — the model captures uncertainty.
The demo wakes as you arrive…

Controller (C): the tiny decision-maker

The Controller is deliberately minimal: a single linear layer that maps the concatenated input [zt,ht][z_t, h_t] to an action vector.

at=Wc[ztht]+bca_t = W_c \begin{bmatrix} z_t \\ h_t \end{bmatrix} + b_c
Controller — a linear policy mapping perception and memory to action — WcW_c and bcb_c are the controller's only parameters. For CarRacing with Nz=32N_z=32 dimensions and Nh=256N_h=256 hidden units, the action has 3 components (steering, acceleration, braking), giving just (32+256)×3+3=867(32+256) \times 3 + 3 = 867 parameters. This is trained with CMA-ES, an evolution strategy that efficiently searches low-dimensional parameter spaces.

Why is the Controller so small? Because all the complexity lives in the world model. The VAE has already extracted the meaningful features. The LSTM has already compressed temporal patterns. All the Controller needs to do is draw a simple linear boundary in this well-organized space. Think of it as a CEO who receives a one-page executive summary from two brilliant analysts — the decision becomes straightforward.

This separation has a practical benefit: since the Controller has so few parameters, it can be trained with evolutionary strategies like instead of gradient descent. Evolution is gradient-free — it does not require through the world model, so V, M, and C can be trained independently.

Learning inside a dream

This is the paper's most striking contribution. Once the world model (V + M) is trained on observations from the real environment, it can generate synthetic experience — the agent can explore a hallucinated version of the environment entirely inside its own Memory model.

The process works like this: at each dream step, the Controller takes an action based on the current latent state and hidden state. The MDN-RNN then samples the next latent state from its predicted distribution. The Controller reacts to this hallucinated state, takes another action, and the cycle continues — all without touching the real environment.

A crucial detail: the τ\tau controls the stochasticity of the dream. At τ=1\tau=1, dreams match the real environment's randomness. At τ>1\tau > 1, dreams become more chaotic and unpredictable — the authors found this acts as a form of regularization, making the Controller more robust because it trains against worst-case scenarios.

Open in Lab
Watch the agent train inside its own dream. Adjust the temperature τ to see how dream stochasticity affects training quality.
The demo wakes as you arrive…

Training pipeline: three stages

The entire system trains in three independent stages, each taking less than an hour on a single GPU.

Stage 1 — Collect rollouts. A random policy interacts with the environment, recording observations and actions. No reward signal is needed for this step — the world model is trained purely on sensory experience.

Stage 2 — Train V and M. The VAE learns to compress frames, then the MDN-RNN learns temporal dynamics in the latent space. Both are trained with unsupervised objectives (reconstruction + KL for the VAE, negative log-likelihood for the MDN-RNN).

Stage 3 — Train C. The Controller is evolved using CMA-ES, either by interacting with the real environment (V + M + C mode) or entirely inside dreams (Dream mode). Each candidate controller is evaluated across multiple rollouts, and CMA-ES updates the parameter distribution toward higher rewards.

Open in Lab
Click through each stage of the training pipeline. Each module trains independently — the world model learns from raw experience, the controller learns from the world model.
The demo wakes as you arrive…

Experiments: CarRacing and VizDoom

The paper evaluates on two environments. CarRacing-v0 is a continuous-control driving task with procedurally generated tracks. The agent sees a top-down view and must visit as many tiles as possible. The full V+M+C model scored 906 (out of ~950 maximum), significantly outperforming the previous best methods. The Vision-only baseline (V+C, no memory) scored 632 — proving that memory and prediction are crucial for handling curves and planning ahead.

VizDoom Take Cover is a first-person survival game where the agent must dodge fireballs thrown by monsters. Here, the paper demonstrated dream training: the controller was trained entirely inside the world model's imagination, then transferred to the real environment, achieving over 1100 timesteps of survival — competitive with agents trained with direct environment access.

Open in Lab
Compare frames from the real environment (left) with the agent's dream reconstruction (right). The dream captures the essential structure while softening fine details.
The demo wakes as you arrive…

Why this matters: from dreaming to modern world models

World Models introduced a blueprint that the entire field followed. The core recipe — learn a latent representation, predict dynamics in that space, plan using imagined trajectories — appears in virtually every modern system.

The paper also introduced a profound conceptual shift: the agent does not directly perceive reality. It only sees what its world model shows it. This "Plato's Cave" perspective highlights both the power and the danger: a good world model enables efficient learning, but a flawed model can create systematic blind spots that the agent cannot detect.

The limitations were clear too. The MDN-RNN struggled with long horizons. The VAE discarded details that sometimes mattered. The CMA-ES controller did not scale to high-dimensional action spaces. But these were engineering limitations, not conceptual ones — and the follow-up work (PlaNet, Dreamer, MuZero) addressed each in turn.

Legacy: from dreaming agents to foundation world models

  1. 2018

    World Models (this paper)

    VAE + MDN-RNN + linear controller. First demonstration of training an agent entirely in its own dream. Solved CarRacing and VizDoom.

  2. 2019

    PlaNet — Learning Latent Dynamics

    Replaced the MDN-RNN with a Recurrent State-Space Model (RSSM), splitting the hidden state into deterministic and stochastic components. Enabled online planning via cross-entropy method in latent space.

  3. 2020

    Dreamer v1 — Dream to Control

    Combined RSSM with actor-critic learning inside imagined trajectories. Achieved human-level performance on 20 continuous control tasks from pixels.

  4. 2020

    MuZero — planning without a simulator

    Learned a latent dynamics model end-to-end with tree search. Mastered Chess, Go, Shogi, and Atari without ever being given the rules of the game.

  5. 2023

    Dreamer v3 — mastering diverse domains

    A single algorithm with fixed hyperparameters that achieves human-level or better across 150+ tasks spanning robotics, Atari, DMLab, and Minecraft — the culmination of the World Models vision.

The trajectory from World Models to Dreamer v3 is a story of engineering the same core idea: learn a latent model of the world, imagine rollouts, and optimize behavior in imagination. Each successor addressed a specific limitation — better state models, better policy learning, better scaling — but the conceptual foundation laid in 2018 remains the blueprint.

CitationHa, Schmidhuber. Recurrent World Models Facilitate Policy Evolution. NeurIPS, 2018.

Terms in this paper