Reinforcement Learning2018intermediate11 min read
World Models
نماذج العالَم
Ha, D. · Schmidhuber, J. — NeurIPS
The problem
By 2018, model-free agents could solve complex tasks but required millions of interactions — expensive, slow, and impractical for real-world robotics. These agents treated the environment as a black box: observe, act, receive , repeat. They had no internal understanding of how the world works, no ability to imagine consequences before acting, and no way to practice offline. Humans, by contrast, build rich mental models of the world and rehearse actions mentally before executing them.
The contribution
A three-component architecture that gives RL agents an internal . The Vision model (V) is a that compresses high-dimensional observations into a compact latent code. The Memory model (M) is an MDN-RNN that learns temporal dynamics in — predicting the probability distribution of future states. The Controller (C) is a tiny linear (867 parameters) trained with evolutionary strategies. The key breakthrough: agents can be trained entirely inside hallucinated "dreams" generated by the world model and then transferred to the real environment.
The impact
World Models revived model-based reinforcement learning and established the template that modern systems still follow: learn a compressed representation, predict dynamics in latent space, and plan using imagined rollouts. It directly inspired PlaNet, the Dreamer family (v1–v3), and MuZero. The idea that agents can "dream" to learn — in imagination rather than reality — became a foundational concept in sample-efficient RL and is now central to robotics, game AI, and autonomous driving research.
Imagine learning to ride a bicycle. Before you even touch the pedals, you can picture yourself pedaling, leaning into a turn, and recovering from a wobble. Your brain runs a mental simulation — a private rehearsal that costs no scraped knees.
Now imagine a robot that can do the same. It watches the world through a camera, compresses each frame into a quick sketch (Vision), uses past sketches to imagine what comes next (Memory), and then a tiny coach decides which way to steer (Controller).
The revolutionary idea in this paper: the robot can train entirely inside its own imagination — pedaling through dreamed roads — and then ride competently in the real world.
The architecture: Vision, Memory, and Controller
The core insight of World Models is to decompose an RL into three specialized modules, each solving a different subproblem. This decomposition mirrors how cognitive scientists describe human perception: we compress sensory input, maintain a predictive model of the world, and make decisions using both.
Vision (V) — a Variational — compresses each raw observation (a 64×64 RGB image) into a compact latent vector of just 32 dimensions. This is like reducing a photograph to a quick sketch that captures the essential layout while discarding pixel-level noise.
Memory (M) — an MDN-RNN (a recurrent network with a output) — receives the current latent and action , then predicts the distribution of the next latent . It maintains a that accumulates temporal context. Think of it as an internal crystal ball that says "given what I've seen and what I just did, here is what might happen next."
Controller (C) — a single linear layer — takes the concatenation of (what I see now) and (what I remember and predict) and outputs the action . With only 867 parameters, it is deliberately tiny so that the world model carries the complexity.
Vision (V): compressing the world into a sketch
The Vision model is a convolutional Variational Autoencoder. The takes a 64×64 RGB frame and maps it to the parameters of a Gaussian distribution in a 32-dimensional latent space: a mean and a log-variance . A sample is drawn from this distribution using the reparameterization trick. The decoder reconstructs the frame from .
Why a VAE and not a plain autoencoder? Because the VAE learns a smooth, continuous latent space. Nearby points in this space correspond to visually similar scenes. This smoothness is critical for the Memory model — it needs to predict trajectories through latent space, and a jumpy, fragmented space would make prediction nearly impossible. The VAE's KL divergence penalty acts like a gentle force that prevents the latent space from collapsing into disconnected islands.
The result: a 32-number "sketch" that captures the essential structure of each frame — the position of the car, the shape of the road, the presence of obstacles — while discarding irrelevant pixel details.
Memory (M): predicting the future in latent space
With a compact representation of the present (), the next challenge is predicting what happens next. The Memory model is an whose output is not a single prediction but a probability distribution over possible futures.
Specifically, the MDN-RNN outputs the parameters of a Mixture of Gaussians: for each of the 32 latent dimensions, it produces Gaussian components (means , standard deviations , and mixing weights ). This means the model can express uncertainty and multimodality — if the car is approaching a fork, the model can assign probability to both branches rather than blurring between them.
At each step, the LSTM receives the concatenation of and and updates its hidden state . The hidden state accumulates temporal context: it remembers the trajectory so far, the velocity, the rhythm of the track. The MDN head then maps to the mixture parameters.
Controller (C): the tiny decision-maker
The Controller is deliberately minimal: a single linear layer that maps the concatenated input to an action vector.
Why is the Controller so small? Because all the complexity lives in the world model. The VAE has already extracted the meaningful features. The LSTM has already compressed temporal patterns. All the Controller needs to do is draw a simple linear boundary in this well-organized space. Think of it as a CEO who receives a one-page executive summary from two brilliant analysts — the decision becomes straightforward.
This separation has a practical benefit: since the Controller has so few parameters, it can be trained with evolutionary strategies like instead of gradient descent. Evolution is gradient-free — it does not require through the world model, so V, M, and C can be trained independently.
Learning inside a dream
This is the paper's most striking contribution. Once the world model (V + M) is trained on observations from the real environment, it can generate synthetic experience — the agent can explore a hallucinated version of the environment entirely inside its own Memory model.
The process works like this: at each dream step, the Controller takes an action based on the current latent state and hidden state. The MDN-RNN then samples the next latent state from its predicted distribution. The Controller reacts to this hallucinated state, takes another action, and the cycle continues — all without touching the real environment.
A crucial detail: the controls the stochasticity of the dream. At , dreams match the real environment's randomness. At , dreams become more chaotic and unpredictable — the authors found this acts as a form of regularization, making the Controller more robust because it trains against worst-case scenarios.
Training pipeline: three stages
The entire system trains in three independent stages, each taking less than an hour on a single GPU.
Stage 1 — Collect rollouts. A random policy interacts with the environment, recording observations and actions. No reward signal is needed for this step — the world model is trained purely on sensory experience.
Stage 2 — Train V and M. The VAE learns to compress frames, then the MDN-RNN learns temporal dynamics in the latent space. Both are trained with unsupervised objectives (reconstruction + KL for the VAE, negative log-likelihood for the MDN-RNN).
Stage 3 — Train C. The Controller is evolved using CMA-ES, either by interacting with the real environment (V + M + C mode) or entirely inside dreams (Dream mode). Each candidate controller is evaluated across multiple rollouts, and CMA-ES updates the parameter distribution toward higher rewards.
Experiments: CarRacing and VizDoom
The paper evaluates on two environments. CarRacing-v0 is a continuous-control driving task with procedurally generated tracks. The agent sees a top-down view and must visit as many tiles as possible. The full V+M+C model scored 906 (out of ~950 maximum), significantly outperforming the previous best methods. The Vision-only baseline (V+C, no memory) scored 632 — proving that memory and prediction are crucial for handling curves and planning ahead.
VizDoom Take Cover is a first-person survival game where the agent must dodge fireballs thrown by monsters. Here, the paper demonstrated dream training: the controller was trained entirely inside the world model's imagination, then transferred to the real environment, achieving over 1100 timesteps of survival — competitive with agents trained with direct environment access.
Why this matters: from dreaming to modern world models
World Models introduced a blueprint that the entire field followed. The core recipe — learn a latent representation, predict dynamics in that space, plan using imagined trajectories — appears in virtually every modern system.
The paper also introduced a profound conceptual shift: the agent does not directly perceive reality. It only sees what its world model shows it. This "Plato's Cave" perspective highlights both the power and the danger: a good world model enables efficient learning, but a flawed model can create systematic blind spots that the agent cannot detect.
The limitations were clear too. The MDN-RNN struggled with long horizons. The VAE discarded details that sometimes mattered. The CMA-ES controller did not scale to high-dimensional action spaces. But these were engineering limitations, not conceptual ones — and the follow-up work (PlaNet, Dreamer, MuZero) addressed each in turn.
Legacy: from dreaming agents to foundation world models
2018
World Models (this paper)
VAE + MDN-RNN + linear controller. First demonstration of training an agent entirely in its own dream. Solved CarRacing and VizDoom.
2019
PlaNet — Learning Latent Dynamics
Replaced the MDN-RNN with a Recurrent State-Space Model (RSSM), splitting the hidden state into deterministic and stochastic components. Enabled online planning via cross-entropy method in latent space.
2020
Dreamer v1 — Dream to Control
Combined RSSM with actor-critic learning inside imagined trajectories. Achieved human-level performance on 20 continuous control tasks from pixels.
2020
MuZero — planning without a simulator
Learned a latent dynamics model end-to-end with tree search. Mastered Chess, Go, Shogi, and Atari without ever being given the rules of the game.
2023
Dreamer v3 — mastering diverse domains
A single algorithm with fixed hyperparameters that achieves human-level or better across 150+ tasks spanning robotics, Atari, DMLab, and Minecraft — the culmination of the World Models vision.
The trajectory from World Models to Dreamer v3 is a story of engineering the same core idea: learn a latent model of the world, imagine rollouts, and optimize behavior in imagination. Each successor addressed a specific limitation — better state models, better policy learning, better scaling — but the conceptual foundation laid in 2018 remains the blueprint.
CitationHa, Schmidhuber. Recurrent World Models Facilitate Policy Evolution. NeurIPS, 2018.
Terms in this paper
- World Modelنموذج العالَم
- Variational Autoencoder (VAE)المرمّز التلقائي المتغير الاحتمالي
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Latent Spaceالفضاء الكامن
- Latent Variableالمتغير الكامن
- Reconstruction Errorخطأ إعادة البناء
- Rolloutالتوليد التجريبي
- Model-Based RLالتعلم بالتعزيز المعتمد على بناء بيئة
- Intrinsic Motivationالدافعية الجوهرية
- Action Spaceفضاء الأفعال
- Hidden Stateالحالة المخفية
- Activation Functionدالة التنشيط
- Mixture Density Networkشبكة كثافة المزيج
- Policyالسياسة
- Rewardالمكافأة