Reinforcement Learning2023advanced12 min read

Mastering Diverse Domains Through World Models

إتقان مجالات متنوّعة عبر نماذج العالم

Hafner, D. · Pasukonis, J. · Ba, J. · Lillicrap, T. — arXiv / Nature 2025

The problem

has achieved superhuman performance in specific domains — Go, Dota 2, StarCraft — but each success required extensive expert tuning of hyperparameters, , and algorithmic modifications tailored to that particular domain. Moving an algorithm from Atari to robotics, or from continuous control to sparse- 3D worlds, typically demands a new round of expensive tuning. This fragility blocks the path toward general-purpose agents that learn across diverse environments without human intervention.

The contribution

DreamerV3 is a single model-based reinforcement learning algorithm with fixed hyperparameters that outperforms specialized methods across more than 150 tasks spanning continuous control, Atari, 3D navigation, and open-world survival games. It achieves this through a suite of robustness techniques — symlog predictions for scale-invariant learning, KL balancing with free bits for stable , unimix categoricals for , and percentile return for domain-agnostic actor updates. DreamerV3 is the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula.

The impact

DreamerV3 demonstrated that a single general-purpose algorithm with no per-domain tuning can compete with or exceed specialists across wildly different reinforcement learning benchmarks. It was published in Nature in 2025, validating the world-model approach as a viable path toward general agents. Its robustness techniques — particularly symlog predictions and percentile return normalization — have been adopted beyond , including in PPO variants. The Minecraft diamond achievement established a new milestone for open-ended exploration from sparse rewards.

Think of a chess grandmaster who never needs a physical board. She visualizes moves, counter-moves, and entire games inside her head — and uses that mental rehearsal to improve. Now imagine she can do the same thing for any game: chess, soccer strategy, flight simulation, survival games — all using the same imagination process, without relearning how to think.

DreamerV3 is that grandmaster. It builds a mental model of whatever world it encounters, then practices thousands of strategies in its imagination. The same thinking process — unchanged — works across over 150 completely different tasks.

The three phases: experience, imagine, improve

DreamerV3 operates in three phases that repeat in a loop. First, the agent interacts with the real environment and stores experiences in a — raw observations, actions, and rewards. Second, the world model learns to compress these experiences into a compact and predict future states and rewards. Third, an pair improves behavior by rehearsing entirely inside the world model's imagination, never touching the real environment again.

The key insight is that phase three — the imagination phase — is vastly cheaper than real interaction. The agent can simulate thousands of future trajectories in the time it would take to run a single real episode. This makes DreamerV3 extraordinarily sample-efficient: it extracts maximum learning from minimal real experience.

Open in Lab
Click through the three phases to see how DreamerV3 cycles between real experience, world model training, and imagination-based policy improvement.
The demo wakes as you arrive…

The world model: RSSM — memory meets uncertainty

The world model is built on the Recurrent State-Space Model (RSSM), which combines two complementary representations. A deterministic path uses a recurrent network to maintain a persistent memory of past events — think of it as a running summary of everything that has happened. A stochastic path captures the inherent uncertainty in what happens next — because the future is never fully predictable.

Together, these form a compact latent state at each timestep. The deterministic state hth_t is the GRU's ; the stochastic state ztz_t is a collection of 32 categorical distributions, each with 32 classes, sampled as one-hot vectors. This discrete representation was shown in DreamerV2 to be more expressive than Gaussian latents for modeling complex environments.

The model has four components working in concert: an compresses raw observations into the stochastic state; a dynamics predictor forecasts the next stochastic state from the deterministic state alone (enabling imagination without observations); a estimates the reward at each state; and a reconstructs observations to keep the latent space grounded in reality.

Open in Lab
Click each RSSM component to see how deterministic memory and stochastic uncertainty combine to form the world model's latent state.
The demo wakes as you arrive…

Symlog predictions: one transform to tame all scales

Different environments produce wildly different numbers. A robotics task might give rewards between -1 and 1. A video game might give rewards in the thousands. Pixel values range from 0 to 255. If the neural network must handle all these scales with the same weights, training becomes unstable — large values dominate gradients and small values get ignored.

Previous solutions included reward clipping (which destroys information) or running statistics normalization (which is non-stationary and fragile). DreamerV3 introduces a simpler approach: the symlog transform.

symlog(x)=sign(x)ln⁡(∣x∣+1)symexp(x)=sign(x)(e∣x∣−1)\text{symlog}(x) = \text{sign}(x) \ln(|x| + 1) \qquad \text{symexp}(x) = \text{sign}(x) \left( e^{|x|} - 1 \right)
Symlog and its inverse symexp — compressing scales while preserving sign — Symlog squashes large values logarithmically while leaving small values (near zero) almost unchanged. This means the network can learn from rewards of 0.01 and 10,000 with the same architecture and loss function. Symexp reverses the transform to recover original-scale predictions. The network predicts in symlog space, then symexp converts back — no clipping, no running statistics, no per-domain tuning.
Open in Lab
Drag the input value slider to see how symlog compresses large values while preserving small ones. Compare with linear and log transforms.
The demo wakes as you arrive…

DreamerV3 applies symlog in three places: the encoder squashes input observations, the reward predictor outputs symlog-space predictions decoded via a twohot distribution, and the critic also predicts values in symlog space. The twohot encoding is a lightweight form of distributional reinforcement learning — instead of predicting a single value, the network predicts a categorical distribution over a discretized range of values, then computes the expected value. This provides richer gradient signals than point estimates.

KL balancing and free bits: keeping the world model honest

The world model must balance two competing pressures. The posterior (encoder) tries to pack as much information as possible into the latent state. The prior (dynamics predictor) tries to predict the next latent state from memory alone, without seeing the observation. The between them measures how much the encoder is smuggling in information that the dynamics can't predict.

If the KL penalty is too strong, the encoder gives up and ignores observations — the latent states become trivially predictable but useless. If it's too weak, the dynamics predictor doesn't learn because the encoder is doing all the work. Previous approaches required careful per-domain tuning of this balance.

DreamerV3 combines two techniques to make this robust. KL balancing uses an 80/20 split: 80% of the gradient trains the dynamics predictor to match the posterior (forward KL), while 20% trains the encoder to stay close to the prior (reverse KL). This asymmetry biases toward training the dynamics predictor, which is the component used during imagination. Free bits clamp the KL below 1 nat, allowing the encoder to encode at least 1 nat of information per timestep without penalty. This prevents the degenerate solution of an empty latent space.

LKL=0.8⋅max⁡ ⁣(KL[sg(qϕ) ∥ pϕ],  1)+0.2⋅max⁡ ⁣(KL[qϕ ∥ sg(pϕ)],  1)\mathcal{L}_{\text{KL}} = 0.8 \cdot \max\!\bigl(\text{KL}[\text{sg}(q_\phi) \,\|\, p_\phi],\; 1\bigr) + 0.2 \cdot \max\!\bigl(\text{KL}[q_\phi \,\|\, \text{sg}(p_\phi)],\; 1\bigr)
Combined KL loss — KL balancing with free bits (1 nat minimum) — sg\text{sg} is the Stop Gradient operator. The first term (80%) trains the prior pϕp_\phi to match a frozen posterior qϕq_\phi. The second term (20%) trains the posterior to stay near a frozen prior. The max⁡(⋅,1)\max(\cdot, 1) clamp is the free-bits mechanism — below 1 nat, no gradient flows, giving the encoder breathing room to encode useful information.

Exploration safety nets: unimix and percentile normalization

Two more robustness techniques complete DreamerV3's domain-agnostic design.

Unimix categoricals. In the RSSM and the actor network, categorical distributions are mixed with 1% uniform noise: pmix=0.99⋅pnet+0.01⋅Uniformp_{\text{mix}} = 0.99 \cdot p_{\text{net}} + 0.01 \cdot \text{Uniform}. This prevents any action or latent class from having zero probability. Even a well-trained policy always retains a 1% chance of trying something unexpected. Think of it as a fire exit that is always unlocked — the agent can never get permanently stuck in a behavioral dead-end.

Percentile return normalization. Traditional return normalization divides by the running standard deviation, but when rewards are sparse and the standard deviation is near zero, this amplifies noise catastrophically. DreamerV3 instead tracks the 5th and 95th percentiles of returns using exponential moving averages, then scales returns by the gap between them (only when the gap exceeds 1). This makes the actor insensitive to the absolute scale of rewards — a reward of 0.01 in one domain and 10,000 in another produce comparable gradient magnitudes.

Rnorm=R−Per5(R)max⁡ ⁣(Per95(R)−Per5(R),  1)R_{\text{norm}} = \frac{R - \text{Per}_5(R)} {\max\!\bigl(\text{Per}_{95}(R) - \text{Per}_5(R),\; 1\bigr)}
Percentile return normalization — scale-free across all domains — Returns are shifted by the 5th percentile and divided by the range between the 5th and 95th percentiles. The max⁡(⋅,1)\max(\cdot, 1) prevents noise amplification when rewards are sparse. Percentiles are tracked as exponential moving averages, making this statistic stable and adaptive.

Learning in imagination: the actor-critic loop

Once the world model is trained, DreamerV3 improves its policy entirely inside the model's imagination. Starting from states sampled from the replay buffer, the agent rolls out imagined trajectories using the dynamics predictor — no real environment interaction is needed. The actor (policy network) selects actions, and the critic (value network) estimates the expected cumulative reward from each state.

The actor maximizes the critic's value estimates through straight-through gradients: the backpropagates through the entire imagined . An regularizer with a fixed scale encourages exploration, and its effectiveness is guaranteed by the percentile normalization — the actor sees similarly scaled returns regardless of the domain.

The critic is trained with symlog twohot predictions and regularized toward an exponential moving average (EMA) of its own past outputs, smoothing training dynamics. A slow provides stable bootstrap targets, preventing the from chasing its own rapidly changing estimates.

Open in Lab
Watch the agent unroll imagined trajectories inside the world model. The actor chooses actions, the dynamics predictor steps forward, and the critic scores each state.
The demo wakes as you arrive…

Results: one algorithm, 150+ tasks, zero tuning

DreamerV3 was evaluated across seven benchmark suites with a single set of hyperparameters: Proprio Control (continuous), Visual Control (pixels), Atari 200M, BSuite, Crafter, DMLab, and Minecraft. In every suite, it matched or outperformed algorithms that were specifically tuned for that domain.

The most dramatic result was in Minecraft. Collecting a diamond requires navigating an open 3D world from pixel observations, crafting a technology tree of 17 items in the correct order, and doing so with extremely sparse rewards (the only reward signal is obtaining the diamond). Previous attempts either used human demonstrations, curriculum learning, or massive compute (VPT used 720 V100 GPUs for 9 days with contractor data). DreamerV3 collected diamonds end-to-end, from scratch, using a single V100 GPU — the first algorithm to do so.

The authors also demonstrated favorable scaling properties: larger models directly translated to higher data efficiency and final performance, with no instability or diminishing returns within the tested range.

Open in Lab
Compare DreamerV3's performance across benchmark suites. A single fixed configuration competes with domain-specific baselines everywhere.
The demo wakes as you arrive…

Architecture details: the full toolkit

Beyond the robustness techniques, DreamerV3 introduces several architectural refinements. The GRU is replaced by a Block GRU that applies the gating mechanism in parallel across blocks rather than element-wise, improving computational efficiency. All layers use LayerNorm (specifically RMSNorm in later versions) followed by SiLU activations, replacing the earlier use of ELU. The optimizer uses adaptive gradient clipping to stabilize training, and the replay buffer includes an online queue for recent experiences alongside the larger uniform-sampling buffer.

These may seem like minor engineering choices, but collectively they enable DreamerV3 to train stably with large networks. The robustness of training with LayerNorm and SiLU was critical for scaling — earlier versions of Dreamer could not use models as large without diverging.

Pseudocode: DreamerV3 training looppython

Simplified to show the idea — not the real implementation.

# Phase 1: Collect experience
obs = env.reset()
for step in range(N):
    action = actor(world_model.encode(obs))  # act in real world
    next_obs, reward, done = env.step(action)
    replay_buffer.add(obs, action, reward, done)
    obs = next_obs

# Phase 2: Train world model
batch = replay_buffer.sample(B, T)  # B sequences of length T
for each timestep t in batch:
    h_t = GRU(h_{t-1}, z_{t-1}, a_{t-1})     # deterministic state
    z_t_prior = dynamics_predictor(h_t)         # predict without obs
    z_t_post  = encoder(h_t, obs_t)             # encode with obs
    loss += reconstruction(decoder(h_t, z_t_post), obs_t)
    loss += reward_loss(reward_pred(h_t, z_t_post), reward_t)
    loss += kl_balanced(z_t_post, z_t_prior)    # 80/20 + free bits

# Phase 3: Improve policy in imagination
h, z = sample_initial_state(replay_buffer)
for horizon in range(H):                        # imagine H steps
    action = actor(h, z)                        # actor picks action
    h = GRU(h, z, action)                       # step dynamics
    z = dynamics_predictor(h)                    # predict next state
    value = critic(h, z)                         # score state

actor_loss  = -sum(normalized_returns)           # maximize returns critic_loss = twohot_symlog_loss(value, target)  # predict returns

The Dreamer lineage: from PlaNet to Nature

  1. 2019

    PlaNet (Hafner et al.)

    Introduced the RSSM architecture for learning world models with both deterministic and stochastic components. Planned actions via online optimization (CEM) in latent space.

  2. 2020

    DreamerV1 (Hafner et al.)

    Replaced online planning with a learned actor-critic, enabling policy improvement through backpropagation through imagined trajectories. First Dreamer to master continuous control tasks from pixels.

  3. 2021

    DreamerV2 (Hafner et al.)

    Switched to categorical latent states (32×32 one-hot vectors) and introduced KL balancing. Achieved human-level performance on Atari with a single GPU.

  4. 2023

    DreamerV3 (this paper)

    Added symlog predictions, free bits, unimix categoricals, and percentile normalization. First fixed-hyperparameter algorithm to master 150+ tasks and collect Minecraft diamonds from scratch.

  5. 2025

    Published in Nature

    DreamerV3 was published in Nature, validating world-model-based reinforcement learning as a serious path toward general-purpose agents.

CitationHafner, Pasukonis, Ba, Lillicrap. Mastering Diverse Domains Through World Models. arXiv preprint arXiv:2301.04104 / Nature 2025, 2023.

Terms in this paper