Dreamer: World Models for RL
Instead of learning through costly trials in the real environment, Dreamer trains an actor-critic agent entirely inside a learned latent world model using imaginary rollouts.
Why Does This Exist?
Model-free reinforcement learning algorithms such as PPO and Soft Actor-Critic achieve impressive asymptotic mastery, but require millions of physical interactions with the environment. In physical robotics, industrial control, and complex virtual worlds, collecting millions of real transitions is prohibitively expensive, slow, and dangerous.
Conversely, traditional model-based RL approaches that plan in observation space (such as pixel-level model predictive control) buckle under high computational costs and compounding pixel blurriness over extended horizons.
Dreamer (Hafner et al., 2019–2023) fundamentally solves this trade-off by decoupling data collection from policy optimization through a learned latent world model:
- The world model continuously learns compact latent dynamics directly from historical sensory observations in an experience replay buffer.
- The actor-critic agent learns behaviors entirely by "dreaming"—rolling out simulated trajectories inside the latent space of the world model without ever rendering pixels.
- Because latent transitions are fast matrix multiplications on GPU hardware, the agent generates tens of thousands of imagined transitions per second, exploring complex behaviors with superhuman sample efficiency.
Think of It Like This
A commercial airline pilot in a full-motion flight simulator
Imagine training an airline pilot to execute hazardous crosswind landings during a storm.
-
Model-free trial-and-error: Putting the student in a real Boeing 777 and letting them crash into the runway dozens of times until they discover correct rudder deflections. The cost, risk, and hardware destruction make this completely impossible.
-
Observational replay: Forcing the student to watch thousands of hours of recorded cockpit videos. The pilot cannot test hypothetical choices ("what happens if I pull the yoke back right now?") because recorded video cannot respond interactively.
-
The Dreamer approach (Flight Simulator): Aerospace engineers build a high-fidelity digital physics engine (the World Model) using historical aerodynamic flight telemetry. The pilot sits in the simulator and logs thousands of simulated flight hours (the Imagined Rollouts).
Inside the simulation, the pilot tests radical maneuvers, explores near-stall conditions, and learns from simulated failures without risking an aircraft. An automated flight instructor (the Critic) grades every touchdown. When the pilot finally steps into the cockpit of a real airplane, their motor reactions are already polished.
Where the analogy stops: Flight simulator physics are hand-coded by human aerodynamicists from known Newtonian equations. In Dreamer, the agent must build the simulator entirely on its own from raw sensory streams (). If the learned world model has inaccurate physics, the policy internalizes dangerous hallucinations—requiring careful information bottleneck regularizers and bounded imagination horizons.
How It Actually Works
The Dreamer Architecture and Evolution
Dreamer operates through three concurrent, asynchronous learning stages:
┌─────────────────────────────────────────────────────────────┐│ 1. REAL ENVIRONMENT INTERACTION ││ Agent executes policy π_ψ(a_t | z_t) in real world ││ Saves transitions (x_t, a_t, r_t, γ_t) to Replay Buffer │└──────────────────────────────┬──────────────────────────────┘ │ Real data ▼┌─────────────────────────────────────────────────────────────┐│ 2. LEARN WORLD MODEL (RSSM) ││ Minimizes ELBO: Visual Recon + Reward Loss + KL Balance ││ Latent state z_t = (h_t, s_t) │└──────────────────────────────┬──────────────────────────────┘ │ Seed states z_0 ▼┌─────────────────────────────────────────────────────────────┐│ 3. LATENT IMAGINATION (ACTOR-CRITIC) ││ Roll out policy π_ψ for H = 15 steps in latent space ││ Compute multi-step λ-returns V^λ(z_τ) ││ Update Actor (analytic gradients) & Critic (two-hot) │└─────────────────────────────────────────────────────────────┘The Evolutionary Milestones: V1, V2, and V3
-
DreamerV1 (2019):
- Parameterized stochastic latents as continuous multivariate Gaussians .
- Introduced analytic gradient backpropagation: because the transition dynamics and reward predictor are fully differentiable neural networks, the Actor receives direct analytic gradients backpropagated backward through time across the latent trajectory.
- Mastered visual continuous control benchmarks (DeepMind Control Suite) directly from pixels.
-
DreamerV2 (2020):
- Categorical Latents: Replaced continuous Gaussian latents with an array of 32 discrete categorical variables, each with 32 classes ( discrete choices). Backpropagation through discrete samples is accomplished using the straight-through estimator: Discrete categorical latents prevented latent representation collapse, accurately modeled multimodal environment branches, and achieved human-level performance across the 55 Atari 200M benchmark games from raw pixels.
- KL Balancing: Addressed the issue of the prior dominating the posterior by splitting the KL loss into separate updates with distinct learning rates: Setting and trains the prior toward the posterior four times faster than the posterior is pulled toward the prior.
-
DreamerV3 (2023):
- Scale-Invariant World Modeling: Designed to run across wildly different domains (Atari, continuous control, DMLab, Minecraft) with fixed hyperparameters.
- Symlog Transformation: Compresses wide dynamic ranges of rewards and observations using a symmetric logarithmic mapping: Its inverse, , maps predictions back to linear scale.
- Two-Hot Categorical Returns: Rather than predicting scalar values with mean-squared error, the Critic predicts probabilities over 255 discrete symlog-spaced bins, preventing gradient explosion in environments with sparse massive rewards.
- DreamerV3 became the first algorithm to mine diamonds in Minecraft from scratch without human demonstration data.
Generalized -Returns for Latent Imagination
During latent imagination, the agent starts from an initial encoded latent state and rolls out the actor policy for a fixed horizon :
To balance the low variance of the Critic with the unbiased nature of long rollouts, Dreamer evaluates imagined trajectories using generalized -returns .
The -step bootstrapped return from state is:
The -return geometrically weights all intermediate -step horizon estimates:
Worked numerical example
Let us compute the 2-step -return and symlog scaling for an imagined trajectory segment.
Trajectory Setup:
- Horizon sub-window: , discount factor , decay weight .
- Imagined reward sequence: , , .
- Critic value estimates at latent states: , , .
Step 1: Compute -step returns for starting state
-
1-step bootstrapped return ():
-
2-step bootstrapped return ():
Step 2: Compute -return at
Weighting the 1-step and 2-step returns via :
Step 3: Symlog Transformation (DreamerV3)
The Critic targets this return using the symlog transform:
Evaluating the inverse to verify exact preservation:
Code
import mathfrom dataclasses import dataclassfrom typing import List, Tuple
def symlog(x: float) -> float: """DreamerV3 symlog transformation: compresses extreme return scales.""" sign = 1.0 if x > 0 else (-1.0 if x < 0 else 0.0) return sign * math.log(abs(x) + 1.0)
def symexp(y: float) -> float: """Inverse symlog: maps compressed representation back to original scale.""" sign = 1.0 if y > 0 else (-1.0 if y < 0 else 0.0) return sign * (math.exp(abs(y)) - 1.0)
@dataclassclass CategoricalLatent: """Represents a discrete categorical latent variable with straight-through estimator."""
one_hot: List[float] probabilities: List[float]
class DreamerImaginationEngine: """Demonstrates Dreamer's latent imagination and return calculation mechanisms:
1. Straight-through categorical sampling (DreamerV2/V3) 2. Multi-step lambda-return calculation (TD-lambda) 3. Scale-invariant symlog mapping (DreamerV3) """
def __init__(self, gamma: float = 0.9, lam: float = 0.8) -> None: self.gamma = gamma self.lam = lam
def sample_categorical_straight_through( self, logits: List[float] ) -> CategoricalLatent: """Samples a discrete category while preserving gradients via straight-through.""" exp_logits = [math.exp(val) for val in logits] denom = sum(exp_logits) probs = [val / denom for val in exp_logits]
# Discrete argmax one-hot selection best_idx = probs.index(max(probs)) one_hot = [1.0 if idx == best_idx else 0.0 for idx in range(len(probs))]
# In a differentiable framework: z_st = z_onehot + probs - probs.detach() return CategoricalLatent(one_hot=one_hot, probabilities=probs)
def compute_two_step_lambda_return( self, r1: float, r2: float, v2: float, v3: float, ) -> Tuple[float, float, float]: """Calculates 1-step, 2-step, and composite lambda-return for state z_1.""" # 1-step bootstrap g1 = r1 + self.gamma * v2
# 2-step bootstrap g2 = r1 + self.gamma * r2 + (self.gamma**2) * v3
# Lambda return blend v_lambda = (1.0 - self.lam) * g1 + self.lam * g2
return g1, g2, v_lambda
# Initialize engine with parameters matching worked numerical exampleengine = DreamerImaginationEngine(gamma=0.9, lam=0.8)
# 1. Straight-Through Categorical Latent Demonstrationtest_logits = [0.2, 1.8, 0.5, 0.1]latent = engine.sample_categorical_straight_through(test_logits)print(f"Sampled one-hot category: {latent.one_hot}")# -> Sampled one-hot category: [0.0, 1.0, 0.0, 0.0]
print( f"Softmax probabilities: {[round(p, 3) for p in latent.probabilities]}")# -> Softmax probabilities: [0.122, 0.603, 0.164, 0.11]
# 2. Lambda-Return Calculationr_1, r_2 = 1.0, 2.0v_2, v_3 = 2.5, 1.0
g_1, g_2, v_lambda_1 = engine.compute_two_step_lambda_return( r1=r_1, r2=r_2, v2=v_2, v3=v_3)
print(f"1-step return G_1^(1): {g_1:.2f}")# -> 1-step return G_1^(1): 3.25
print(f"2-step return G_1^(2): {g_2:.2f}")# -> 2-step return G_1^(2): 3.61
print(f"Lambda-return V^lambda(z_1): {v_lambda_1:.3f}")# -> Lambda-return V^lambda(z_1): 3.538
# 3. Symlog Transformationsymlog_target = symlog(v_lambda_1)print(f"Symlog return target: {symlog_target:.4f}")# -> Symlog return target: 1.5125
restored_target = symexp(symlog_target)print(f"Symexp restored target: {restored_target:.3f}")# -> Symexp restored target: 3.538
# Verification assertionsassert round(g_1, 2) == 3.25assert round(g_2, 2) == 3.61assert round(v_lambda_1, 3) == 3.538assert round(symlog_target, 4) == 1.5125assert round(restored_target, 3) == 3.538Watch Out For
Imagination Horizon Drift and Latent Value Hallucination
A frequent failure mode among practitioners implementing Dreamer is arbitrarily extending the imagination rollout horizon (e.g., or ) under the assumption that longer rollouts necessarily discover better long-term strategies.
In practice, learned world models inevitably harbor subtle errors. When unrolled over long horizons without real sensor corrections, small prediction inaccuracies compound exponentially. The actor policy quickly learns to "exploit the physics engine"—navigating into hallucinated states where the model falsely predicts infinite rewards or zero risk. The policy achieves sky-high scores inside its internal imagination, but crashes instantly when evaluated in the real environment. Conversely, setting cripples the model into myopic greedy behavior.
The Fix:
- Calibrated Horizon Window: Keep the imagination horizon fixed at , which empirically balances planning depth against compounding dynamics error.
- TD() Bootstrapping: Always anchor long-term imagination returns using the Critic's value estimate weighted by . The -return acts as a principled regularizer, decaying reliance on distant, noisy simulation steps.
The Quick Version
- Dreamer decouples reinforcement learning into two concurrent loops: learning an internal world model from replay data, and training an actor-critic agent entirely inside imagined latent rollouts.
- DreamerV2 replaced continuous Gaussian latents with arrays of categorical latents () using straight-through estimators and KL balancing, preventing representation collapse.
- DreamerV3 introduced scale-invariant symlog transformations and two-hot return distributions, enabling identical hyperparameters across continuous control, Atari, and Minecraft.
- Policy optimization runs across step latent rollouts evaluated via -returns, achieving 10,000+ simulated steps per second without decoding pixels.