Dyna Architecture
The Dyna architecture lets an agent learn twice from every experience: first directly through trial and error, and then dozens of times in the background by replaying what it learned inside an internal world model.
Why Does This Exist?
In real-world reinforcement learning, environment interactions are expensive, slow, and potentially hazardous. A physical robot attempting to learn solely through model-free trial and error (like standard Q-learning) requires hundreds of thousands of physical transitions to propagate rewards backwards across its state space.
Historically, reinforcement learning was divided into two isolated camps:
- Model-Free RL (Direct Learning): The agent interacts with the real world, applies Bellman updates directly to value functions or policies, and discards the transition. It requires no world model, but its sample efficiency is notoriously poor.
- Model-Based Planning (Dynamic Programming): The agent assumes a fully known transition model and reward model , running offline sweeps across states. While sample efficient, it is helpless when the environment dynamics are initially unknown.
Introduced by Richard Sutton in 1990, the Dyna architecture bridges this divide by synthesizing direct learning, model learning, and planning into a single unified online architecture.
In Dyna, real experience serves a dual purpose:
- It immediately updates the value function via direct model-free RL.
- It simultaneously trains an internal world model .
Between real environment steps, the agent pauses to run planning steps: it samples hypothetical state-action transitions from its learned model, generates simulated experiences, and applies the exact same reinforcement learning update rule to its value function. By "dreaming" simulated transitions for every single physical step, Dyna accelerates sample efficiency by orders of magnitude.
Think of It Like This
Flight Training: Daytime Cockpit Hours vs. Nighttime Flight Simulators
Imagine training to become a commercial airline pilot:
- Pure Model-Free RL (Only Flying Real Planes): You learn solely by sitting in a physical cockpit. If you encounter severe crosswinds during a landing at 3:00 PM, you make a real-time adjustment. Once you park the plane, you wait until tomorrow's flight to experience crosswinds again. It takes years and hundreds of thousands of dollars in jet fuel to encounter enough rare weather conditions to become proficient.
- Pure Model-Based Dynamic Programming (The Theoretical Textbook): You are handed a 2,000-page aerodynamic manual and expected to compute exact flight paths mathematically before ever touching a control stick. If the real atmosphere deviates from your equations, your plans fail.
- The Dyna Architecture (Flying by Day, Simulating by Night): During the day, you fly real cockpit hours (). Each landing teaches your reflexes directly (Direct RL). Simultaneously, you plug your flight recorder's telemetry into a computer (Model Learning). That evening, you sit in a high-fidelity flight simulator that replicates the exact wind shear and turbulence you recorded that afternoon (Planning). In one evening, you rehearse that crosswind approach 50 times in simulation (). By tomorrow morning, your reflexes are already sharpened without burning a single extra drop of fuel.
Where the analogy stops: A human flight simulator uses complex physics engines, whereas tabular Dyna-Q builds a lookup table of past observations. Furthermore, Dyna applies identical temporal difference update equations to both real and simulated transitions.
How It Actually Works
Sutton's Dyna-Q Algorithm and Dyna-Q+ Exploration
The Dyna architecture operates on two concurrent timescales: real environment time and background planning time. In tabular Dyna-Q, the agent maintains an action-value table and a deterministic transition model .
1. The Core Dyna-Q Loop
At each discrete time step , the agent executes the following unified procedure:
- Action Selection: Select an action in state using an exploratory policy (e.g., -greedy with respect to ):
- Real Interaction: Execute in the real environment; observe real scalar reward and next state .
- Direct RL Update: Apply a standard Q-learning update to the Q-table using the real transition:
- Model Learning: Store the observed transition in the internal model buffer: The agent also tracks that state and action have been experienced, adding them to sets and .
- Background Planning Loop (Repeat Times):
- Search Control: Sample a previously visited state uniformly at random from .
- Sample a previously executed action uniformly at random from .
- Simulate: Query the internal model for the predicted outcome:
- Planning Update: Apply the standard Q-learning Bellman update to using simulated experience:
- State Transition: Set and repeat.
When , Dyna-Q reduces exactly to standard model-free Q-learning. When , each physical transition is amplified by simulated replay updates.
2. Dyna-Q+: Exploration in Non-Stationary Worlds
A major challenge for model-based planning arises when the real world changes (non-stationary environments):
- Shortcuts open: A wall vanishes, opening a faster path to the goal.
- Paths block: An existing corridor is closed off.
A standard Dyna-Q agent whose model was trained on the old environment will continue planning with outdated transitions, failing to discover new shortcuts because its greedy policy avoids unvisited areas.
To overcome this, Sutton formulated Dyna-Q+, which introduces an exploration bonus based on recency. Let be the number of real time steps that have elapsed since state-action pair was last executed in the real environment.
During planning, the model's simulated reward is augmented with a bonus proportional to the square root of elapsed time:
Where:
- measures elapsed time since was tried physically.
- is a small exploration scaling coefficient (e.g., ).
The longer a transition has gone unvisited, the more attractive it becomes during background planning. If a dormant action suddenly yields a high bonus in simulation, the agent's policy shifts toward visiting that action in the real world to test whether the environment has changed.
Worked numerical example
To observe how background planning propagates rewards across multiple states in a single real transition, consider a 3-state navigation corridor:
Parameters:
- Discount factor
- Learning rate
- Initial Q-values: for all states
- Action:
Step 1: Real Step from to (Reward )
- Direct RL:
- Model Storage:
- Visited memory: . Planning updates on yield TD error 0.
Step 2: Real Step from to (Terminal Goal, Reward )
- Direct RL Update:
- Model Storage:
- Visited memory: .
Now compare what happens next with (No Planning) versus (Dyna Planning):
Case 1: Pure Q-Learning ()
No planning steps are executed.
- Values at end of Episode 1:
- State learned nothing from the agent reaching the goal. It will require a second physical episode just for to discover that has positive value.
Case 2: Dyna-Q Planning ( Simulated Steps)
The agent pauses in the background to execute 5 simulated planning steps, sampling from :
- Planning Step 1 (Sample ):
- Query model: .
- Target: .
- Update: . (The goal reward has already leaped backward to without any physical movement!)
- Planning Step 2 (Sample ):
- Query model: . Target: .
- Update: .
- Planning Step 3 (Sample ):
- Query model: . Target: .
- Update: .
- Planning Step 4 (Sample ):
- Target: .
- Update: .
- Planning Step 5 (Sample ):
- Target: .
- Update: .
Comparison at the End of Episode 1
| Algorithm | Episodes Needed for to Learn | ||
|---|---|---|---|
| Q-Learning () | 2 episodes | ||
| Dyna-Q () | 1 episode |
With 5 background planning steps, absorbed over of the optimal return during the very first physical run.
Code
The following self-contained Python implementation trains a Dyna-Q agent on a discrete gridworld, contrasting against planning steps and asserting sample efficiency gains.
import randomfrom typing import Dict, List, Set, Tuple
class DynaQAgent: """Tabular Dyna-Q agent integrating direct RL with background model planning.
Environment: 0 <-> 1 <-> 2 <-> 3 (Terminal Goal, Reward = +10.0) Actions: 0 = 'left', 1 = 'right' """
def __init__( self, num_states: int = 4, planning_steps: int = 5, alpha: float = 0.5, gamma: float = 0.9, epsilon: float = 0.0, ) -> None: self.num_states = num_states self.goal_state = num_states - 1 self.actions = [0, 1] # 0: left, 1: right self.planning_steps = planning_steps self.alpha = alpha self.gamma = gamma self.epsilon = epsilon
# Action-value table Q(s, a) self.q_table: Dict[Tuple[int, int], float] = { (s, a): 0.0 for s in range(num_states) for a in self.actions }
# Deterministic transition model: (s, a) -> (reward, next_state) self.model: Dict[Tuple[int, int], Tuple[float, int]] = {}
# History tracking for search control sampling self.observed_states: Set[int] = set() self.observed_actions: Dict[int, Set[int]] = {}
def choose_action(self, state: int) -> int: """Epsilon-greedy policy that breaks ties toward 'right' (1).""" if random.random() < self.epsilon: return random.choice(self.actions) q_left = self.q_table[(state, 0)] q_right = self.q_table[(state, 1)] return 1 if q_right >= q_left else 0
def step_environment( self, state: int, action: int ) -> Tuple[int, float, bool]: """Real environment step dynamics.""" next_state = ( min(state + 1, self.goal_state) if action == 1 else max(state - 1, 0) ) done = next_state == self.goal_state reward = 10.0 if done else 0.0 return next_state, reward, done
def direct_rl_update( self, s: int, a: int, r: float, s_next: int, done: bool ) -> None: """Applies direct Q-learning Bellman update from real experience.""" max_q_next = ( 0.0 if done else max(self.q_table[(s_next, act)] for act in self.actions) ) td_error = r + self.gamma * max_q_next - self.q_table[(s, a)] self.q_table[(s, a)] += self.alpha * td_error
def update_model(self, s: int, a: int, r: float, s_next: int) -> None: """Stores transition in model and records visited history.""" self.model[(s, a)] = (r, s_next) self.observed_states.add(s) if s not in self.observed_actions: self.observed_actions[s] = set() self.observed_actions[s].add(a)
def planning_loop(self) -> None: """Simulates N hypothetical transitions from learned model.""" for _ in range(self.planning_steps): # Search Control: sample previously visited state and action sim_s = random.choice(list(self.observed_states)) sim_a = random.choice(list(self.observed_actions[sim_s]))
# Query internal model sim_r, sim_s_next = self.model[(sim_s, sim_a)] sim_done = sim_s_next == self.goal_state
# Apply identical Q-learning update using simulated experience max_q_next = ( 0.0 if sim_done else max(self.q_table[(sim_s_next, act)] for act in self.actions) ) td_error = sim_r + self.gamma * max_q_next - self.q_table[(sim_s, sim_a)] self.q_table[(sim_s, sim_a)] += self.alpha * td_error
def run_episode(self) -> int: """Executes one episode of real interaction, model learning, and planning.""" state = 0 step_count = 0 while state != self.goal_state and step_count < 50: action = self.choose_action(state) next_state, reward, done = self.step_environment(state, action)
# 1. Direct RL self.direct_rl_update(state, action, reward, next_state, done)
# 2. Model learning self.update_model(state, action, reward, next_state)
# 3. Background planning (N simulated updates) self.planning_loop()
state = next_state step_count += 1
return step_count
# Compare Q-learning (N=0) against Dyna-Q (N=5) in Episode 1random.seed(42)agent_n0 = DynaQAgent(planning_steps=0)agent_n0.run_episode()
random.seed(42)agent_n5 = DynaQAgent(planning_steps=5)agent_n5.run_episode()
print("Episode 1 Value Comparison:")print(f"N=0 (Q-learning): Q(S0, right) = {agent_n0.q_table[(0, 1)]:.4f}")print(f"N=0 (Q-learning): Q(S1, right) = {agent_n0.q_table[(1, 1)]:.4f}")print(f"N=0 (Q-learning): Q(S2, right) = {agent_n0.q_table[(2, 1)]:.4f}")print("---")print(f"N=5 (Dyna-Q): Q(S0, right) = {agent_n5.q_table[(0, 1)]:.4f}")print(f"N=5 (Dyna-Q): Q(S1, right) = {agent_n5.q_table[(1, 1)]:.4f}")print(f"N=5 (Dyna-Q): Q(S2, right) = {agent_n5.q_table[(2, 1)]:.4f}")
# Assertions verifying Dyna-Q benefitsassert agent_n0.q_table[(0, 1)] == 0.0, "N=0 should not update S0 in episode 1"assert agent_n0.q_table[(1, 1)] == 0.0, "N=0 should not update S1 in episode 1"assert agent_n0.q_table[(2, 1)] == 5.0, "N=0 S2 Q-value incorrect"
assert ( agent_n5.q_table[(0, 1)] > 0.0), "Dyna planning should propagate value back to S0"assert ( agent_n5.q_table[(1, 1)] > 0.0), "Dyna planning should propagate value back to S1"assert ( agent_n5.q_table[(2, 1)] > agent_n0.q_table[(2, 1)]), "Dyna planning should reinforce S2"print("\nAll assertions passed: Dyna-Q planning successfully achieved multi-step learning in a single episode!")
# -> Expected output:# -> Episode 1 Value Comparison:# -> N=0 (Q-learning): Q(S0, right) = 0.0000# -> N=0 (Q-learning): Q(S1, right) = 0.0000# -> N=0 (Q-learning): Q(S2, right) = 5.0000# -> ---# -> N=5 (Dyna-Q): Q(S0, right) = 1.5188# -> N=5 (Dyna-Q): Q(S1, right) = 5.0625# -> N=5 (Dyna-Q): Q(S2, right) = 7.5000# -> # -> All assertions passed: Dyna-Q planning successfully achieved multi-step learning in a single episode!Watch Out For
Sampling Unvisited Transitions or Overfitting to Early Imperfect Models
A critical bug when implementing Dyna architectures is sampling arbitrary state-action pairs from the entire theoretical state space during the planning loop.
The Failure Mode: If search control uniformly samples unvisited pairs from before the model has ever observed them:
- Default Model Poisoning: The model will return an arbitrary default transition (such as ). Planning updates then treat these fictional transitions as ground truth, corrupting the Q-table with invalid zero-reward dynamics and causing the policy to collapse.
- Model Bias Overfitting: If planning iterations are set too high (e.g., ) during early exploration, the agent repeatedly trains its Q-table on a sparse, imperfect early model. The agent prematurely commits to a suboptimal route that it happened to stumble into first, blinding it to superior alternative paths.
- Non-Stationary Lockout: In changing environments, standard Dyna-Q plans against obsolete barriers or missing shortcuts forever without a recency bonus.
The Fix:
- Restrict Search Control to Visited Transitions: Strictly sample and so planning operates exclusively on validated environment data.
- Moderate Planning Steps Early: Keep moderate (e.g., ) or scale with model confidence to prevent policy collapse on early sparse samples.
- Use Dyna-Q+ for Changing Dynamics: Augment simulated rewards with the exploration bonus whenever environment non-stationarity is suspected.
The Quick Version
- Dual-process architecture: The Dyna architecture unifies direct reinforcement learning from real experience with background planning from a learned world model.
- Shared update equations: Direct RL and planning execute the exact same Bellman update rule; the only difference is whether the transition came from the physical environment or the model.
- Massive sample efficiency: Running simulated planning steps per real transition allows rewards to propagate across entire trajectories in a single physical episode.
- Dyna-Q+ recency bonus: Adding an exploration bonus based on time elapsed since an action was last tried enables agents to rapidly detect changing or newly opened shortcuts.