Skip to content
AI360Xpert
Beta

Dyna Architecture

The Dyna architecture lets an agent learn twice from every experience: first directly through trial and error, and then dozens of times in the background by replaying what it learned inside an internal world model.

The Dyna architecture integrates direct trial-and-error reinforcement learning with model learning and background planning to update value functions.
The Dyna architecture integrates direct trial-and-error reinforcement learning with model learning and background planning to update value functions.

Why Does This Exist?

In real-world reinforcement learning, environment interactions are expensive, slow, and potentially hazardous. A physical robot attempting to learn solely through model-free trial and error (like standard Q-learning) requires hundreds of thousands of physical transitions to propagate rewards backwards across its state space.

Historically, reinforcement learning was divided into two isolated camps:

  1. Model-Free RL (Direct Learning): The agent interacts with the real world, applies Bellman updates directly to value functions or policies, and discards the transition. It requires no world model, but its sample efficiency is notoriously poor.
  2. Model-Based Planning (Dynamic Programming): The agent assumes a fully known transition model P(s′∣s,a)P(s' \mid s, a) and reward model R(s,a)R(s, a), running offline sweeps across states. While sample efficient, it is helpless when the environment dynamics are initially unknown.

Introduced by Richard Sutton in 1990, the Dyna architecture bridges this divide by synthesizing direct learning, model learning, and planning into a single unified online architecture.

In Dyna, real experience serves a dual purpose:

  • It immediately updates the value function via direct model-free RL.
  • It simultaneously trains an internal world model Model(s,a)→(r,s′)\text{Model}(s, a) \to (r, s').

Between real environment steps, the agent pauses to run planning steps: it samples hypothetical state-action transitions from its learned model, generates simulated experiences, and applies the exact same reinforcement learning update rule to its value function. By "dreaming" NN simulated transitions for every single physical step, Dyna accelerates sample efficiency by orders of magnitude.

Think of It Like This

Flight Training: Daytime Cockpit Hours vs. Nighttime Flight Simulators

Imagine training to become a commercial airline pilot:

  • Pure Model-Free RL (Only Flying Real Planes): You learn solely by sitting in a physical cockpit. If you encounter severe crosswinds during a landing at 3:00 PM, you make a real-time adjustment. Once you park the plane, you wait until tomorrow's flight to experience crosswinds again. It takes years and hundreds of thousands of dollars in jet fuel to encounter enough rare weather conditions to become proficient.
  • Pure Model-Based Dynamic Programming (The Theoretical Textbook): You are handed a 2,000-page aerodynamic manual and expected to compute exact flight paths mathematically before ever touching a control stick. If the real atmosphere deviates from your equations, your plans fail.
  • The Dyna Architecture (Flying by Day, Simulating by Night): During the day, you fly real cockpit hours (S→A→R→S′S \to A \to R \to S'). Each landing teaches your reflexes directly (Direct RL). Simultaneously, you plug your flight recorder's telemetry into a computer (Model Learning). That evening, you sit in a high-fidelity flight simulator that replicates the exact wind shear and turbulence you recorded that afternoon (Planning). In one evening, you rehearse that crosswind approach 50 times in simulation (N=50N = 50). By tomorrow morning, your reflexes are already sharpened without burning a single extra drop of fuel.

Where the analogy stops: A human flight simulator uses complex physics engines, whereas tabular Dyna-Q builds a lookup table of past (s,a)→(r,s′)(s, a) \to (r, s') observations. Furthermore, Dyna applies identical temporal difference update equations to both real and simulated transitions.

How It Actually Works

Sutton's Dyna-Q Algorithm and Dyna-Q+ Exploration

The Dyna architecture operates on two concurrent timescales: real environment time and background planning time. In tabular Dyna-Q, the agent maintains an action-value table Q(s,a)Q(s, a) and a deterministic transition model Model(s,a)\text{Model}(s, a).

1. The Core Dyna-Q Loop

At each discrete time step tt, the agent executes the following unified procedure:

  1. Action Selection: Select an action AA in state SS using an exploratory policy (e.g., ε\varepsilon-greedy with respect to QQ): A←π(S)A \leftarrow \pi(S)
  2. Real Interaction: Execute AA in the real environment; observe real scalar reward RR and next state S′S'.
  3. Direct RL Update: Apply a standard Q-learning update to the Q-table using the real transition: Q(S,A)←Q(S,A)+α[R+γmax⁡a′Q(S′,a′)−Q(S,A)]Q(S, A) \leftarrow Q(S, A) + \alpha \left[ R + \gamma \max_{a'} Q(S', a') - Q(S, A) \right]
  4. Model Learning: Store the observed transition in the internal model buffer: Model(S,A)←(R,S′)\text{Model}(S, A) \leftarrow (R, S') The agent also tracks that state SS and action AA have been experienced, adding them to sets Svisited\mathcal{S}_{\text{visited}} and Avisited(S)\mathcal{A}_{\text{visited}}(S).
  5. Background Planning Loop (Repeat NN Times):
    • Search Control: Sample a previously visited state ss uniformly at random from Svisited\mathcal{S}_{\text{visited}}.
    • Sample a previously executed action aa uniformly at random from Avisited(s)\mathcal{A}_{\text{visited}}(s).
    • Simulate: Query the internal model for the predicted outcome: (r,s′)←Model(s,a)(r, s') \leftarrow \text{Model}(s, a)
    • Planning Update: Apply the standard Q-learning Bellman update to Q(s,a)Q(s, a) using simulated experience: Q(s,a)←Q(s,a)+α[r+γmax⁡a′Q(s′,a′)−Q(s,a)]Q(s, a) \leftarrow Q(s, a) + \alpha \left[ r + \gamma \max_{a'} Q(s', a') - Q(s, a) \right]
  6. State Transition: Set S←S′S \leftarrow S' and repeat.

When N=0N = 0, Dyna-Q reduces exactly to standard model-free Q-learning. When N>0N > 0, each physical transition is amplified by NN simulated replay updates.

2. Dyna-Q+: Exploration in Non-Stationary Worlds

A major challenge for model-based planning arises when the real world changes (non-stationary environments):

  • Shortcuts open: A wall vanishes, opening a faster path to the goal.
  • Paths block: An existing corridor is closed off.

A standard Dyna-Q agent whose model was trained on the old environment will continue planning with outdated transitions, failing to discover new shortcuts because its greedy policy avoids unvisited areas.

To overcome this, Sutton formulated Dyna-Q+, which introduces an exploration bonus based on recency. Let τ(s,a)\tau(s, a) be the number of real time steps that have elapsed since state-action pair (s,a)(s, a) was last executed in the real environment.

During planning, the model's simulated reward is augmented with a bonus proportional to the square root of elapsed time:

rbonus=r+κτ(s,a)r_{\text{bonus}} = r + \kappa \sqrt{\tau(s, a)}

Where:

  • τ(s,a)\tau(s, a) measures elapsed time since (s,a)(s, a) was tried physically.
  • κ>0\kappa > 0 is a small exploration scaling coefficient (e.g., 10−410^{-4}).

The longer a transition has gone unvisited, the more attractive it becomes during background planning. If a dormant action suddenly yields a high bonus in simulation, the agent's policy shifts toward visiting that action in the real world to test whether the environment has changed.

Worked numerical example

To observe how background planning propagates rewards across multiple states in a single real transition, consider a 3-state navigation corridor:

S0→rightS1→rightS2 (Terminal Goal, Reward +10.0)S_0 \xrightarrow{\text{right}} S_1 \xrightarrow{\text{right}} S_2 \text{ (Terminal Goal, Reward } +10.0\text{)}

Parameters:

  • Discount factor γ=0.9\gamma = 0.9
  • Learning rate α=0.5\alpha = 0.5
  • Initial Q-values: Q(s,a)=0.0Q(s, a) = 0.0 for all states
  • Action: A="right"A = \text{"right"}

Step 1: Real Step from S0S_0 to S1S_1 (Reward R1=0.0R_1 = 0.0)

  • Direct RL: Q(S0,right)←0.0+0.5[0.0+0.9max⁡Q(S1)−0.0]=0.0000Q(S_0, \text{right}) \leftarrow 0.0 + 0.5 \left[ 0.0 + 0.9 \max Q(S_1) - 0.0 \right] = 0.0000
  • Model Storage: Model(S0,right)←(0.0,S1)\text{Model}(S_0, \text{right}) \leftarrow (0.0, S_1)
  • Visited memory: Svisited={S0}\mathcal{S}_{\text{visited}} = \{S_0\}. Planning updates on (S0,right)(S_0, \text{right}) yield TD error 0.

Step 2: Real Step from S1S_1 to S2S_2 (Terminal Goal, Reward R2=10.0R_2 = 10.0)

  • Direct RL Update: Q(S1,right)←0.0+0.5[10.0+0.9(0.0)−0.0]=5.0000Q(S_1, \text{right}) \leftarrow 0.0 + 0.5 \left[ 10.0 + 0.9(0.0) - 0.0 \right] = \mathbf{5.0000}
  • Model Storage: Model(S1,right)←(10.0,S2)\text{Model}(S_1, \text{right}) \leftarrow (10.0, S_2)
  • Visited memory: Svisited={S0,S1}\mathcal{S}_{\text{visited}} = \{S_0, S_1\}.

Now compare what happens next with N=0N = 0 (No Planning) versus N=5N = 5 (Dyna Planning):

Case 1: Pure Q-Learning (N=0N = 0)

No planning steps are executed.

  • Values at end of Episode 1: Q(S1,right)=5.0000,Q(S0,right)=0.0000Q(S_1, \text{right}) = 5.0000, \quad Q(S_0, \text{right}) = \mathbf{0.0000}
  • State S0S_0 learned nothing from the agent reaching the goal. It will require a second physical episode just for S0S_0 to discover that S1S_1 has positive value.

Case 2: Dyna-Q Planning (N=5N = 5 Simulated Steps)

The agent pauses in the background to execute 5 simulated planning steps, sampling from Svisited={S0,S1}\mathcal{S}_{\text{visited}} = \{S_0, S_1\}:

  1. Planning Step 1 (Sample S0S_0):
    • Query model: (r=0.0,s′=S1)(r=0.0, s'=S_1).
    • Target: r+γmax⁡Q(S1)=0.0+0.9(5.0000)=4.5000r + \gamma \max Q(S_1) = 0.0 + 0.9(5.0000) = 4.5000.
    • Update: Q(S0,right)←0.0+0.5(4.5000−0.0)=2.2500Q(S_0, \text{right}) \leftarrow 0.0 + 0.5(4.5000 - 0.0) = \mathbf{2.2500}. (The goal reward has already leaped backward to S0S_0 without any physical movement!)
  2. Planning Step 2 (Sample S1S_1):
    • Query model: (r=10.0,s′=S2)(r=10.0, s'=S_2). Target: 10.0+0.0=10.010.0 + 0.0 = 10.0.
    • Update: Q(S1,right)←5.0+0.5(10.0−5.0)=7.5000Q(S_1, \text{right}) \leftarrow 5.0 + 0.5(10.0 - 5.0) = \mathbf{7.5000}.
  3. Planning Step 3 (Sample S0S_0):
    • Query model: (r=0.0,s′=S1)(r=0.0, s'=S_1). Target: 0.9(7.5000)=6.75000.9(7.5000) = 6.7500.
    • Update: Q(S0,right)←2.2500+0.5(6.7500−2.2500)=4.5000Q(S_0, \text{right}) \leftarrow 2.2500 + 0.5(6.7500 - 2.2500) = \mathbf{4.5000}.
  4. Planning Step 4 (Sample S1S_1):
    • Target: 10.010.0.
    • Update: Q(S1,right)←7.5000+0.5(10.0−7.5000)=8.7500Q(S_1, \text{right}) \leftarrow 7.5000 + 0.5(10.0 - 7.5000) = \mathbf{8.7500}.
  5. Planning Step 5 (Sample S0S_0):
    • Target: 0.9(8.7500)=7.87500.9(8.7500) = 7.8750.
    • Update: Q(S0,right)←4.5000+0.5(7.8750−4.5000)=6.1875Q(S_0, \text{right}) \leftarrow 4.5000 + 0.5(7.8750 - 4.5000) = \mathbf{6.1875}.

Comparison at the End of Episode 1

AlgorithmQ(S0,right)Q(S_0, \text{right})Q(S1,right)Q(S_1, \text{right})Episodes Needed for S0S_0 to Learn
Q-Learning (N=0N = 0)0.00000.00005.00005.00002 episodes
Dyna-Q (N=5N = 5)6.1875\mathbf{6.1875}8.7500\mathbf{8.7500}1 episode

With 5 background planning steps, S0S_0 absorbed over 60%60\% of the optimal return during the very first physical run.

Code

The following self-contained Python implementation trains a Dyna-Q agent on a discrete gridworld, contrasting N=0N=0 against N=5N=5 planning steps and asserting sample efficiency gains.

import randomfrom typing import Dict, List, Set, Tuple

class DynaQAgent:    """Tabular Dyna-Q agent integrating direct RL with background model planning.
    Environment: 0 <-> 1 <-> 2 <-> 3 (Terminal Goal, Reward = +10.0)    Actions: 0 = 'left', 1 = 'right'    """
    def __init__(        self,        num_states: int = 4,        planning_steps: int = 5,        alpha: float = 0.5,        gamma: float = 0.9,        epsilon: float = 0.0,    ) -> None:        self.num_states = num_states        self.goal_state = num_states - 1        self.actions = [0, 1]  # 0: left, 1: right        self.planning_steps = planning_steps        self.alpha = alpha        self.gamma = gamma        self.epsilon = epsilon
        # Action-value table Q(s, a)        self.q_table: Dict[Tuple[int, int], float] = {            (s, a): 0.0 for s in range(num_states) for a in self.actions        }
        # Deterministic transition model: (s, a) -> (reward, next_state)        self.model: Dict[Tuple[int, int], Tuple[float, int]] = {}
        # History tracking for search control sampling        self.observed_states: Set[int] = set()        self.observed_actions: Dict[int, Set[int]] = {}
    def choose_action(self, state: int) -> int:        """Epsilon-greedy policy that breaks ties toward 'right' (1)."""        if random.random() < self.epsilon:            return random.choice(self.actions)        q_left = self.q_table[(state, 0)]        q_right = self.q_table[(state, 1)]        return 1 if q_right >= q_left else 0
    def step_environment(        self, state: int, action: int    ) -> Tuple[int, float, bool]:        """Real environment step dynamics."""        next_state = (            min(state + 1, self.goal_state)            if action == 1            else max(state - 1, 0)        )        done = next_state == self.goal_state        reward = 10.0 if done else 0.0        return next_state, reward, done
    def direct_rl_update(        self, s: int, a: int, r: float, s_next: int, done: bool    ) -> None:        """Applies direct Q-learning Bellman update from real experience."""        max_q_next = (            0.0            if done            else max(self.q_table[(s_next, act)] for act in self.actions)        )        td_error = r + self.gamma * max_q_next - self.q_table[(s, a)]        self.q_table[(s, a)] += self.alpha * td_error
    def update_model(self, s: int, a: int, r: float, s_next: int) -> None:        """Stores transition in model and records visited history."""        self.model[(s, a)] = (r, s_next)        self.observed_states.add(s)        if s not in self.observed_actions:            self.observed_actions[s] = set()        self.observed_actions[s].add(a)
    def planning_loop(self) -> None:        """Simulates N hypothetical transitions from learned model."""        for _ in range(self.planning_steps):            # Search Control: sample previously visited state and action            sim_s = random.choice(list(self.observed_states))            sim_a = random.choice(list(self.observed_actions[sim_s]))
            # Query internal model            sim_r, sim_s_next = self.model[(sim_s, sim_a)]            sim_done = sim_s_next == self.goal_state
            # Apply identical Q-learning update using simulated experience            max_q_next = (                0.0                if sim_done                else max(self.q_table[(sim_s_next, act)] for act in self.actions)            )            td_error = sim_r + self.gamma * max_q_next - self.q_table[(sim_s, sim_a)]            self.q_table[(sim_s, sim_a)] += self.alpha * td_error
    def run_episode(self) -> int:        """Executes one episode of real interaction, model learning, and planning."""        state = 0        step_count = 0        while state != self.goal_state and step_count < 50:            action = self.choose_action(state)            next_state, reward, done = self.step_environment(state, action)
            # 1. Direct RL            self.direct_rl_update(state, action, reward, next_state, done)
            # 2. Model learning            self.update_model(state, action, reward, next_state)
            # 3. Background planning (N simulated updates)            self.planning_loop()
            state = next_state            step_count += 1
        return step_count

# Compare Q-learning (N=0) against Dyna-Q (N=5) in Episode 1random.seed(42)agent_n0 = DynaQAgent(planning_steps=0)agent_n0.run_episode()
random.seed(42)agent_n5 = DynaQAgent(planning_steps=5)agent_n5.run_episode()
print("Episode 1 Value Comparison:")print(f"N=0 (Q-learning): Q(S0, right) = {agent_n0.q_table[(0, 1)]:.4f}")print(f"N=0 (Q-learning): Q(S1, right) = {agent_n0.q_table[(1, 1)]:.4f}")print(f"N=0 (Q-learning): Q(S2, right) = {agent_n0.q_table[(2, 1)]:.4f}")print("---")print(f"N=5 (Dyna-Q):     Q(S0, right) = {agent_n5.q_table[(0, 1)]:.4f}")print(f"N=5 (Dyna-Q):     Q(S1, right) = {agent_n5.q_table[(1, 1)]:.4f}")print(f"N=5 (Dyna-Q):     Q(S2, right) = {agent_n5.q_table[(2, 1)]:.4f}")
# Assertions verifying Dyna-Q benefitsassert agent_n0.q_table[(0, 1)] == 0.0, "N=0 should not update S0 in episode 1"assert agent_n0.q_table[(1, 1)] == 0.0, "N=0 should not update S1 in episode 1"assert agent_n0.q_table[(2, 1)] == 5.0, "N=0 S2 Q-value incorrect"
assert (    agent_n5.q_table[(0, 1)] > 0.0), "Dyna planning should propagate value back to S0"assert (    agent_n5.q_table[(1, 1)] > 0.0), "Dyna planning should propagate value back to S1"assert (    agent_n5.q_table[(2, 1)] > agent_n0.q_table[(2, 1)]), "Dyna planning should reinforce S2"print("\nAll assertions passed: Dyna-Q planning successfully achieved multi-step learning in a single episode!")
# -> Expected output:# -> Episode 1 Value Comparison:# -> N=0 (Q-learning): Q(S0, right) = 0.0000# -> N=0 (Q-learning): Q(S1, right) = 0.0000# -> N=0 (Q-learning): Q(S2, right) = 5.0000# -> ---# -> N=5 (Dyna-Q):     Q(S0, right) = 1.5188# -> N=5 (Dyna-Q):     Q(S1, right) = 5.0625# -> N=5 (Dyna-Q):     Q(S2, right) = 7.5000# -> # -> All assertions passed: Dyna-Q planning successfully achieved multi-step learning in a single episode!

Watch Out For

Sampling Unvisited Transitions or Overfitting to Early Imperfect Models

A critical bug when implementing Dyna architectures is sampling arbitrary state-action pairs from the entire theoretical state space during the planning loop.

The Failure Mode: If search control uniformly samples unvisited pairs (s,a)(s, a) from S×A\mathcal{S} \times \mathcal{A} before the model has ever observed them:

  1. Default Model Poisoning: The model will return an arbitrary default transition (such as s′=0,r=0s' = 0, r = 0). Planning updates then treat these fictional transitions as ground truth, corrupting the Q-table with invalid zero-reward dynamics and causing the policy to collapse.
  2. Model Bias Overfitting: If planning iterations NN are set too high (e.g., N=100N=100) during early exploration, the agent repeatedly trains its Q-table on a sparse, imperfect early model. The agent prematurely commits to a suboptimal route that it happened to stumble into first, blinding it to superior alternative paths.
  3. Non-Stationary Lockout: In changing environments, standard Dyna-Q plans against obsolete barriers or missing shortcuts forever without a recency bonus.

The Fix:

  • Restrict Search Control to Visited Transitions: Strictly sample s∈Svisiteds \in \mathcal{S}_{\text{visited}} and a∈Avisited(s)a \in \mathcal{A}_{\text{visited}}(s) so planning operates exclusively on validated environment data.
  • Moderate Planning Steps Early: Keep NN moderate (e.g., N∈[5,20]N \in [5, 20]) or scale NN with model confidence to prevent policy collapse on early sparse samples.
  • Use Dyna-Q+ for Changing Dynamics: Augment simulated rewards with the exploration bonus r+κτ(s,a)r + \kappa \sqrt{\tau(s, a)} whenever environment non-stationarity is suspected.

The Quick Version

  • Dual-process architecture: The Dyna architecture unifies direct reinforcement learning from real experience with background planning from a learned world model.
  • Shared update equations: Direct RL and planning execute the exact same Bellman update rule; the only difference is whether the transition came from the physical environment or the model.
  • Massive sample efficiency: Running NN simulated planning steps per real transition allows rewards to propagate across entire trajectories in a single physical episode.
  • Dyna-Q+ recency bonus: Adding an exploration bonus r+κτ(s,a)r + \kappa \sqrt{\tau(s, a)} based on time elapsed since an action was last tried enables agents to rapidly detect changing or newly opened shortcuts.