Skip to content
AI360Xpert
Beta

Planning and Learning

Learning improves decisions through direct trial-and-error in the real world, while planning mentally simulates prospective scenarios using an internal model. Combining both enables rapid learning without risking costly real-world mistakes.

Planning and learning unite around a shared value function, combining real experience and simulated rollouts to accelerate policy optimization.
Planning and learning unite around a shared value function, combining real experience and simulated rollouts to accelerate policy optimization.

Why Does This Exist?

In classic reinforcement learning, algorithms are often split into two seemingly disconnected paradigms:

  1. Model-free learning: The agent takes physical actions in the real environment, observes rewards, and updates its value estimates through direct trial-and-error (e.g., Q-learning or SARSA).
  2. Model-based planning: The agent possesses an explicit model of transition dynamics and rewards, computing optimal policies through forward search or state-space sweeps (e.g., Dynamic Programming or Monte Carlo Tree Search).

In physical domains—such as robotics, autonomous driving, or industrial process automation—relying purely on real-world trial-and-error is prohibitively expensive, agonizingly slow, and physically dangerous. An autonomous car cannot afford to crash into barriers hundreds of times just to learn that collisions carry a negative reward.

The unified framework of planning and learning solves this fundamental bottleneck. It reveals that learning and planning are not mutually exclusive alternatives; they are complementary processes operating on the exact same core currency: Bellman value updates. By integrating both into a unified closed loop, real experience is used both to update values immediately and to learn an internal predictive model of the environment. That model then generates synthetic transitions to update policies hundreds of times faster than real time permits.

Think of It Like This

The Grandmaster's Mental Simulation

Imagine a world-class chess grandmaster competing in a high-stakes championship tournament.

When it is the grandmaster's turn:

  1. Real Experience: She reaches out, picks up a rook, and places it on square e4. She presses the clock. That single physical move is an interaction with the external environment. She observes her opponent's real reaction and face-to-face clock pressure.
  2. Mental Planning: As her opponent deliberates, the grandmaster does not sit with an empty mind waiting for the next turn. She closes her eyes and simulates ten moves ahead in her imagination: "If he moves the bishop to f5, I strike with knight to d5; if he pushes the c-pawn, I castle queenside."

None of those imagined board states are physically placed on the wooden table, yet every mental rehearsal updates her internal evaluation of which board formations are advantageous. When her opponent finally moves, she already knows the winning response because she has already experienced it dozens of times in her head.

Where the analogy stops: Human grandmasters simulate complex, deep tree structures with tactical pruning and heuristic pattern recognition. In computational reinforcement learning, planning often takes the form of randomized state-action background replays drawn from a learned lookup model or parametric neural network. Furthermore, if the grandmaster's mental model misremembers the rules of piece movement, her imaginary simulations will teach her disastrous blunders.

How It Actually Works

The Shared Foundation: Bellman Updates

At the heart of the unified perspective lies a profound insight formulated by Richard Sutton: planning and learning use the exact same mathematical updates to modify value functions and policies.

Both paradigms estimate state values V(s)V(s) or action values Q(s,a)Q(s, a) by applying Bellman backups. For instance, in Q-learning:

Q(S,A)←Q(S,A)+α[R+γmax⁡a′Q(S′,a′)−Q(S,A)]Q(S, A) \leftarrow Q(S, A) + \alpha \left[ R + \gamma \max_{a'} Q(S', a') - Q(S, A) \right]

The only difference between learning and planning is the origin of the transition quadruple (S,A,R,S′)(S, A, R, S'):

  • Learning (Model-Free RL): The transition (St,At,Rt+1,St+1)(S_t, A_t, R_{t+1}, S_{t+1}) is generated by the real external environment.
  • Planning (Model-Based RL): The transition (S,A,R,S′)(S, A, R, S') is generated by an internal environment model M\mathcal{M}.

Internal Models: Sample Models vs. Distribution Models

An internal model M\mathcal{M} is any computational process that mimics environmental dynamics. Given a state SS and action AA, a model predicts possible future outcomes:

  1. Distribution Models: A distribution model produces the complete probability distribution over all possible next states and expected rewards: Mdist={P(s′∣s,a),  R(s,a,s′)}\mathcal{M}_{\text{dist}} = \{P(s' \mid s, a), \; R(s, a, s')\} Planning method: Exhaustive expected Bellman backups, as performed in classic Dynamic Programming. While mathematically rigorous, distribution models require massive memory and cannot be sampled cheaply in high-dimensional or continuous state spaces.

  2. Sample Models: A sample model acts as an internal simulator. Given current state SS and chosen action AA, it draws a single concrete sample according to underlying transition probabilities: (R,S′)∼Msample(S,A)(R, S') \sim \mathcal{M}_{\text{sample}}(S, A) Planning method: Sample backups and simulated rollouts, identical in format to real experience. Sample models can be queried cheaply and parallelized effortlessly on modern hardware.

The Unified Feedback Architecture (Dyna)

The synthesis of planning, acting, and model learning forms the canonical Dyna architecture:

                       +-------------------+                       |    Environment    |                       +-------------------+                         ^               |                 Action  |               | Real Experience                   A_t   |               | (S_t, A_t, R_{t+1}, S_{t+1})                         |               v                   +-----------+   +-----------+                   |  Policy / |<--|   Model   |                   |   Value   |   |  Learning |                   +-----------+   +-----------+                         ^               |                         |               | Internal Model M = (P_hat, R_hat)                         |               v                         |      +-----------------+                         +------|    Planning     |                     Simulated  | (Simulated Exp) |                      Backups   +-----------------+

On every time step:

  1. Act: The agent selects action At∼π(⋅∣St)A_t \sim \pi(\cdot \mid S_t) based on its current value function and executes it in the real environment.
  2. Direct RL: The real transition (St,At,Rt+1,St+1)(S_t, A_t, R_{t+1}, S_{t+1}) updates the value function directly via a standard TD backup.
  3. Model Learning: The real transition updates the internal model M=(P^,R^)\mathcal{M} = (\hat{P}, \hat{R}) to better approximate real-world physics.
  4. Planning: The agent pauses its physical actions and executes NN simulated planning steps:
    • Sample previously visited state-action pair (Sk,Ak)(S_k, A_k).
    • Query model: (Rk,Sk′)←M(Sk,Ak)(R_k, S'_k) \leftarrow \mathcal{M}(S_k, A_k).
    • Apply standard Bellman backup to Q(Sk,Ak)Q(S_k, A_k) using the simulated outcome.

Worked numerical comparison of convergence speed

Consider a linear corridor task with three states {S0,S1,S2}\{S_0, S_1, S_2\}, where S2S_2 is a rewarding terminal goal: S0→ForwardS1→ForwardS2(Goal Reward R=10.0)S_0 \xrightarrow{\text{Forward}} S_1 \xrightarrow{\text{Forward}} S_2 \quad (\text{Goal Reward } R = 10.0)

Hyperparameters:

  • Initial action values: Q(S0)=0.0,  Q(S1)=0.0,  Q(S2)=0.0Q(S_0) = 0.0, \; Q(S_1) = 0.0, \; Q(S_2) = 0.0.
  • Step size α=0.5\alpha = 0.5, discount factor γ=0.9\gamma = 0.9.

Case 1: Pure Direct RL (No Planning, N=0N=0)

  • Step 0: Agent executes (S0,Forward)→S1(S_0, \text{Forward}) \to S_1 with reward R1=0R_1 = 0. Q(S0)←Q(S0)+0.5[0+0.9Q(S1)−Q(S0)]=0.0+0.5[0+0−0]=0.0Q(S_0) \leftarrow Q(S_0) + 0.5 \left[ 0 + 0.9 Q(S_1) - Q(S_0) \right] = 0.0 + 0.5[0 + 0 - 0] = \mathbf{0.0}
  • Step 1: Agent executes (S1,Forward)→S2(S_1, \text{Forward}) \to S_2 with reward R2=10.0R_2 = 10.0. Q(S1)←Q(S1)+0.5[10.0+0.9(0)−Q(S1)]=0.0+0.5[10.0−0]=5.0Q(S_1) \leftarrow Q(S_1) + 0.5 \left[ 10.0 + 0.9(0) - Q(S_1) \right] = 0.0 + 0.5[10.0 - 0] = \mathbf{5.0}
  • Result at episode end: Q(S1)=5.0Q(S_1) = 5.0, but Q(S0)=0.0Q(S_0) = \mathbf{0.0}. The start state learned nothing about the goal during this entire episode. Information requires a second physical episode to propagate back to S0S_0.

Case 2: Integrated Planning & Learning (N=2N=2 Planning Steps)

  • Step 0: Real transition S0→S1,R=0S_0 \to S_1, R=0. Direct RL updates Q(S0)=0.0Q(S_0) = 0.0. Model stores M(S0)=(S1,0)\mathcal{M}(S_0) = (S_1, 0).
  • Step 1: Real transition S1→S2,R=10.0S_1 \to S_2, R=10.0. Direct RL updates Q(S1)=5.0Q(S_1) = \mathbf{5.0}. Model stores M(S1)=(S2,10.0)\mathcal{M}(S_1) = (S_2, 10.0).
  • Immediate Planning Phase (mental simulation):
    • Planning Step 1: Query model for state S1S_1: M(S1)=(S2,10.0)\mathcal{M}(S_1) = (S_2, 10.0). Q(S1)←5.0+0.5[10.0+0−5.0]=7.5Q(S_1) \leftarrow 5.0 + 0.5 [10.0 + 0 - 5.0] = \mathbf{7.5}
    • Planning Step 2: Query model for state S0S_0: M(S0)=(S1,0.0)\mathcal{M}(S_0) = (S_1, 0.0). Q(S0)←0.0+0.5[0.0+0.9Q(S1)−0.0]=0.5[0.9×7.5]=3.375Q(S_0) \leftarrow 0.0 + 0.5 \left[ 0.0 + 0.9 Q(S_1) - 0.0 \right] = 0.5 [0.9 \times 7.5] = \mathbf{3.375}

In the exact same physical episode, the integrated planning agent propagated goal information all the way back to the initial start state S0S_0 (Q(S0)=3.375Q(S_0) = 3.375 vs. 0.00.0), saving real-world exploration time.

Code

import randomfrom typing import Dict, List, Tuple

def step_environment(state: int, action: str) -> Tuple[int, float]:    """Simulate a 5-state grid corridor (0: start -> 4: terminal goal)."""    if action == "right":        next_state = min(state + 1, 4)    else:        next_state = max(state - 1, 0)    reward = 10.0 if next_state == 4 else 0.0    return next_state, reward

def evaluate_planning_vs_learning(    num_episodes: int = 5,    planning_steps: int = 5,    alpha: float = 0.5,    gamma: float = 0.9,    seed: int = 42,) -> Tuple[float, float]:    """Compare pure Direct Q-learning against Dyna-style Integrated Planning & Learning.
    Args:        num_episodes: Number of training episodes.        planning_steps: Number of simulated planning updates per real step.        alpha: Learning rate step size.        gamma: Discount factor.        seed: Random seed for deterministic reproducibility.
    Returns:        Tuple of (direct_rl_start_q, integrated_planning_start_q).    """    states = [0, 1, 2, 3, 4]    actions = ["left", "right"]
    # -------------------------------------------------------------    # 1. Pure Direct RL (Model-Free Q-learning, N = 0 planning steps)    # -------------------------------------------------------------    random.seed(seed)    q_direct: Dict[Tuple[int, str], float] = {        (s, a): 0.0 for s in states for a in actions    }
    for ep in range(num_episodes):        s = 0        while s != 4:            a = "right" if random.random() < 0.8 else "left"            next_s, r = step_environment(s, a)            best_next_q = (                max(q_direct[(next_s, act)] for act in actions)                if next_s != 4                else 0.0            )            td_target = r + gamma * best_next_q            q_direct[(s, a)] += alpha * (td_target - q_direct[(s, a)])            s = next_s
    # -------------------------------------------------------------    # 2. Integrated Planning and Learning (Dyna-Q with N planning steps)    # -------------------------------------------------------------    random.seed(seed)    q_plan: Dict[Tuple[int, str], float] = {        (s, a): 0.0 for s in states for a in actions    }    model: Dict[Tuple[int, str], Tuple[int, float]] = {}
    for ep in range(num_episodes):        s = 0        while s != 4:            a = "right" if random.random() < 0.8 else "left"            next_s, r = step_environment(s, a)
            # A. Direct RL update from real experience            best_next_q = (                max(q_plan[(next_s, act)] for act in actions)                if next_s != 4                else 0.0            )            q_plan[(s, a)] += alpha * (r + gamma * best_next_q - q_plan[(s, a)])
            # B. Model learning: store observed deterministic transition            model[(s, a)] = (next_s, r)
            # C. Planning: perform N simulated updates from internal model            for _ in range(planning_steps):                sim_s, sim_a = random.choice(list(model.keys()))                sim_next, sim_r = model[(sim_s, sim_a)]                best_sim_next_q = (                    max(q_plan[(sim_next, act)] for act in actions)                    if sim_next != 4                    else 0.0                )                q_plan[(sim_s, sim_a)] += alpha * (                    sim_r + gamma * best_sim_next_q - q_plan[(sim_s, sim_a)]                )
            s = next_s
    q_direct_start = q_direct[(0, "right")]    q_plan_start = q_plan[(0, "right")]    return q_direct_start, q_plan_start

if __name__ == "__main__":    direct_val, plan_val = evaluate_planning_vs_learning(        num_episodes=5, planning_steps=5    )
    print("=== Convergence Comparison at Start State Q(0, 'right') ===")    print(f"Direct RL (N=0) Value:        {direct_val:.4f}")    print(f"Integrated Planning (N=5) Value: {plan_val:.4f}")
    # Assertion confirming planning drastically accelerates value propagation    assert (        plan_val > direct_val    ), "Planning must accelerate credit assignment back to start state!"    print(        f"\nValidation passed: Planning accelerated value propagation by {plan_val / direct_val:.2f}x."    )
# Expected Output:# === Convergence Comparison at Start State Q(0, 'right') ===# Direct RL (N=0) Value:        1.8225# Integrated Planning (N=5) Value: 7.2896## Validation passed: Planning accelerated value propagation by 4.00x.

Watch Out For

Compounding Model Bias and Hallucinated Planning Updates

The central vulnerability of planning is that the value function converges to the optimal policy for the learned internal model, which may diverge sharply from the real environment.

When an internal model is trained on limited data, is overparameterized, or encounters non-stationary dynamics, it develops subtle prediction biases. Because planning iterates simulated updates repeatedly through this flawed model, initial small errors compound exponentially across imaginary rollout steps:

  • The agent uncovers "phantom rewards" or hallucinated shortcuts that do not exist in physical reality.
  • The policy overfits to model artifacts, resulting in catastrophic failure when tested back in the real world.

The Fix:

  1. Dyna-Q+ Exploration Bonuses: Reward the agent with an exploration bonus (κτ\kappa \sqrt{\tau}) for visiting state-action pairs that have not been tested in real life for τ\tau steps, actively keeping the internal model fresh.
  2. Short Planning Horizons: Limit simulated rollouts to short horizons (k∈[1,5]k \in [1, 5] steps) rather than full imaginary episodes, constraining accumulated transition error (e.g., Model-Based Policy Optimization / MBPO).
  3. Ensemble Models: Maintain an ensemble of models and plan only over transitions where ensemble disagreement is low (uncertainty-aware planning).

The Quick Version

  • Shared algorithmic currency: Both learning and planning update state-action values using the identical Bellman equation; they differ solely in whether transitions originate from the real environment or an internal simulator.
  • The Dyna feedback loop: Real experience serves a dual purpose: it updates the value function directly (Direct RL) and trains the internal environment model M=(P^,R^)\mathcal{M} = (\hat{P}, \hat{R}).
  • Sample vs. distribution models: Distribution models describe complete probability distributions for exhaustive DP sweeps; sample models draw concrete transitions, enabling fast, parallelized Monte Carlo planning.
  • Sample efficiency vs. model bias: Planning dramatically reduces the number of physical real-world interactions required to learn, but risks policy failure if the internal model develops compounding simulation errors.