Skip to content
AI360Xpert
Beta

Transition Dynamics

Transition dynamics define the probability distribution over where an agent will end up next after taking an action in a given state, modeling the uncertain physics of the environment.

Transition dynamics map state-action pairs to probability distributions over successor states, capturing the stochastic physics of the environment.
Transition dynamics map state-action pairs to probability distributions over successor states, capturing the stochastic physics of the environment.

Why Does This Exist?

In idealized computing environments, taking an action produces a deterministic outcome: if you write a byte to memory, it lands in that memory address with certainty. Real physical systems and competitive environments rarely behave this predictably. A robot issuing a drive command may hit slippery ice, an autonomous car steering right can encounter sudden wind gusts or traction loss, a network router forwarding a packet may suffer buffer drops, and a trading algorithm submitting a market order faces slippage and stochastic order fills.

Without a formal mechanism for transition dynamics, decision algorithms assume deterministic transitions—that taking action aa in state ss always leads to a single known next state s′s'. Under this naive assumption, an agent designs brittle plans that catastrophic noise dismantles within milliseconds. Furthermore, foundational reinforcement learning frameworks—such as Markov Decision Processes, Bellman Equations, and Value Functions—cannot formulate expected returns without knowing how likely each outcome is.

Transition dynamics provide the exact mathematical engine that models environmental physics. By mapping each state-action pair (s,a)(s, a) to a rigorous probability distribution over all possible successor states s′s', the dynamics function lets an agent compute expected returns, hedge against downside risks, and plan optimal behaviors in uncertain worlds.

Think of It Like This

Driving on an icy road

Imagine driving down a winter highway approaching an intersection. You turn your steering wheel 30 degrees to the right: that physical command is your action aa, and your current vehicle position and velocity make up your state ss.

On dry, clean asphalt, your steering input determines your next state with near 100% certainty: you cleanly enter the right lane. But on an icy road, traction physics decouples your intention from the physical outcome. Turning the wheel right produces a probability distribution:

  • 70% probability: The front tires bite through the frost and track into the right lane (s1′s'_1).
  • 15% probability: The front wheels lose grip entirely, and momentum skids the vehicle straight forward into the center of the intersection (s2′s'_2).
  • 15% probability: Road slush catches the left tires, pulling the car onto the shoulder (s3′s'_3).

You do not control which of these three outcomes occurs on any single attempt. The icy road, the tire rubber, and the laws of friction—the environment dynamics—decide the actual successor state. Your only lever of control is choosing which action distribution to gamble on.

Where the analogy stops: A human driver continuously corrects their steering in real time using high-frequency sensory feedback. In a discrete-time Markov process, the agent must commit to a single discrete action ata_t for the duration of the time step, and the transition kernel samples the successor state st+1s_{t+1} in one atomic transition before the agent receives its next observation.

How It Actually Works

The Conditional Dynamics Kernel and Joint Dynamics Function

In a Markov Decision Process, the environment's physics are formalized through conditional probability distributions. Let S\mathcal{S} denote the set of all valid states, A\mathcal{A} denote the set of all actions, and R⊂R\mathcal{R} \subset \mathbb{R} denote the set of possible rewards.

1. State Transition Probability Function

The standard state transition probability P(s′∣s,a)P(s' \mid s, a) (often written as T(s,a,s′)T(s, a, s') or Pss′a\mathcal{P}_{ss'}^a) defines the conditional probability that the environment transitions to state s′s' at time t+1t+1, given that the agent observed state ss and selected action aa at time tt:

P(s′∣s,a)≐Pr⁡(St+1=s′∣St=s,At=a)P(s' \mid s, a) \doteq \Pr(S_{t+1} = s' \mid S_t = s, A_t = a)

This formulation relies on the Markov Property: the transition distribution depends exclusively on the current state StS_t and immediate action AtA_t, making the full historical sequence of prior states and actions (S0,A0,S1,…,St−1,At−1)(S_0, A_0, S_1, \dots, S_{t-1}, A_{t-1}) conditionally independent:

Pr⁡(St+1=s′∣St=s,At=a,St−1=st−1,…,S0=s0)=Pr⁡(St+1=s′∣St=s,At=a)\Pr(S_{t+1} = s' \mid S_t = s, A_t = a, S_{t-1} = s_{t-1}, \dots, S_0 = s_0) = \Pr(S_{t+1} = s' \mid S_t = s, A_t = a)

2. The Four-Argument Joint Dynamics Function

In modern literature (notably Sutton & Barto), the environment is completely characterized by a four-argument joint dynamics function p:S×R×S×A→[0,1]p: \mathcal{S} \times \mathcal{R} \times \mathcal{S} \times \mathcal{A} \to [0, 1]:

p(s′,r∣s,a)≐Pr⁡(St+1=s′,Rt+1=r∣St=s,At=a)p(s', r \mid s, a) \doteq \Pr(S_{t+1} = s', R_{t+1} = r \mid S_t = s, A_t = a)

The state transition probability is recovered by marginalizing out the reward:

P(s′∣s,a)=∑r∈Rp(s′,r∣s,a)P(s' \mid s, a) = \sum_{r \in \mathcal{R}} p(s', r \mid s, a)

Similarly, the expected immediate reward received from taking action aa in state ss is the expectation across all possible rewards and successor states:

R(s,a)≐E[Rt+1∣St=s,At=a]=∑r∈Rr∑s′∈Sp(s′,r∣s,a)R(s, a) \doteq \mathbb{E}[R_{t+1} \mid S_t = s, A_t = a] = \sum_{r \in \mathcal{R}} r \sum_{s' \in \mathcal{S}} p(s', r \mid s, a)

3. Row Stochasticity and Transition Matrices

Because P(⋅∣s,a)P(\cdot \mid s, a) represents a valid probability distribution over mutual events, it must satisfy the axioms of probability:

∑s′∈SP(s′∣s,a)=1∀s∈S,  a∈A,andP(s′∣s,a)≥0∀s′∈S\sum_{s' \in \mathcal{S}} P(s' \mid s, a) = 1 \quad \forall s \in \mathcal{S}, \; a \in \mathcal{A}, \qquad \text{and} \qquad P(s' \mid s, a) \ge 0 \quad \forall s' \in \mathcal{S}

For a finite state space with ∣S∣=n|\mathcal{S}| = n, each action aa defines an n×nn \times n transition probability matrix Pa\mathbf{P}^a:

Pa=[P(s1∣s1,a)P(s2∣s1,a)⋯P(sn∣s1,a)P(s1∣s2,a)P(s2∣s2,a)⋯P(sn∣s2,a)⋮⋮⋱⋮P(s1∣sn,a)P(s2∣sn,a)⋯P(sn∣sn,a)]\mathbf{P}^a = \begin{bmatrix} P(s_1 \mid s_1, a) & P(s_2 \mid s_1, a) & \cdots & P(s_n \mid s_1, a) \\ P(s_1 \mid s_2, a) & P(s_2 \mid s_2, a) & \cdots & P(s_n \mid s_2, a) \\ \vdots & \vdots & \ddots & \vdots \\ P(s_1 \mid s_n, a) & P(s_2 \mid s_n, a) & \cdots & P(s_n \mid s_n, a) \end{bmatrix}

Every row of Pa\mathbf{P}^a sums to exactly 1, making Pa\mathbf{P}^a a right-stochastic matrix. If the agent follows a fixed policy π(a∣s)\pi(a \mid s), the induced transition matrix Pπ\mathbf{P}^\pi is:

[Pπ]i,j=∑a∈Aπ(a∣si)P(sj∣si,a)[\mathbf{P}^\pi]_{i, j} = \sum_{a \in \mathcal{A}} \pi(a \mid s_i) P(s_j \mid s_i, a)

Given an initial state probability distribution vector μ0∈R1×n\boldsymbol{\mu}_0 \in \mathbb{R}^{1 \times n}, the state distribution at step tt evolves through vector-matrix multiplication:

μt=μ0(Pπ)t\boldsymbol{\mu}_t = \boldsymbol{\mu}_0 (\mathbf{P}^\pi)^t

Worked numerical example

Consider a 3-state stochastic corridor environment:

  • s1s_1: Safe Road
  • s2s_2: Icy Patch
  • s3s_3: Goal Destination (absorbing terminal state)

The agent chooses action a=Forwarda = \text{Forward}. The transition dynamics matrix PForward\mathbf{P}^{\text{Forward}} is:

PForward=[0.100.700.200.200.300.500.000.001.00]\mathbf{P}^{\text{Forward}} = \begin{bmatrix} 0.10 & 0.70 & 0.20 \\ 0.20 & 0.30 & 0.50 \\ 0.00 & 0.00 & 1.00 \end{bmatrix}

Row checks confirm row-stochasticity:

  • Row 1: 0.10+0.70+0.20=1.000.10 + 0.70 + 0.20 = 1.00
  • Row 2: 0.20+0.30+0.50=1.000.20 + 0.30 + 0.50 = 1.00
  • Row 3: 0.00+0.00+1.00=1.000.00 + 0.00 + 1.00 = 1.00

Step 1: Forward distribution propagation over time

Assume the agent starts with absolute certainty in s1s_1, giving initial state vector:

μ0=[1.000.000.00]\boldsymbol{\mu}_0 = \begin{bmatrix} 1.00 & 0.00 & 0.00 \end{bmatrix}

After 1 step, the state distribution μ1=μ0PForward\boldsymbol{\mu}_1 = \boldsymbol{\mu}_0 \mathbf{P}^{\text{Forward}} is:

μ1,1=1.0(0.10)+0.0(0.20)+0.0(0.00)=0.10\mu_{1, 1} = 1.0(0.10) + 0.0(0.20) + 0.0(0.00) = 0.10 μ1,2=1.0(0.70)+0.0(0.30)+0.0(0.00)=0.70\mu_{1, 2} = 1.0(0.70) + 0.0(0.30) + 0.0(0.00) = 0.70 μ1,3=1.0(0.20)+0.0(0.50)+0.0(1.00)=0.20\mu_{1, 3} = 1.0(0.20) + 0.0(0.50) + 0.0(1.00) = 0.20 μ1=[0.100.700.20]\boldsymbol{\mu}_1 = \begin{bmatrix} 0.10 & 0.70 & 0.20 \end{bmatrix}

After 2 steps, the distribution μ2=μ1PForward\boldsymbol{\mu}_2 = \boldsymbol{\mu}_1 \mathbf{P}^{\text{Forward}} becomes:

μ2,1=0.10(0.10)+0.70(0.20)+0.20(0.00)=0.01+0.14+0.00=0.15\mu_{2, 1} = 0.10(0.10) + 0.70(0.20) + 0.20(0.00) = 0.01 + 0.14 + 0.00 = 0.15 μ2,2=0.10(0.70)+0.70(0.30)+0.20(0.00)=0.07+0.21+0.00=0.28\mu_{2, 2} = 0.10(0.70) + 0.70(0.30) + 0.20(0.00) = 0.07 + 0.21 + 0.00 = 0.28 μ2,3=0.10(0.20)+0.70(0.50)+0.20(1.00)=0.02+0.35+0.20=0.57\mu_{2, 3} = 0.10(0.20) + 0.70(0.50) + 0.20(1.00) = 0.02 + 0.35 + 0.20 = 0.57 μ2=[0.150.280.57]\boldsymbol{\mu}_2 = \begin{bmatrix} 0.15 & 0.28 & 0.57 \end{bmatrix}

Notice that the sum 0.15+0.28+0.57=1.000.15 + 0.28 + 0.57 = 1.00, preserving the total probability mass, with a 57% chance of already having reached the Goal.

Step 2: Weighting Bellman value expectations

Now let current state values be estimated as:

  • V(s1)=2.00V(s_1) = 2.00
  • V(s2)=5.00V(s_2) = 5.00
  • V(s3)=10.00V(s_3) = 10.00

Taking action a=Forwarda = \text{Forward} from s1s_1 yields an immediate step reward R(s1,Forward)=−1.00R(s_1, \text{Forward}) = -1.00, with discount factor γ=0.90\gamma = 0.90. The expected successor state value is computed by weighting each possible future state by its transition probability:

Es′∼P(⋅∣s1,Forward)[V(s′)]=∑s′∈{s1,s2,s3}P(s′∣s1,Forward)V(s′)\mathbb{E}_{s' \sim P(\cdot \mid s_1, \text{Forward})}[V(s')] = \sum_{s' \in \{s_1, s_2, s_3\}} P(s' \mid s_1, \text{Forward}) V(s') E[V(s′)]=0.10(2.00)+0.70(5.00)+0.20(10.00)=0.20+3.50+2.00=5.70\mathbb{E}[V(s')] = 0.10(2.00) + 0.70(5.00) + 0.20(10.00) = 0.20 + 3.50 + 2.00 = 5.70

The Bellman action-value is therefore:

Q(s1,Forward)=R(s1,Forward)+γE[V(s′)]=−1.00+0.90×5.70=−1.00+5.13=4.13Q(s_1, \text{Forward}) = R(s_1, \text{Forward}) + \gamma \mathbb{E}[V(s')] = -1.00 + 0.90 \times 5.70 = -1.00 + 5.13 = 4.13

Without the transition probabilities 0.10,0.70,0.200.10, 0.70, 0.20, computing this expectation would be impossible.

Code

import randomfrom typing import Dict, List, Tuple

class TransitionModel:    """Represents an MDP transition dynamics kernel P(s' | s, a) with invariant verification."""
    def __init__(        self,        states: List[str],        actions: List[str],        transitions: Dict[Tuple[str, str], Dict[str, float]],    ) -> None:        self.states = states        self.actions = actions        self.transitions = transitions        self._validate_stochasticity()
    def _validate_stochasticity(self, tolerance: float = 1e-6) -> None:        """Verifies row-stochasticity: for every (s, a), sum_s' P(s'|s,a) == 1.0."""        for s in self.states:            for a in self.actions:                dist = self.transitions.get((s, a), {})                total_prob = sum(dist.values())                if abs(total_prob - 1.0) > tolerance:                    raise ValueError(                        f"Row-stochasticity violated for ({s}, {a}): sum = {total_prob:.6f}"                    )                for s_prime, prob in dist.items():                    if prob < 0.0 or prob > 1.0:                        raise ValueError(                            f"Invalid probability {prob} for transition ({s}, {a}) -> {s_prime}"                        )
    def transition_matrix(self, action: str) -> List[List[float]]:        """Constructs the square transition matrix P^a where entry (i, j) is P(s_j | s_i, a)."""        return [            [self.transitions[(s, action)].get(s_next, 0.0) for s_next in self.states]            for s in self.states        ]
    def sample_next_state(self, state: str, action: str, rng: random.Random) -> str:        """Simulates environment physics by sampling s' ~ P(. | s, a)."""        dist = self.transitions[(state, action)]        candidates = list(dist.keys())        weights = list(dist.values())        return rng.choices(candidates, weights=weights, k=1)[0]

def propagate_distribution(    distribution: List[float], transition_matrix: List[List[float]]) -> List[float]:    """Computes mu_{t+1} = mu_t @ P for a state distribution vector mu."""    n_states = len(transition_matrix)    return [        sum(distribution[i] * transition_matrix[i][j] for i in range(n_states))        for j in range(n_states)    ]

if __name__ == "__main__":    # Define states and actions for a 3-state icy corridor    states = ["SafeRoad", "IcyPatch", "Goal"]    actions = ["Forward"]
    # Dynamics mapping: (state, action) -> {next_state: probability}    transitions = {        ("SafeRoad", "Forward"): {"SafeRoad": 0.10, "IcyPatch": 0.70, "Goal": 0.20},        ("IcyPatch", "Forward"): {"SafeRoad": 0.20, "IcyPatch": 0.30, "Goal": 0.50},        ("Goal", "Forward"): {"SafeRoad": 0.00, "IcyPatch": 0.00, "Goal": 1.00},    }
    env = TransitionModel(states, actions, transitions)    p_mat = env.transition_matrix("Forward")
    # 1. Exact theoretical distribution evolution: mu_{t+1} = mu_t @ P    mu_0 = [1.0, 0.0, 0.0]  # Agent begins in SafeRoad with certainty    print("Initial distribution (t=0):", [round(p, 4) for p in mu_0])
    mu_1 = propagate_distribution(mu_0, p_mat)    print("Step 1 distribution  (t=1):", [round(p, 4) for p in mu_1])
    mu_2 = propagate_distribution(mu_1, p_mat)    print("Step 2 distribution  (t=2):", [round(p, 4) for p in mu_2])
    # 2. Monte Carlo environment sampling: draw 10,000 transitions from SafeRoad    rng = random.Random(42)    sample_counts = {s: 0 for s in states}    trials = 10000
    for _ in range(trials):        s_prime = env.sample_next_state("SafeRoad", "Forward", rng)        sample_counts[s_prime] += 1
    empirical_dist = {s: round(count / trials, 4) for s, count in sample_counts.items()}    print("\nEmpirical frequencies after 10,000 trials from SafeRoad:")    print(empirical_dist)
    # 3. Bellman expectation evaluation: E[V(S_{t+1}) | S_t = SafeRoad, A_t = Forward]    state_values = {"SafeRoad": 2.0, "IcyPatch": 5.0, "Goal": 10.0}    reward = -1.0    gamma = 0.90
    expected_next_v = sum(        prob * state_values[s_next]        for s_next, prob in transitions[("SafeRoad", "Forward")].items()    )    q_value = reward + gamma * expected_next_v
    print(f"\nExpected successor value E[V(S')]: {expected_next_v:.2f}")    print(f"Action value Q(SafeRoad, Forward):  {q_value:.2f}")
    # Explicit test assertions    assert round(mu_1[1], 2) == 0.70, "Step 1 IcyPatch probability must equal 0.70"    assert round(mu_2[2], 2) == 0.57, "Step 2 Goal probability must equal 0.57"    assert round(q_value, 2) == 4.13, "Q-value must equal 4.13"
Initial distribution (t=0): [1.0, 0.0, 0.0]Step 1 distribution  (t=1): [0.1, 0.7, 0.2]Step 2 distribution  (t=2): [0.15, 0.28, 0.57]
Empirical frequencies after 10,000 trials from SafeRoad:{'SafeRoad': 0.0986, 'IcyPatch': 0.7045, 'Goal': 0.1969}
Expected successor value E[V(S')]: 5.70Action value Q(SafeRoad, Forward):  4.13

Watch Out For

Assuming deterministic transitions in real-world systems

A common pitfall is constructing simulators or planning models that treat physics as deterministic (P(s′∣s,a)=1.0P(s' \mid s, a) = 1.0 for a single successor state). When algorithms like Monte Carlo Tree Search (MCTS) or Value Iteration plan under deterministic dynamics, they find fragile policies that cut hair-thin margins next to hazards (e.g. driving right along a cliff edge because the planned path technically never touches it).

When transferred to physical hardware, micro-disturbances, motor friction variances, wind resistance, and sensor latency disrupt the intended trajectory, causing catastrophic failure. To prevent this sim-to-real gap, practitioners must explicitly incorporate transition stochasticity through domain randomization (randomly jittering friction, mass, and delay parameters during rollout generation), using robust MDP formulations, or enforcing entropy-regularized exploration.

Violating the row-stochasticity constraint

When building custom simulators, learning environment dynamics with neural networks, or manually designing transition probability tables, practitioners frequently introduce subtle normalization errors where rows do not sum to unity (∑s′P^(s′∣s,a)≠1\sum_{s'} \hat{P}(s' \mid s, a) \neq 1) or produce negative probability mass.

If row sums exceed 1.0, Bellman updates artificially inflate future value estimates, causing value functions to diverge to positive infinity. Conversely, if row sums fall below 1.0, probability mass leaks on every iteration, shrinking value estimates as if an artificial discount were applied. For learned dynamics models, always apply a strictly normalized softmax projection over successor state logits:

P^(s′∣s,a)=exp⁡(fθ(s,a,s′))∑s′′∈Sexp⁡(fθ(s,a,s′′))\hat{P}(s' \mid s, a) = \frac{\exp\left(f_\theta(s, a, s')\right)}{\sum_{s'' \in \mathcal{S}} \exp\left(f_\theta(s, a, s'')\right)}

In tabular environments, add assertion checks confirming that every row sums to 1.0±10−71.0 \pm 10^{-7} and contains zero negative elements.

The Quick Version

  • Transition dynamics P(s′∣s,a)=Pr⁡(St+1=s′∣St=s,At=a)P(s' \mid s, a) = \Pr(S_{t+1} = s' \mid S_t = s, A_t = a) govern the environment's physics by assigning a probability distribution over all possible next states given the current state and chosen action.
  • The Markov property guarantees that future transitions depend solely on the current state-action pair (s,a)(s, a), making the preceding trajectory history conditionally independent.
  • Every valid transition matrix must be row-stochastic: all probabilities must be non-negative and every row must sum to exactly 1.01.0 (∑s′P(s′∣s,a)=1\sum_{s'} P(s' \mid s, a) = 1).
  • Transition dynamics weight future value estimates during Bellman updates; model-based RL plans directly with P(s′∣s,a)P(s' \mid s, a), while model-free methods sample from it via experience.