Skip to content
AI360Xpert
Beta

Off-Policy Importance Sampling

Evaluate a target policy without ever running it by observing a different behavior policy and re-weighting returns by action probability ratios.

Off-policy importance sampling adjusts returns collected under an exploratory behavior policy by multiplying action probability ratios to estimate target policy value.
Off-policy importance sampling adjusts returns collected under an exploratory behavior policy by multiplying action probability ratios to estimate target policy value.

Why Does This Exist?

In reinforcement learning, learning on-policy creates a persistent conflict between exploration and control. To find an optimal policy, an agent must take exploratory, potentially suboptimal actions. But to evaluate or execute an optimal policy, the agent must act greedily.

In on-policy methods, an agent evaluates the very policy that dictates its behavior. If that behavior explores, the agent evaluates an exploratory policy rather than the true optimal policy. Furthermore, all historical data collected under previous policies becomes obsolete and must be discarded whenever the agent updates its policy.

Off-policy prediction resolves this dilemma by decoupling the learning process into two distinct roles:

  1. Target Policy (π\pi): The policy being evaluated, optimized, or learned (often a deterministic, greedy, or candidate control policy).
  2. Behavior Policy (bb): The policy used to generate actions and gather experience in the environment (often exploratory, safe, or an existing legacy controller).

Because the agent follows behavior policy bb, the raw observed returns reflect the distribution induced by bb, not π\pi. Without correction, averaging these returns yields a heavily biased estimate of π\pi's value. Importance sampling provides the exact statistical re-weighting needed to transform sample returns gathered under bb into an unbiased estimate of the value under π\pi.

Think of It Like This

Demographic Polling Adjustments

Imagine a polling organization wants to predict how the entire voting electorate (the target policy π\pi) will vote on upcoming referendums. Due to collection constraints, their field survey only managed to interview college students on campus (the behavior policy bb).

In the surveyed student sample, 80% support public transit expansion and 20% support highway expansions. However, historical census data reveals that in the general electorate, only 40% prioritize public transit while 60% prioritize highway expansions.

If the pollsters simply averaged the raw student responses, their forecast for the general election would be hopelessly skewed. To correct for this sampling bias, they apply an adjustment factor to each response:

  • A vote for public transit is multiplied by 0.400.80=0.50\frac{0.40}{0.80} = 0.50 (down-weighted because transit supporters were overrepresented).
  • A vote for highway expansion is multiplied by 0.600.20=3.00\frac{0.60}{0.20} = 3.00 (up-weighted because highway supporters were underrepresented).

By multiplying every recorded response by the ratio π(group)b(group)\frac{\pi(\text{group})}{b(\text{group})}, the pollsters compute an unbiased estimate of the general population's preferences using data from a completely different group.

Where the analogy stops: Demographic polling evaluates a single, independent choice per respondent. In reinforcement learning, an episode is a sequential chain of decisions over time. The agent must multiply these probability ratios across every successive time step of the trajectory, creating a compounding product that can rapidly fluctuate.

How It Actually Works

Target Policy vs. Behavior Policy and the Importance Sampling Ratio

Let π(a∣s)\pi(a \mid s) denote the probability of choosing action aa in state ss under the target policy, and let b(a∣s)b(a \mid s) denote the corresponding probability under the behavior policy.

To estimate the value function Vπ(s)V^\pi(s) from episodes generated by bb, we must satisfy the coverage assumption: π(a∣s)>0  ⟹  b(a∣s)>0∀s∈S,a∈A(s)\pi(a \mid s) > 0 \implies b(a \mid s) > 0 \quad \forall s \in \mathcal{S}, a \in \mathcal{A}(s)

If the target policy π\pi could take an action in a state that the behavior policy bb never attempts (b(a∣s)=0b(a \mid s) = 0), the agent will never observe data for that transition, making it impossible to evaluate π\pi.

Consider an episode trajectory segment starting at time step tt and terminating at time TT: τ=(At,St+1,At+1,…,ST)\tau = (A_t, S_{t+1}, A_{t+1}, \dots, S_T)

Under behavior policy bb, the probability of observing this sequence of actions and states given starting state StS_t is: P(τ∣St,b)=∏k=tT−1b(Ak∣Sk)p(Sk+1∣Sk,Ak)\mathbb{P}(\tau \mid S_t, b) = \prod_{k=t}^{T-1} b(A_k \mid S_k) p(S_{k+1} \mid S_k, A_k) where p(Sk+1∣Sk,Ak)p(S_{k+1} \mid S_k, A_k) represents the environment's transition dynamics.

Under target policy π\pi, the probability of that identical trajectory would be: P(τ∣St,π)=∏k=tT−1π(Ak∣Sk)p(Sk+1∣Sk,Ak)\mathbb{P}(\tau \mid S_t, \pi) = \prod_{k=t}^{T-1} \pi(A_k \mid S_k) p(S_{k+1} \mid S_k, A_k)

The importance sampling ratio ρt:T−1\rho_{t:T-1} is defined as the relative probability of the trajectory under π\pi compared to bb: ρt:T−1=∏k=tT−1π(Ak∣Sk)p(Sk+1∣Sk,Ak)∏k=tT−1b(Ak∣Sk)p(Sk+1∣Sk,Ak)=∏k=tT−1π(Ak∣Sk)b(Ak∣Sk)\rho_{t:T-1} = \frac{\prod_{k=t}^{T-1} \pi(A_k \mid S_k) p(S_{k+1} \mid S_k, A_k)}{\prod_{k=t}^{T-1} b(A_k \mid S_k) p(S_{k+1} \mid S_k, A_k)} = \prod_{k=t}^{T-1} \frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)}

The transition probabilities p(Sk+1∣Sk,Ak)p(S_{k+1} \mid S_k, A_k) are identical in numerator and denominator and cancel out completely. This makes importance sampling fully model-free: we do not need to know the environment's transition probabilities.

The expected value of the return GtG_t under behavior policy bb, scaled by ρt:T−1\rho_{t:T-1}, matches the true expectation under target policy π\pi: Eb[ρt:T−1Gt∣St=s]=∑τP(τ∣s,b)(P(τ∣s,π)P(τ∣s,b)Gt(τ))=∑τP(τ∣s,π)Gt(τ)=Vπ(s)\mathbb{E}_b[\rho_{t:T-1} G_t \mid S_t = s] = \sum_{\tau} \mathbb{P}(\tau \mid s, b) \left( \frac{\mathbb{P}(\tau \mid s, \pi)}{\mathbb{P}(\tau \mid s, b)} G_t(\tau) \right) = \sum_{\tau} \mathbb{P}(\tau \mid s, \pi) G_t(\tau) = V^\pi(s)

In ordinary importance sampling, an agent samples NN episodes starting from state ss under behavior policy bb and computes the empirical mean: V(s)=1N∑i=1Nρt(i):T(i)−1(i)Gt(i)V(s) = \frac{1}{N} \sum_{i=1}^N \rho_{t(i):T(i)-1}^{(i)} G_t^{(i)}

Worked numerical example

Consider a 3-step episode generated under behavior policy bb with discount factor γ=1.0\gamma = 1.0: τ=(S0,A0,R1=2,S1,A1,R2=1,S2,A2,R3=5,S3)\tau = (S_0, A_0, R_1=2, S_1, A_1, R_2=1, S_2, A_2, R_3=5, S_3)

The observed unweighted return from S0S_0 is: G0=R1+R2+R3=2.0+1.0+5.0=8.0G_0 = R_1 + R_2 + R_3 = 2.0 + 1.0 + 5.0 = 8.0

The target policy π\pi and behavior policy bb assign the following probabilities to the actions chosen at each time step:

  1. Step k=0k = 0 at state S0S_0 taking action A0A_0: π(A0∣S0)=0.80,b(A0∣S0)=0.50  ⟹  π(A0∣S0)b(A0∣S0)=0.800.50=1.60\pi(A_0 \mid S_0) = 0.80, \quad b(A_0 \mid S_0) = 0.50 \implies \frac{\pi(A_0 \mid S_0)}{b(A_0 \mid S_0)} = \frac{0.80}{0.50} = 1.60

  2. Step k=1k = 1 at state S1S_1 taking action A1A_1: π(A1∣S1)=0.40,b(A1∣S1)=0.80  ⟹  π(A1∣S1)b(A1∣S1)=0.400.80=0.50\pi(A_1 \mid S_1) = 0.40, \quad b(A_1 \mid S_1) = 0.80 \implies \frac{\pi(A_1 \mid S_1)}{b(A_1 \mid S_1)} = \frac{0.40}{0.80} = 0.50

  3. Step k=2k = 2 at state S2S_2 taking action A2A_2: π(A2∣S2)=0.50,b(A2∣S2)=0.25  ⟹  π(A2∣S2)b(A2∣S2)=0.500.25=2.00\pi(A_2 \mid S_2) = 0.50, \quad b(A_2 \mid S_2) = 0.25 \implies \frac{\pi(A_2 \mid S_2)}{b(A_2 \mid S_2)} = \frac{0.50}{0.25} = 2.00

The cumulative trajectory importance sampling ratio is the product across all three steps: ρ0:2=1.60×0.50×2.00=1.60\rho_{0:2} = 1.60 \times 0.50 \times 2.00 = 1.60

The re-weighted return contribution from this trajectory is: ρ0:2G0=1.60×8.0=12.80\rho_{0:2} G_0 = 1.60 \times 8.0 = 12.80

Now suppose a second episode starting from S0S_0 yields raw return G0(2)=4.0G_0^{(2)} = 4.0 with cumulative ratio ρ0:T−1(2)=0.75\rho_{0:T-1}^{(2)} = 0.75, giving a re-weighted return of 0.75×4.0=3.000.75 \times 4.0 = 3.00.

The ordinary importance sampling estimate of Vπ(S0)V^\pi(S_0) across these two episodes is: V(S0)=12.80+3.002=7.90V(S_0) = \frac{12.80 + 3.00}{2} = 7.90

Even though the target policy π\pi was never executed in the environment, the scaled returns provide an unbiased estimate of how well π\pi performs.

Code

from typing import List, NamedTuple
class Transition(NamedTuple):    state: str    action: str    reward: float    pi_prob: float  # Target policy probability: pi(a|s)    b_prob: float   # Behavior policy probability: b(a|s)
def evaluate_trajectory_importance_sampling(    trajectory: List[Transition],     gamma: float = 1.0) -> tuple[float, float, float]:    """    Computes raw return G_0, cumulative importance sampling ratio rho,    and the re-weighted return rho * G_0 for off-policy prediction.    """    rho = 1.0    for step in trajectory:        assert step.b_prob > 0.0, "Coverage violation: b(a|s) must be > 0"        ratio = step.pi_prob / step.b_prob        rho *= ratio
    # Calculate discounted return G_0    g = 0.0    for t, step in enumerate(trajectory):        g += (gamma ** t) * step.reward
    weighted_return = rho * g    return g, rho, weighted_return
def ordinary_importance_sampling_estimate(    episodes: List[List[Transition]],     gamma: float = 1.0) -> float:    """Computes the sample average of weighted returns across episodes."""    total_weighted = sum(        evaluate_trajectory_importance_sampling(ep, gamma)[2]         for ep in episodes    )    return total_weighted / len(episodes)
# Episode 1 (from worked example)ep1 = [    Transition(state="S0", action="A0", reward=2.0, pi_prob=0.80, b_prob=0.50),    Transition(state="S1", action="A1", reward=1.0, pi_prob=0.40, b_prob=0.80),    Transition(state="S2", action="A2", reward=5.0, pi_prob=0.50, b_prob=0.25),]
# Episode 2ep2 = [    Transition(state="S0", action="A1", reward=1.0, pi_prob=0.30, b_prob=0.40),    Transition(state="S1", action="A0", reward=3.0, pi_prob=1.00, b_prob=1.00),]
g1, rho1, weighted1 = evaluate_trajectory_importance_sampling(ep1)g2, rho2, weighted2 = evaluate_trajectory_importance_sampling(ep2)
print(f"Episode 1: Raw Return = {g1:.2f}, rho = {rho1:.2f}, Weighted = {weighted1:.2f}")print(f"Episode 2: Raw Return = {g2:.2f}, rho = {rho2:.2f}, Weighted = {weighted2:.2f}")
v_estimate = ordinary_importance_sampling_estimate([ep1, ep2])print(f"Estimated V^pi(S0) = {v_estimate:.2f}")
# -> Expected output:# Episode 1: Raw Return = 8.00, rho = 1.60, Weighted = 12.80# Episode 2: Raw Return = 4.00, rho = 0.75, Weighted = 3.00# Estimated V^pi(S0) = 7.90

Watch Out For

Exponential Variance Explosion Along Long Trajectories

Ordinary importance sampling is strictly unbiased, but its variance can be astronomically high or even theoretically infinite. Because the trajectory weight ρt:T−1=∏k=tT−1π(Ak∣Sk)b(Ak∣Sk)\rho_{t:T-1} = \prod_{k=t}^{T-1} \frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)} is a product of ratios, small differences between π\pi and bb compound exponentially with trajectory length TT.

In long horizons, most trajectories under bb take at least one action that π\pi considers unlikely, causing ρ\rho to drop to 00. Conversely, an occasional trajectory aligns with high-probability actions under π\pi, producing an enormous ρ\rho value. The resulting empirical average oscillates wildly, requiring impractical numbers of samples to converge.

The Fix:

  1. Switch to Weighted Importance Sampling (WIS), which divides by the sum of importance ratios (∑iρ(i)\sum_i \rho^{(i)}) rather than the episode count NN. WIS introduces a mild bias but guarantees bounded weights and dramatically lower variance.
  2. Use discounting-aware importance sampling or temporal-difference (TD) methods (such as Q-learning or Expected SARSA), which bootstrap after 1 step to avoid cumulative multi-step product expansion.

The Quick Version

  • Off-policy prediction evaluates a target policy π\pi using trajectory data generated by a separate behavior policy bb.
  • The coverage assumption requires b(a∣s)>0b(a \mid s) > 0 for every state-action pair where π(a∣s)>0\pi(a \mid s) > 0.
  • The importance sampling ratio ρt:T−1=∏k=tT−1π(Ak∣Sk)b(Ak∣Sk)\rho_{t:T-1} = \prod_{k=t}^{T-1} \frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)} cancels environment transition dynamics, enabling model-free re-weighting.
  • Ordinary importance sampling is strictly unbiased (Eb[ρGt]=Vπ(St)\mathbb{E}_b[\rho G_t] = V^\pi(S_t)), but suffers from extreme variance over long horizons due to compounding products of ratios.