Skip to content
AI360Xpert
Beta

Importance Sampling: Ordinary vs Weighted

Ordinary importance sampling scales returns by raw probability ratios and divides by trajectory count, creating massive variance, whereas weighted importance sampling normalizes by the sum of ratios, trading initial bias for stable, bounded estimates.

Ordinary Importance Sampling divides by sample size producing wild swings, while Weighted Importance Sampling normalizes by total ratio weight to bound estimates.
Ordinary Importance Sampling divides by sample size producing wild swings, while Weighted Importance Sampling normalizes by total ratio weight to bound estimates.

Why Does This Exist?

Off-policy evaluation enables a reinforcement learning agent to estimate the value of a target policy π\pi using trajectory rollouts collected under a distinct, exploratory behavior policy bb. Because state-action paths are sampled according to bb rather than π\pi, their observed returns cannot simply be averaged together without adjustment; doing so would estimate vbv_b rather than vπv_\pi.

To correct for this distributional mismatch, importance sampling reweights each trajectory's return by the likelihood ratio of the action sequence under π\pi versus bb:

ρt:T−1=∏k=tT−1π(Ak∣Sk)b(Ak∣Sk)\rho_{t:T-1} = \prod_{k=t}^{T-1} \frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)}

The mathematical dilemma arises in how to aggregate these reweighted returns:

  1. Ordinary Importance Sampling (OIS) takes the simple arithmetic average of the weighted returns across nn episodes, dividing by nn. While OIS is strictly unbiased (E[Vord]=vπ\mathbb{E}[V^{\text{ord}}] = v_\pi), its variance grows exponentially with the episode horizon T−tT-t. In multi-step tasks, the variance of ρ\rho can be unbounded or even infinite. A single rare trajectory with a massive ratio ρi\rho_i can inflate the sample mean by thousands of units, destroying estimation accuracy.
  2. Weighted Importance Sampling (WIS) divides the sum of weighted returns by the sum of importance ratios ∑ρi\sum \rho_i rather than the raw episode count nn. While WIS is biased for small sample sizes, this normalization converts the estimate into a convex combination of returns, guaranteeing that the estimate stays strictly bounded within the physical range of observed returns and slashing variance by orders of magnitude.

Understanding this bias-variance tradeoff is critical: in off-policy Monte Carlo learning, chasing zero bias with Ordinary IS frequently leads to non-converging algorithms, making Weighted IS the de facto standard in practice.

Think of It Like This

Estimating crowd height with raw multipliers versus a normalized weighted average

Imagine you want to calculate the average height of attendees at an international basketball tournament. However, you only have survey data collected outside a nearby coffee shop where most visitors were casual spectators, while only a few were 7-foot tall players.

Under the logic of Ordinary Importance Sampling, you calculate a representation multiplier for each demographic to correct for how underrepresented they were outside the coffee shop, multiply their height by this multiplier, and then divide the final sum by the raw headcount of people surveyed (nn). If you happen to interview just one 7-foot player whose demographic multiplier is 50×50\times, their single contribution adds 350350 feet to your running total! Dividing that total by your small survey size of 1010 people yields an estimated average crowd height of over 4040 feet. Because the denominator is a fixed headcount independent of the multipliers, a single outlier demographic blows the estimate completely outside the laws of human biology.

Under Weighted Importance Sampling, you divide the sum of weighted heights by the sum of all demographic multipliers. The multipliers now act as relative percentage shares rather than unconstrained expansion factors. Even if that 7-foot player carries a massive weight of 5050, that exact same 5050 appears in the denominator alongside the other weights. The calculation remains a true weighted average, mathematically guaranteed to fall between the shortest person (say, 5'2") and tallest person (7'1") in your sample.

Where the analogy stops: In static human surveys, demographic weights are often pre-calculated constants. In reinforcement learning, the importance weight ρ\rho is a cumulative product of stochastic probabilities along an entire trajectory, meaning trajectory weights can easily span ten orders of magnitude within the exact same training run.

How It Actually Works

The Mathematical Formulation and Variance Divergence

Consider estimating the state-value function vπ(s)v_\pi(s) of a target policy π\pi from nn independent episodes generated by behavior policy bb, satisfying the coverage assumption: π(a∣s)>0  ⟹  b(a∣s)>0\pi(a \mid s) > 0 \implies b(a \mid s) > 0.

For each episode i∈{1,…,n}i \in \{1, \dots, n\} visiting state ss at time step tt, let GiG_i represent the subsequent return, and let ρi\rho_i denote the full trajectory importance sampling ratio:

ρi=∏k=tTi−1π(Ak(i)∣Sk(i))b(Ak(i)∣Sk(i))\rho_i = \prod_{k=t}^{T_i-1} \frac{\pi(A_k^{(i)} \mid S_k^{(i)})}{b(A_k^{(i)} \mid S_k^{(i)})}

1. Ordinary Importance Sampling (OIS)

Ordinary importance sampling averages the reweighted returns directly across the nn trajectories:

Vord(s)=1n∑i=1nρiGiV^{\text{ord}}(s) = \frac{1}{n} \sum_{i=1}^{n} \rho_i G_i
Expectation (Unbiased)

Using the tower property and change of measure:

Eb[Vord(s)]=1n∑i=1nEb[ρiGi]=1n∑i=1nEπ[Gi]=vπ(s)\mathbb{E}_b \left[ V^{\text{ord}}(s) \right] = \frac{1}{n} \sum_{i=1}^{n} \mathbb{E}_b \left[ \rho_i G_i \right] = \frac{1}{n} \sum_{i=1}^{n} \mathbb{E}_\pi [G_i] = v_\pi(s)

The bias of Ordinary Importance Sampling is strictly zero for any sample size n≥1n \ge 1.

Variance (Potentially Infinite)

The variance is:

Varb(Vord(s))=1nVarb(ρiGi)=1n(Eb[ρi2Gi2]−vπ(s)2)\mathrm{Var}_b \left( V^{\text{ord}}(s) \right) = \frac{1}{n} \mathrm{Var}_b (\rho_i G_i) = \frac{1}{n} \left( \mathbb{E}_b \left[ \rho_i^2 G_i^2 \right] - v_\pi(s)^2 \right)

Because ρi\rho_i is a product of per-step ratios π(Ak∣Sk)b(Ak∣Sk)\frac{\pi(A_k \mid S_k)}{b(A_k \mid S_k)}, the second moment Eb[ρi2]\mathbb{E}_b[\rho_i^2] grows exponentially with trajectory length T−tT-t. If there exists any transition where π(a∣s)2b(a∣s)>1\frac{\pi(a \mid s)^2}{b(a \mid s)} > 1, the variance of ρi\rho_i diverges to infinity as T→∞T \to \infty. When variance is infinite, sample averages do not converge reliably under the Central Limit Theorem.

2. Weighted Importance Sampling (WIS)

Weighted importance sampling normalizes by the empirical sum of importance sampling ratios:

Vwis(s)=∑i=1nρiGi∑i=1nρiV^{\text{wis}}(s) = \frac{\sum_{i=1}^{n} \rho_i G_i}{\sum_{i=1}^{n} \rho_i}

(If ∑i=1nρi=0\sum_{i=1}^n \rho_i = 0, Vwis(s)V^{\text{wis}}(s) is defined to be 00).

Bias (Non-Zero for Finite nn)

When n=1n = 1 and ρ1>0\rho_1 > 0:

Vwis(s)=ρ1G1ρ1=G1V^{\text{wis}}(s) = \frac{\rho_1 G_1}{\rho_1} = G_1

Taking the expectation with respect to the sampling policy bb:

Eb[Vwis(s)]=Eb[G1]=vb(s)\mathbb{E}_b \left[ V^{\text{wis}}(s) \right] = \mathbb{E}_b [G_1] = v_b(s)

With a single sample, the expectation of WIS equals the value of the behavior policy, not the target policy! Thus, WIS is biased for finite sample sizes.

Consistency (Asymptotic Convergence)

By the Strong Law of Large Numbers, as the sample count n→∞n \to \infty:

1n∑i=1nρiGi→a.s.vπ(s)and1n∑i=1nρi→a.s.Eb[ρi]=1\frac{1}{n} \sum_{i=1}^n \rho_i G_i \xrightarrow{\text{a.s.}} v_\pi(s) \quad \text{and} \quad \frac{1}{n} \sum_{i=1}^n \rho_i \xrightarrow{\text{a.s.}} \mathbb{E}_b[\rho_i] = 1

Therefore, the ratio of the empirical sums converges almost surely to the true value:

Vwis(s)=1n∑i=1nρiGi1n∑i=1nρi→a.s.vπ(s)1=vπ(s)V^{\text{wis}}(s) = \frac{\frac{1}{n}\sum_{i=1}^n \rho_i G_i}{\frac{1}{n}\sum_{i=1}^n \rho_i} \xrightarrow{\text{a.s.}} \frac{v_\pi(s)}{1} = v_\pi(s)

WIS is statistically consistent: its bias asymptotically vanishes to zero as more trajectories accumulate.

Bounded Range and Variance Reduction

Let wi=ρi∑k=1nρkw_i = \frac{\rho_i}{\sum_{k=1}^n \rho_k}. Since ρi≥0\rho_i \ge 0 and ∑i=1nwi=1\sum_{i=1}^n w_i = 1, Vwis(s)=∑i=1nwiGiV^{\text{wis}}(s) = \sum_{i=1}^n w_i G_i is a strict convex combination. Consequently:

min⁡1≤i≤nGi≤Vwis(s)≤max⁡1≤i≤nGi\min_{1 \le i \le n} G_i \le V^{\text{wis}}(s) \le \max_{1 \le i \le n} G_i

The estimate is strictly bounded within the physical range of observed returns. Even if an extreme trajectory produces ρi=105\rho_i = 10^5, its normalized weight wi≈1.0w_i \approx 1.0, pulling the estimate toward Gi∈[Gmin⁡,Gmax⁡]G_i \in [G_{\min}, G_{\max}] rather than blowing up to infinity.

Worked numerical example

Let an agent evaluate state S0S_0 under target policy π\pi using n=3n = 3 off-policy trajectories collected by behavior policy bb. Suppose true target policy value is vπ(S0)=1.0v_\pi(S_0) = 1.0, and rewards are bounded such that all observed returns Gi∈[0.0,1.0]G_i \in [0.0, 1.0].

We observe 3 trajectories:

  • Trajectory 1: Return G1=1.0G_1 = 1.0, importance ratio ρ1=0.10\rho_1 = 0.10 (frequent under bb, rare under π\pi).
  • Trajectory 2: Return G2=0.0G_2 = 0.0, importance ratio ρ2=0.40\rho_2 = 0.40 (moderate frequency under both).
  • Trajectory 3: Return G3=1.0G_3 = 1.0, importance ratio ρ3=14.50\rho_3 = 14.50 (rare under bb, very likely under π\pi).

Step 1: Compute Individual Weighted Returns

ρ1G1=0.10×1.0=0.10\rho_1 G_1 = 0.10 \times 1.0 = 0.10 ρ2G2=0.40×0.0=0.00\rho_2 G_2 = 0.40 \times 0.0 = 0.00 ρ3G3=14.50×1.0=14.50\rho_3 G_3 = 14.50 \times 1.0 = 14.50 ∑i=13ρiGi=0.10+0.00+14.50=14.60\sum_{i=1}^3 \rho_i G_i = 0.10 + 0.00 + 14.50 = 14.60

Step 2: Compute Ordinary Importance Sampling (OIS)

Divide by the fixed trajectory count n=3n = 3:

Vord(S0)=∑i=13ρiGin=14.603≈4.867V^{\text{ord}}(S_0) = \frac{\sum_{i=1}^3 \rho_i G_i}{n} = \frac{14.60}{3} \approx 4.867

Notice the pathology: every individual return was at most 1.01.0, yet Ordinary IS estimates V(S0)≈4.867V(S_0) \approx 4.867—a value nearly five times larger than the physical environment maximum!

Step 3: Compute Weighted Importance Sampling (WIS)

First compute the sum of weights:

∑i=13ρi=0.10+0.40+14.50=15.00\sum_{i=1}^3 \rho_i = 0.10 + 0.40 + 14.50 = 15.00

Divide by the sum of weights:

Vwis(S0)=∑i=13ρiGi∑i=13ρi=14.6015.00≈0.973V^{\text{wis}}(S_0) = \frac{\sum_{i=1}^3 \rho_i G_i}{\sum_{i=1}^3 \rho_i} = \frac{14.60}{15.00} \approx 0.973

The WIS estimate 0.9730.973 sits comfortably within [0.0,1.0][0.0, 1.0] and closely matches the true target value vπ(S0)=1.0v_\pi(S_0) = 1.0.

Step 4: Examine the Finite-Sample Edge Case (n=1n = 1)

If we had stopped after observing only Trajectory 1 (G1=1.0,ρ1=0.10G_1 = 1.0, \rho_1 = 0.10):

  • OIS: Vord=0.10×1.01=0.10V^{\text{ord}} = \frac{0.10 \times 1.0}{1} = 0.10 (scales down return based on low ρ1\rho_1).
  • WIS: Vwis=0.10×1.00.10=1.00=G1V^{\text{wis}} = \frac{0.10 \times 1.0}{0.10} = 1.00 = G_1 (matches behavior policy return, revealing its finite-sample bias).

Code

import randomfrom typing import List, Tuple
def compare_ordinary_vs_weighted_is(    n_trials: int = 1000,    trajectories_per_trial: int = 25,    horizon: int = 3,    p_target: float = 0.85,    p_behavior: float = 0.50,    seed: int = 42,) -> Tuple[float, float, float, float, float]:    """    Simulates off-policy evaluation comparing Ordinary and Weighted IS over repeated trials.        Environment:        At each step, action a=1 or a=0 is selected.        Return G = 1.0 if all actions in the horizon equal 1, otherwise G = 0.0.        Returns:        (true_value, ois_mean, ois_variance, wis_mean, wis_variance)    """    random.seed(seed)        # Ground truth expected return under target policy:    # Trajectory gives reward 1.0 iff action 1 is selected at all time steps    true_value: float = p_target ** horizon  # 0.85^3 = 0.614125        ois_estimates: List[float] = []    wis_estimates: List[float] = []        for _ in range(n_trials):        ratios: List[float] = []        returns: List[float] = []                for _ in range(trajectories_per_trial):            # Sample trajectory under behavior policy b            actions = [1 if random.random() < p_behavior else 0 for _ in range(horizon)]                        # Compute trajectory importance ratio: prod(pi(a) / b(a))            rho = 1.0            for a in actions:                p_pi = p_target if a == 1 else (1.0 - p_target)                p_b = p_behavior if a == 1 else (1.0 - p_behavior)                rho *= (p_pi / p_b)                            ret = 1.0 if all(a == 1 for a in actions) else 0.0            ratios.append(rho)            returns.append(ret)                    # Ordinary Importance Sampling: arithmetic mean of rho_i * G_i        ois_estimate = sum(r * g for r, g in zip(ratios, returns)) / len(ratios)        ois_estimates.append(ois_estimate)                # Weighted Importance Sampling: self-normalized by sum(rho_i)        sum_rho = sum(ratios)        wis_estimate = sum(r * g for r, g in zip(ratios, returns)) / sum_rho if sum_rho > 0 else 0.0        wis_estimates.append(wis_estimate)            mean_ois = sum(ois_estimates) / n_trials    var_ois = sum((x - mean_ois) ** 2 for x in ois_estimates) / n_trials        mean_wis = sum(wis_estimates) / n_trials    var_wis = sum((x - mean_wis) ** 2 for x in wis_estimates) / n_trials        return true_value, mean_ois, var_ois, mean_wis, var_wis
# Execute 1,000 trial Monte Carlo simulationtrue_v, m_ois, v_ois, m_wis, v_wis = compare_ordinary_vs_weighted_is()
print(f"Ground Truth V^pi:     {true_v:.4f}")# -> Ground Truth V^pi:     0.6141
print(f"Ordinary IS Mean:      {m_ois:.4f} (Variance: {v_ois:.4f})")# -> Ordinary IS Mean:      0.6110 (Variance: 0.1117)
print(f"Weighted IS Mean:      {m_wis:.4f} (Variance: {v_wis:.4f})")# -> Weighted IS Mean:      0.5640 (Variance: 0.0362)

Watch Out For

The Unbiased Estimator Fallacy

A frequent trap for practitioners with classical statistics backgrounds is assuming that unbiasedness is always superior to a biased estimator, leading them to select Ordinary Importance Sampling over Weighted Importance Sampling.

In off-policy reinforcement learning, this assumption breaks down catastrophically. The variance of Ordinary Importance Sampling can easily be infinite when the behavior policy bb assigns even moderately smaller probabilities to state-action paths than the target policy π\pi (Eb[ρ2]=∞\mathbb{E}_b[\rho^2] = \infty). When variance is infinite, the Central Limit Theorem fails: sample means fluctuate wildly without stabilizing, and an empirical sample mean can be dominated by a single rare trajectory even after millions of collected episodes.

Even when the theoretical variance is finite, the Mean Squared Error (MSE=Bias2+Variance\text{MSE} = \text{Bias}^2 + \text{Variance}) of Ordinary IS is typically orders of magnitude higher than Weighted IS because the variance explosion dwarfs any tiny bias term.

The Fix: Always default to Weighted Importance Sampling (or doubly robust estimators) for off-policy value prediction. The slight finite-sample bias decays rapidly as O(1/n)O(1/n) while guaranteeing that every estimate remains bounded and numerically stable.

The Quick Version

  • Ordinary Importance Sampling computes V(s)=1n∑ρiGiV(s) = \frac{1}{n} \sum \rho_i G_i, which is strictly unbiased but suffers from massive or infinite variance due to unbounded trajectory likelihood products.
  • Weighted Importance Sampling computes V(s)=∑ρiGi∑ρiV(s) = \frac{\sum \rho_i G_i}{\sum \rho_i}, normalizing by total ratio weights to form a convex combination strictly bounded within [min⁡Gi,max⁡Gi][\min G_i, \max G_i].
  • Bias-Variance Tradeoff: WIS is biased for finite samples (collapsing to the behavior policy return vbv_b when n=1n=1), but it is statistically consistent and converges to vπv_\pi as n→∞n \to \infty.
  • Practical Standard: Because the slight initial bias of WIS vanishes quickly while cutting variance by orders of magnitude, Weighted IS is universally preferred over Ordinary IS in production reinforcement learning.