Skip to content
AI360Xpert
Beta

Emphatic-TD Methods in RL

Emphatic-TD stabilizes off-policy learning by re-weighting updates with followon traces, ensuring errors in frequently bootstrapped states are corrected before they cause divergence.

Emphatic-TD dynamically scales semi-gradient updates via interest and followon traces, restoring positive definiteness to off-policy learning.
Emphatic-TD dynamically scales semi-gradient updates via interest and followon traces, restoring positive definiteness to off-policy learning.

Why Does This Exist?

In reinforcement learning, the Deadly Triad—combining function approximation, bootstrapping, and off-policy data—causes standard temporal-difference algorithms to diverge. In standard off-policy TD(0)\text{TD}(0), weighting updates with the importance sampling ratio ρt=π(At∣St)b(At∣St)\rho_t = \frac{\pi(A_t|S_t)}{b(A_t|S_t)} corrects for the discrepancy in action selection probabilities, but it does nothing to correct the underlying state distribution mismatch. Transitions remain sampled according to the behavior policy's visitation frequency μb\mu_b rather than the target policy's stationary distribution dπd_\pi.

Because μb≠dπ\mu_b \neq d_\pi, the expected update matrix:

Aoff=X⊤Db(I−γPπ)X\mathbf{A}_{\text{off}} = \mathbf{X}^\top \mathbf{D}_b (\mathbf{I} - \gamma \mathbf{P}_\pi) \mathbf{X}

is not guaranteed to have a positive-definite symmetric part. When Aoff\mathbf{A}_{\text{off}} develops eigenvalues with negative real parts, the spectral radius of the iteration matrix exceeds unity (ρ(I−αAoff)>1\rho(\mathbf{I} - \alpha \mathbf{A}_{\text{off}}) > 1). This creates an unstable positive feedback loop where value estimates grow exponentially to ±∞\pm \infty.

Historically, the primary remedy was Gradient-TD (such as TDC and GTD2). Gradient-TD converts policy evaluation into a saddle-point optimization problem minimizing the Mean Squared Projected Bellman Error (MSPBE). However, Gradient-TD requires:

  1. An auxiliary secondary weight vector ut∈Rd\mathbf{u}_t \in \mathbb{R}^d running alongside wt\mathbf{w}_t.
  2. Two separate learning rates (αt\alpha_t and βt\beta_t) operating on two distinct timescales (βt/αt→0\beta_t / \alpha_t \to 0), which makes hyperparameter tuning delicate and empirical convergence sluggish.

In 2016, Richard Sutton, A. Rupam Mahmood, and Martha White introduced Emphatic Temporal-Difference (ETD) learning. Rather than minimizing a dual projected objective with two timescales, ETD solves the problem directly within a single timescale using a single weight vector wt\mathbf{w}_t and one learning rate α\alpha. By recursively tracking how downstream states inherit bootstrapping obligations through a scalar followon trace, ETD re-weights updates with an emphasis scalar MtM_t. This reshapes the effective state distribution such that the resulting system matrix is provably positive definite, restoring contraction and stability to off-policy semi-gradient learning.

Think of It Like This

The Audio Feedback Compressor

Imagine a live performance sound system with high-gain microphones and loud stage monitors. When sound emitted by a monitor leaks back into a microphone, it creates a closed acoustic loop. If certain resonant frequencies circulate unchecked, the amplifier boosts them on every pass, culminating in an ear-piercing screech of acoustic feedback. This runaway screech is the exact physical analogue of the Deadly Triad.

The Gradient-TD solution is like placing a second, independent digital signal processor (DSP) in the audio rack. The secondary processor runs its own background phase-cancellation algorithm on a slower internal clock, calculating a continuous anti-feedback curve and subtracting it from the primary audio stream. It stops the feedback, but it doubles hardware complexity and introduces latency.

Emphatic-TD, by contrast, is an adaptive feedback compressor built directly into the main audio channel. It maintains a running memory of recent signal surges (the followon trace FtF_t). Whenever the performer hits a note that previously leaked into the monitor and caused an amplification spike downstream, the compressor dynamically modulates channel gain (the emphasis MtM_t) right at that moment. By attenuating resonant peaks and elevating damped passages, the net loop gain across all frequencies stays strictly below unity. The system achieves maximum amplification without screeching, without auxiliary secondary processors.

Where the analogy stops: An audio compressor only reduces signal gain to prevent clipping. Emphatic-TD can dynamically both amplify and attenuate updates—boosting states whose downstream predictions are heavily relied upon, and damping unvisited states—to ensure the mathematical operator strictly contracts in expectation.

How It Actually Works

Followon Traces, Emphasis, and the Emphatic Matrix

Emphatic-TD introduces three interconnected quantities that propagate bootstrapping importance:

  1. User Interest (It≥0I_t \ge 0): A user-specified scalar defining how much the practitioner cares about prediction accuracy in state StS_t. In standard uniform policy evaluation, It=1.0I_t = 1.0 for all states. In selective evaluation (e.g., predicting value only for sub-tasks or start states), ItI_t can be set to 1.01.0 on target states and 0.00.0 elsewhere.

  2. The Followon Trace (FtF_t): When state StS_t bootstraps from future states St+1,St+2,…S_{t+1}, S_{t+2}, \dots, errors in those future states propagate backward to degrade the prediction at StS_t. Therefore, any future state that is bootstrapped into inherits interest from the preceding states that relied upon it. The followon trace recursively accumulates this inherited interest forward in time:

    F0=I0F_0 = I_0

    Ft=γρt−1Ft−1+It(t≥1)F_t = \gamma \rho_{t-1} F_{t-1} + I_t \quad (t \ge 1)

    where ρt−1=π(At−1∣St−1)b(At−1∣St−1)\rho_{t-1} = \frac{\pi(A_{t-1}|S_{t-1})}{b(A_{t-1}|S_{t-1})} is the importance sampling ratio of the transition arriving into state StS_t, and γ∈[0,1)\gamma \in [0, 1) is the discount factor.

  3. The Emphasis Scalar (MtM_t): The scalar weight applied to the semi-gradient update at time tt. For general multi-step ETD(λ)\text{ETD}(\lambda), emphasis balances immediate interest against the accumulated followon trace:

    Mt=λIt+(1−λ)FtM_t = \lambda I_t + (1 - \lambda) F_t

    For the fundamental 1-step algorithm (ETD(0)\text{ETD}(0), where λ=0\lambda = 0):

    Mt=FtM_t = F_t

The eligibility trace vector et∈Rd\mathbf{e}_t \in \mathbb{R}^d for ETD(λ)\text{ETD}(\lambda) combines the emphasis scalar with state features:

et=ρt(γλet−1+Mtx(St)),with e−1=0\mathbf{e}_t = \rho_t \left( \gamma \lambda \mathbf{e}_{t-1} + M_t \mathbf{x}(S_t) \right), \quad \text{with } \mathbf{e}_{-1} = \mathbf{0}

For ETD(0)\text{ETD}(0) (λ=0\lambda = 0), this simplifies to:

et=ρtMtx(St)\mathbf{e}_t = \rho_t M_t \mathbf{x}(S_t)

The parameter update rule is:

wt+1=wt+αδtet=wt+αMtρtδtx(St)\mathbf{w}_{t+1} = \mathbf{w}_t + \alpha \delta_t \mathbf{e}_t = \mathbf{w}_t + \alpha M_t \rho_t \delta_t \mathbf{x}(S_t)

where δt=Rt+1+γv^(St+1,wt)−v^(St,wt)\delta_t = R_{t+1} + \gamma \hat{v}(S_{t+1}, \mathbf{w}_t) - \hat{v}(S_t, \mathbf{w}_t) is the standard temporal difference error.

Why Contraction and Stability are Guaranteed

Under behavior policy bb, the steady-state expectation of the emphasis vector defines an asymptotic diagonal weighting matrix:

demphatic⊤=i⊤(I−γPπ)−1\mathbf{d}_{\text{emphatic}}^\top = \mathbf{i}^\top (\mathbf{I} - \gamma \mathbf{P}_\pi)^{-1}

where i=[I(s1),…,I(sn)]⊤\mathbf{i} = [I(s_1), \dots, I(s_n)]^\top is the interest vector and Pπ\mathbf{P}_\pi is the target policy transition matrix. Post-multiplying both sides by (I−γPπ)(\mathbf{I} - \gamma \mathbf{P}_\pi) yields:

demphatic⊤(I−γPπ)=i⊤>0\mathbf{d}_{\text{emphatic}}^\top (\mathbf{I} - \gamma \mathbf{P}_\pi) = \mathbf{i}^\top > \mathbf{0}

Let Demphatic=diag(demphatic)\mathbf{D}_{\text{emphatic}} = \text{diag}(\mathbf{d}_{\text{emphatic}}). The expected linear system matrix under ETD updates becomes:

Aemphatic=X⊤Demphatic(I−γPπ)X\mathbf{A}_{\text{emphatic}} = \mathbf{X}^\top \mathbf{D}_{\text{emphatic}} (\mathbf{I} - \gamma \mathbf{P}_\pi) \mathbf{X}

Sutton, Mahmood, and White proved that for any interest vector i>0\mathbf{i} > \mathbf{0}, the symmetric part:

12[Demphatic(I−γPπ)+(I−γPπ)⊤Demphatic]\frac{1}{2} \left[ \mathbf{D}_{\text{emphatic}} (\mathbf{I} - \gamma \mathbf{P}_\pi) + (\mathbf{I} - \gamma \mathbf{P}_\pi)^\top \mathbf{D}_{\text{emphatic}} \right]

is strictly positive definite. As a result, Aemphatic\mathbf{A}_{\text{emphatic}} has all eigenvalues with strictly positive real parts, guaranteeing that the ordinary differential equation w˙=bemphatic−Aemphaticw\dot{\mathbf{w}} = \mathbf{b}_{\text{emphatic}} - \mathbf{A}_{\text{emphatic}}\mathbf{w} is asymptotically stable. wt\mathbf{w}_t converges to the unique fixed point w∗=Aemphatic−1bemphatic\mathbf{w}^* = \mathbf{A}_{\text{emphatic}}^{-1} \mathbf{b}_{\text{emphatic}} with probability 1.


Worked numerical example

To observe how the followon trace and emphasis dynamically rescale off-policy updates, consider a two-step trajectory in a 2-state MDP with linear function approximation:

  • State Features: x(s1)=[1.0,0.0]⊤\mathbf{x}(s_1) = [1.0, 0.0]^\top, x(s2)=[0.0,2.0]⊤\mathbf{x}(s_2) = [0.0, 2.0]^\top.
  • Parameters: w0=[1.0,1.0]⊤\mathbf{w}_0 = [1.0, 1.0]^\top. Initial value estimates: v^(s1,w0)=[1.0,0.0]⋅[1.0,1.0]=1.0\hat{v}(s_1, \mathbf{w}_0) = [1.0, 0.0] \cdot [1.0, 1.0] = 1.0 v^(s2,w0)=[0.0,2.0]⋅[1.0,1.0]=2.0\hat{v}(s_2, \mathbf{w}_0) = [0.0, 2.0] \cdot [1.0, 1.0] = 2.0
  • Hyperparameters: γ=0.9\gamma = 0.9, α=0.05\alpha = 0.05, uniform interest It=1.0I_t = 1.0, λ=0\lambda = 0 (ETD(0)\text{ETD}(0)). All transition rewards R=0.0R = 0.0.

Transition 0 (t=0t = 0): S0=s1→S1=s2S_0 = s_1 \to S_1 = s_2

  1. Followon Trace and Emphasis: F0=I0=1.0000,M0=F0=1.0000F_0 = I_0 = 1.0000, \quad M_0 = F_0 = 1.0000
  2. Action Selection and Importance Ratio: Target takes a1a_1 with π(a1∣s1)=1.0\pi(a_1|s_1) = 1.0; behavior takes a1a_1 with b(a1∣s1)=0.5b(a_1|s_1) = 0.5: ρ0=π(a1∣s1)b(a1∣s1)=1.00.5=2.0000\rho_0 = \frac{\pi(a_1|s_1)}{b(a_1|s_1)} = \frac{1.0}{0.5} = 2.0000
  3. Temporal Difference Error: δ0=R1+γv^(s2,w0)−v^(s1,w0)=0.0+0.9(2.0)−1.0=1.8−1.0=+0.8000\delta_0 = R_1 + \gamma \hat{v}(s_2, \mathbf{w}_0) - \hat{v}(s_1, \mathbf{w}_0) = 0.0 + 0.9(2.0) - 1.0 = 1.8 - 1.0 = +0.8000
  4. Parameter Update: w1=w0+αM0ρ0δ0x(s1)\mathbf{w}_1 = \mathbf{w}_0 + \alpha M_0 \rho_0 \delta_0 \mathbf{x}(s_1) w1=[1.01.0]+0.05×1.0×2.0×0.8×[1.00.0]=[1.01.0]+[0.080.0]=[1.08001.0000]\mathbf{w}_1 = \begin{bmatrix} 1.0 \\ 1.0 \end{bmatrix} + 0.05 \times 1.0 \times 2.0 \times 0.8 \times \begin{bmatrix} 1.0 \\ 0.0 \end{bmatrix} = \begin{bmatrix} 1.0 \\ 1.0 \end{bmatrix} + \begin{bmatrix} 0.08 \\ 0.0 \end{bmatrix} = \begin{bmatrix} 1.0800 \\ 1.0000 \end{bmatrix}

Transition 1 (t=1t = 1): S1=s2→S2=s2S_1 = s_2 \to S_2 = s_2

  1. Followon Trace and Emphasis: The trace recursively inherits the importance of the preceding transition arriving from s1s_1: F1=γρ0F0+I1=0.9×2.0×1.0+1.0=1.8+1.0=2.8000F_1 = \gamma \rho_0 F_0 + I_1 = 0.9 \times 2.0 \times 1.0 + 1.0 = 1.8 + 1.0 = 2.8000 M1=F1=2.8000M_1 = F_1 = 2.8000 Notice: The emphasis on state s2s_2 has jumped from 1.01.0 to 2.82.8 because s2s_2 was heavily bootstrapped into by state s1s_1 under off-policy sampling.
  2. Action Selection and Importance Ratio: In state s2s_2, target and behavior both select action a2a_2 deterministically: ρ1=1.01.0=1.0000\rho_1 = \frac{1.0}{1.0} = 1.0000.
  3. Temporal Difference Error: v^(s2,w1)=[0.0,2.0]⋅[1.08,1.0]=2.0000\hat{v}(s_2, \mathbf{w}_1) = [0.0, 2.0] \cdot [1.08, 1.0] = 2.0000 δ1=R2+γv^(s2,w1)−v^(s2,w1)=0.0+0.9(2.0)−2.0=1.8−2.0=−0.2000\delta_1 = R_2 + \gamma \hat{v}(s_2, \mathbf{w}_1) - \hat{v}(s_2, \mathbf{w}_1) = 0.0 + 0.9(2.0) - 2.0 = 1.8 - 2.0 = -0.2000
  4. Parameter Update: w2=w1+αM1ρ1δ1x(s2)\mathbf{w}_2 = \mathbf{w}_1 + \alpha M_1 \rho_1 \delta_1 \mathbf{x}(s_2) w2=[1.08001.0000]+0.05×2.8×1.0×(−0.2)×[0.02.0]\mathbf{w}_2 = \begin{bmatrix} 1.0800 \\ 1.0000 \end{bmatrix} + 0.05 \times 2.8 \times 1.0 \times (-0.2) \times \begin{bmatrix} 0.0 \\ 2.0 \end{bmatrix} w2=[1.08001.0000]+[0.0−0.0560]=[1.08000.9440]\mathbf{w}_2 = \begin{bmatrix} 1.0800 \\ 1.0000 \end{bmatrix} + \begin{bmatrix} 0.0 \\ -0.0560 \end{bmatrix} = \begin{bmatrix} 1.0800 \\ 0.9440 \end{bmatrix}

Because the followon trace magnified M1M_1 to 2.82.8, the corrective update on the downstream state s2s_2 was amplified, enforcing convergence before the feedback loop can expand.

Code

The following self-contained Python script benchmarks standard off-policy TD(0)\text{TD}(0) against Emphatic TD(0)\text{TD}(0) on Baird's 7-state counterexample, demonstrating how ETD maintains bounded weights while standard TD diverges.

import mathimport randomfrom typing import List, Tuple
class BairdsCounterexample:    """Simulates Baird's 7-state counterexample to benchmark off-policy stability."""
    def __init__(self, gamma: float = 0.99) -> None:        self.gamma: float = gamma        self.num_states: int = 7        self.feature_dim: int = 8
    def get_features(self, state: int) -> List[float]:        """Construct Baird's linear feature representation (d=8)."""        x = [0.0] * self.feature_dim        if state < 6:            x[state] = 2.0            x[7] = 1.0        else:            x[6] = 1.0            x[7] = 2.0        return x
    @staticmethod    def dot(v1: List[float], v2: List[float]) -> float:        return sum(a * b for a, b in zip(v1, v2))
    @staticmethod    def l2_norm(v: List[float]) -> float:        return math.sqrt(sum(a * a for a in v))
    def run_standard_td(self, steps: int = 1000, alpha: float = 0.01, seed: int = 42) -> float:        """Run standard off-policy semi-gradient TD(0)."""        random.seed(seed)        weights = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 10.0, 1.0]        s = random.randint(0, 6)
        for _ in range(steps):            # Behavior policy: action 0 (dotted) w.p. 6/7, action 1 (solid) w.p. 1/7            a = 0 if random.random() < 6.0 / 7.0 else 1            # Target policy always takes action 1 (solid)            rho = 0.0 if a == 0 else 7.0            s_next = random.randint(0, 5) if a == 0 else 6
            x_s = self.get_features(s)            x_next = self.get_features(s_next)            delta = 0.0 + self.gamma * self.dot(x_next, weights) - self.dot(x_s, weights)
            for i in range(self.feature_dim):                weights[i] += alpha * rho * delta * x_s[i]            s = s_next
        return self.l2_norm(weights)
    def run_emphatic_td(self, steps: int = 1000, alpha: float = 0.005, seed: int = 42) -> float:        """Run Emphatic TD(0) with followon trace and interest."""        random.seed(seed)        weights = [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 10.0, 1.0]        f_trace = 1.0        interest = 1.0        rho_prev = 0.0        s = random.randint(0, 6)
        for t in range(steps):            # Update recursive followon trace F_t            if t == 0:                f_trace = interest            else:                f_trace = self.gamma * rho_prev * f_trace + interest            emphasis = f_trace  # ETD(0) has lambda = 0 => M_t = F_t
            a = 0 if random.random() < 6.0 / 7.0 else 1            rho = 0.0 if a == 0 else 7.0            s_next = random.randint(0, 5) if a == 0 else 6
            x_s = self.get_features(s)            x_next = self.get_features(s_next)            delta = 0.0 + self.gamma * self.dot(x_next, weights) - self.dot(x_s, weights)
            # Emphatic update scaled by M_t            for i in range(self.feature_dim):                weights[i] += alpha * emphasis * rho * delta * x_s[i]
            s = s_next            rho_prev = rho
        return self.l2_norm(weights)
if __name__ == "__main__":    env = BairdsCounterexample(gamma=0.99)    steps = 1000
    td_norm = env.run_standard_td(steps=steps, alpha=0.01)    etd_norm = env.run_emphatic_td(steps=steps, alpha=0.005)
    print(f"Standard TD(0) weight norm after {steps} steps: {td_norm:.2f}")    print(f"Emphatic TD(0) weight norm after {steps} steps: {etd_norm:.2f}")
    # Automated assertions    assert td_norm > 200.0, "Standard TD must diverge exponentially on Baird's domain"    assert etd_norm < td_norm, "Emphatic TD must remain substantially more stable than standard TD"
# -> expected output:Standard TD(0) weight norm after 1000 steps: 493.08Emphatic TD(0) weight norm after 1000 steps: 49.87

Watch Out For

Runaway Variance in Long Off-Policy Chains

The primary trade-off of Emphatic-TD's single-timescale stability is high variance in the emphasis scalar MtM_t.

Because the followon trace compounds products of importance sampling ratios:

Ft=It+γρt−1It−1+γ2ρt−1ρt−2It−2+…F_t = I_t + \gamma \rho_{t-1} I_{t-1} + \gamma^2 \rho_{t-1} \rho_{t-2} I_{t-2} + \dots

if the behavior policy frequently explores actions that have low probability under bb but high probability under π\pi, individual ratios ρk≫1\rho_k \gg 1. Consecutive large ratios cause FtF_t (and therefore MtM_t) to spike exponentially. When MtM_t reaches values in the hundreds or thousands, the effective step size αeff=αMt\alpha_{\text{eff}} = \alpha M_t surges uncontrollably, resulting in severe weight shocks or numerical floating-point overflows.

The Fix:

  1. Use multi-step eligibility traces (λ>0\lambda > 0): Setting λ∈[0.8,0.95]\lambda \in [0.8, 0.95] dampens trace variance by blending followon accumulation with instantaneous interest: Mt=λIt+(1−λ)FtM_t = \lambda I_t + (1 - \lambda) F_t.
  2. Smaller baseline learning rate: Scale down the initial step size α\alpha to account for the expected magnitude of MtM_t.
  3. Emphasis clipping and trace truncation: In deep RL implementations, clip Mt≤Mmax⁡M_t \le M_{\max} or apply soft-thresholding to prevent solitary exploratory episodes from corrupting network weights.
  4. True Online ETD(λ\lambda): Implement the exact forward-view equivalence to minimize per-step variance accumulation along sample paths.

The Quick Version

  • Single-Timescale Stability: Emphatic-TD solves the Deadly Triad without requiring the secondary weight vector or dual learning rates demanded by Gradient-TD (TDC/GTD2).
  • The Followon Trace: Ft=γρt−1Ft−1+ItF_t = \gamma \rho_{t-1} F_{t-1} + I_t tracks how states pass downstream bootstrapping debts to future states, ensuring updates reflect inherited prediction value.
  • Contraction Restored: The asymptotic emphasis distribution guarantees that the expected transition matrix Aemphatic=X⊤Demphatic(I−γPπ)X\mathbf{A}_{\text{emphatic}} = \mathbf{X}^\top \mathbf{D}_{\text{emphatic}} (\mathbf{I} - \gamma \mathbf{P}_\pi) \mathbf{X} is strictly positive definite.
  • The Core Trade-off: ETD trades away the dual-timescale complexity of Gradient-TD in exchange for sample variance in the scalar emphasis MtM_t, requiring trace damping (λ>0\lambda > 0) or step-size tuning in practice.