Skip to content
AI360Xpert
Beta

Policy Objective Functions

Policy gradient methods define scalar objective functions measuring policy performance across episodic, average value, and continuous average reward settings.

Reinforcement learning formalizes policy optimization through three canonical objectives—episodic start-state, continuing average value, and continuous average reward—which unify under the Policy Gradient Theorem.
Reinforcement learning formalizes policy optimization through three canonical objectives—episodic start-state, continuing average value, and continuous average reward—which unify under the Policy Gradient Theorem.

Why Does This Exist?

In value-based reinforcement learning (such as Q-learning or DQN), the learning goal is implicit: minimize Bellman error until the action-value function converges to Q∗Q^*.

In policy-based methods, however, the agent directly optimizes a parameterized policy πθ(a∣s)\pi_\theta(a \mid s) via gradient ascent:

θt+1=θt+α∇θJ(θ)\theta_{t+1} = \theta_t + \alpha \nabla_\theta J(\theta)

To compute that gradient, we must define the scalar performance objective J(θ)J(\theta) being maximized. In real-world reinforcement learning, problems exhibit fundamentally different temporal horizons:

  1. Episodic tasks that naturally terminate at a terminal goal or timeout (games, robotic grasping).
  2. Continuing discounted tasks where an ongoing process is evaluated across stationary state visitations with time-preference discounting (server load balancing).
  3. Continuing undiscounted tasks where an infinite-horizon process runs indefinitely and success is measured purely by steady-state throughput (chemical refining, continuous industrial heating, high-frequency trading).

Selecting the wrong objective distorts learning. Applying discounted objectives to continuing problems introduces artificial horizon truncation, while episodic objectives fail on open-ended streams.

Furthermore, defining J(θ)J(\theta) appears to create a major theoretical dilemma: J(θ)J(\theta) depends on θ\theta both through action selection probabilities πθ(a∣s)\pi_\theta(a \mid s) and through the resulting state distribution dπθ(s)d^{\pi_\theta}(s). Without a special mathematical structure, computing ∇θJ(θ)\nabla_\theta J(\theta) would require differentiating the environmental state transition dynamics ∇θdπθ(s)\nabla_\theta d^{\pi_\theta}(s)—an intractable model-dependent quantity. Understanding policy objective functions reveals the central miracle of policy gradients: across all three formulations, ∇θdπθ(s)\nabla_\theta d^{\pi_\theta}(s) cancels out entirely.

Think of It Like This

Performance Metrics for an Investment Portfolio Manager

Imagine evaluating the executive performance of an investment portfolio manager whose algorithmic trading parameters θ\theta execute continuous financial allocations.

Depending on the fund's contractual charter, investors evaluate the manager using three distinct performance metrics:

  1. Metric 1: Final Maturity Payout (J1J_1 - Episodic Start-State Value): In a 5-year fixed-term venture fund starting with an initial capital deposit on Day 1 (s0s_0), the manager is evaluated strictly on total accumulated wealth upon liquidation at Year 5. All interim quarterly valuations only matter insofar as they maximize total discounted return from start state s0s_0.

  2. Metric 2: Average Portfolio Net Asset Value (JavVJ_{\text{avV}} - Continuing Average Value): In an ongoing wealth management trust with quarterly discounting, the trustee measures the expected account balance averaged across all quarters. The score weights bull, bear, and sideways market regimes by how frequently the portfolio naturally visits them (dπ(s)d^\pi(s)).

  3. Metric 3: Steady-State Dividend Yield (JavRJ_{\text{avR}} - Continuous Average Reward): In a perpetual endowment fund that never liquidates, investors care only about the steady-state cash flow generated per day (dividend yield rate r(π)r(\pi)). Immediate fluctuations in principal are irrelevant; the sole objective is maximizing the sustainable dividend rate per unit time.

The unifying insight: Even though the three accounting metrics measure distinct financial horizons, the trades the manager must execute (buying undervalued assets and selling overvalued assets) share the exact same directional gradient.

Where the analogy stops: Financial markets suffer from non-stationary macro shocks; in reinforcement learning, the environment transition dynamics P(s′∣s,a)P(s' \mid s, a) are stationary Markov Decision Processes governed by ergodic chains.

How It Actually Works

The Three Canonical Mathematical Formulations

1. Episodic Start-State Objective J1(θ)J_1(\theta)

In episodic environments with a designated start state s0s_0, the objective is the expected discounted cumulative return from s0s_0:

J1(θ)≐Vπθ(s0)=Eπθ[∑t=0TγtRt+1  |  S0=s0]J_1(\theta) \doteq V^{\pi_\theta}(s_0) = \mathbb{E}_{\pi_\theta} \left[ \sum_{t=0}^T \gamma^t R_{t+1} \;\middle|\; S_0 = s_0 \right]

If the initial state is drawn from a starting distribution d0(s)d_0(s), the objective generalizes to the expectation over starting states:

J1(θ)≐∑s∈Sd0(s)Vπθ(s)J_1(\theta) \doteq \sum_{s \in \mathcal{S}} d_0(s) V^{\pi_\theta}(s)

Under this formulation, the state visitation distribution is the unnormalized discounted on-policy state visitation frequency:

dπθ(s)≐∑t=0∞γtP(St=s∣S0=s0,πθ)d^{\pi_\theta}(s) \doteq \sum_{t=0}^\infty \gamma^t P(S_t = s \mid S_0 = s_0, \pi_\theta)

2. Continuing Average Value Objective JavV(θ)J_{\text{avV}}(\theta)

In continuing environments evaluated with discount factor γ∈[0,1)\gamma \in [0, 1), we measure the average state value weighted by the stationary distribution dπθ(s)d^{\pi_\theta}(s):

JavV(θ)≐∑s∈Sdπθ(s)Vπθ(s)J_{\text{avV}}(\theta) \doteq \sum_{s \in \mathcal{S}} d^{\pi_\theta}(s) V^{\pi_\theta}(s)

Where dπθ(s)≐lim⁡t→∞P(St=s∣S0,πθ)d^{\pi_\theta}(s) \doteq \lim_{t \to \infty} P(S_t = s \mid S_0, \pi_\theta) is the stationary distribution under policy πθ\pi_\theta, satisfying the balance equations:

dπθ(s)=∑s′∈Sdπθ(s′)∑a∈Aπθ(a∣s′)P(s∣s′,a)d^{\pi_\theta}(s) = \sum_{s' \in \mathcal{S}} d^{\pi_\theta}(s') \sum_{a \in \mathcal{A}} \pi_\theta(a \mid s') P(s \mid s', a)

3. Continuous Average Reward Objective JavR(θ)=r(πθ)J_{\text{avR}}(\theta) = r(\pi_\theta)

In continuing environments without discounting (γ=1\gamma = 1), accumulated returns diverge to infinity. The canonical metric is the average reward per time step (the reward rate r(π)r(\pi)):

JavR(θ)≐r(πθ)≐lim⁡T→∞1T∑t=1TE[Rt∣πθ]=∑s∈Sdπθ(s)∑a∈Aπθ(a∣s)RsaJ_{\text{avR}}(\theta) \doteq r(\pi_\theta) \doteq \lim_{T \to \infty} \frac{1}{T} \sum_{t=1}^T \mathbb{E}[R_t \mid \pi_\theta] = \sum_{s \in \mathcal{S}} d^{\pi_\theta}(s) \sum_{a \in \mathcal{A}} \pi_\theta(a \mid s) \mathcal{R}_s^a

Where Rsa≐E[Rt+1∣St=s,At=a]\mathcal{R}_s^a \doteq \mathbb{E}[R_{t+1} \mid S_t = s, A_t = a] is the expected immediate reward.


The Unifying Policy Gradient Theorem

Differentiating any of these objectives appears daunting because the state distribution dπθ(s)d^{\pi_\theta}(s) depends directly on policy parameters θ\theta:

∇θJ(θ)=∇θ(∑sdπθ(s)Vπθ(s))=∑s(∇θdπθ(s)Vπθ(s)+dπθ(s)∇θVπθ(s))\nabla_\theta J(\theta) = \nabla_\theta \left( \sum_s d^{\pi_\theta}(s) V^{\pi_\theta}(s) \right) = \sum_s \left( \nabla_\theta d^{\pi_\theta}(s) V^{\pi_\theta}(s) + d^{\pi_\theta}(s) \nabla_\theta V^{\pi_\theta}(s) \right)

Evaluating ∇θdπθ(s)\nabla_\theta d^{\pi_\theta}(s) would require knowing the environment's internal transition probabilities P(s′∣s,a)P(s' \mid s, a), making model-free learning impossible.

The Policy Gradient Theorem (Sutton et al., 1999) establishes that for all three objectives, the state distribution gradient miraculously vanishes:

∇θJ(θ)∝∑s∈Sdπθ(s)∑a∈A∇θπθ(a∣s)Qπθ(s,a)\nabla_\theta J(\theta) \propto \sum_{s \in \mathcal{S}} d^{\pi_\theta}(s) \sum_{a \in \mathcal{A}} \nabla_\theta \pi_\theta(a \mid s) Q^{\pi_\theta}(s, a)

Using the identity ∇θπθ(a∣s)=πθ(a∣s)∇θlog⁡πθ(a∣s)\nabla_\theta \pi_\theta(a \mid s) = \pi_\theta(a \mid s) \nabla_\theta \log \pi_\theta(a \mid s), this converts directly to an expectation over trajectory samples:

∇θJ(θ)=Eπθ[∇θlog⁡πθ(At∣St)Qπθ(St,At)]\nabla_\theta J(\theta) = \mathbb{E}_{\pi_\theta} \left[ \nabla_\theta \log \pi_\theta(A_t \mid S_t) Q^{\pi_\theta}(S_t, A_t) \right]

This single unifying result allows REINFORCE, Actor-Critic, PPO, and TRPO to optimize policies without modeling environmental dynamics.


Worked Numerical Example

Consider a 2-state cyclic MDP with states S={1,2}\mathcal{S} = \{1, 2\} and discount factor γ=0.8\gamma = 0.8:

  • State 1: The agent chooses between:
    • Action a0a_0 (Safe): Loops 1→11 \to 1 with reward R=1.0R = 1.0.
    • Action a1a_1 (Switch): Transitions 1→21 \to 2 with reward R=3.0R = 3.0.
  • State 2: Forced deterministic transition 2→12 \to 1 with reward R=0.0R = 0.0.
  • Policy parameter: Let p≐πθ(a1∣1)∈[0,1]p \doteq \pi_\theta(a_1 \mid 1) \in [0, 1] be the probability of taking the Switch action a1a_1 in state 1.

1. Stationary Distribution dπ=[d1,d2]d^\pi = [d_1, d_2]

Transition probability matrix: P=[1−pp1.00.0]P = \begin{bmatrix} 1 - p & p \\ 1.0 & 0.0 \end{bmatrix}

The stationary distribution satisfies d1(1−p)+d2(1.0)=d1  ⟹  d2=pd1d_1 (1 - p) + d_2 (1.0) = d_1 \implies d_2 = p d_1. Since d1+d2=1  ⟹  d1(1+p)=1d_1 + d_2 = 1 \implies d_1 (1 + p) = 1: d1=11+p,d2=p1+pd_1 = \frac{1}{1 + p}, \quad d_2 = \frac{p}{1 + p}

2. Continuous Average Reward Objective JavR(p)=r(π)J_{\text{avR}}(p) = r(\pi)

Expected reward in State 1: Rˉ1=(1−p)(1.0)+p(3.0)=1.0+2p\bar{R}_1 = (1 - p)(1.0) + p(3.0) = 1.0 + 2p. Expected reward in State 2: Rˉ2=0.0\bar{R}_2 = 0.0. r(π)=d1Rˉ1+d2Rˉ2=(11+p)(1+2p)+0=1+2p1+pr(\pi) = d_1 \bar{R}_1 + d_2 \bar{R}_2 = \left( \frac{1}{1 + p} \right)(1 + 2p) + 0 = \frac{1 + 2p}{1 + p}

3. Episodic Start-State Value J1(p)=V(1)J_1(p) = V(1) with γ=0.8\gamma = 0.8

From State 2: V(2)=0+γV(1)=0.8V(1)V(2) = 0 + \gamma V(1) = 0.8 V(1). From State 1: Q(1,a0)=1.0+γV(1)=1.0+0.8V(1)Q(1, a_0) = 1.0 + \gamma V(1) = 1.0 + 0.8 V(1) Q(1,a1)=3.0+γV(2)=3.0+0.8(0.8V(1))=3.0+0.64V(1)Q(1, a_1) = 3.0 + \gamma V(2) = 3.0 + 0.8 (0.8 V(1)) = 3.0 + 0.64 V(1) V(1)=(1−p)Q(1,a0)+pQ(1,a1)=(1+2p)+(0.8−0.16p)V(1)V(1) = (1 - p) Q(1, a_0) + p Q(1, a_1) = (1 + 2p) + (0.8 - 0.16p) V(1) V(1)[0.2+0.16p]=1+2p  ⟹  V(1)=1+2p0.2+0.16pV(1) [0.2 + 0.16p] = 1 + 2p \implies V(1) = \frac{1 + 2p}{0.2 + 0.16p}

4. Continuing Average Value Objective JavV(p)J_{\text{avV}}(p)

JavV(p)=d1V(1)+d2V(2)=(11+p+0.8p1+p)V(1)=(1+0.8p1+p)V(1)J_{\text{avV}}(p) = d_1 V(1) + d_2 V(2) = \left( \frac{1}{1 + p} + \frac{0.8p}{1 + p} \right) V(1) = \left( \frac{1 + 0.8p}{1 + p} \right) V(1)

Numerical Evaluation Across Policy Values:

Policy p=π(a1∣1)p = \pi(a_1 \mid 1)Stationary dπd^\piEpisodic J1(p)=V(1)J_1(p) = V(1)Average Value JavV(p)J_{\text{avV}}(p)Average Reward JavR(p)=r(π)J_{\text{avR}}(p) = r(\pi)
p=0.0p = 0.0 (Always Safe)[1.0000,0.0000][1.0000, 0.0000]5.00005.00005.00005.00001.00001.0000
p=0.5p = 0.5 (Balanced)[0.6667,0.3333][0.6667, 0.3333]7.14297.14296.66676.66671.33331.3333
p=1.0p = 1.0 (Always Switch)[0.5000,0.5000][0.5000, 0.5000]8.33338.33337.50007.50001.50001.5000

All three objective functions strictly increase as pp goes from 0→10 \to 1, confirming that all three formulations agree on the optimal policy (p∗=1.0p^* = 1.0) while operating under distinct mathematical units.

Code

from typing import Dict, Tuple

def evaluate_policy_objectives(    p: float,    gamma: float = 0.8,) -> Dict[str, float]:    """Compute analytical policy objective functions for the 2-state cyclic MDP.
    Args:        p: Probability of taking action a1 (Switch) in State 1.        gamma: Discount factor for episodic and average value objectives.
    Returns:        Dictionary containing stationary distribution and all 3 objective values.    """    if not 0.0 <= p <= 1.0:        raise ValueError("Policy parameter p must be in [0, 1].")
    # 1. Stationary Distribution: d1 = 1 / (1 + p), d2 = p / (1 + p)    d1 = 1.0 / (1.0 + p)    d2 = p / (1.0 + p)
    # 2. Continuous Average Reward Objective: J_avR = (1 + 2p) / (1 + p)    j_avr = (1.0 + 2.0 * p) / (1.0 + p)
    # 3. Episodic Start-State Objective: J_1 = V(1) = (1 + 2p) / (0.2 + 0.16p)    # Derived from Bellman system: V(2) = gamma * V(1)    denom = (1.0 - gamma) + p * gamma * (1.0 - gamma)    v1 = (1.0 + 2.0 * p) / (denom)    v2 = gamma * v1
    # 4. Continuing Average Value Objective: J_avV = d1 * V(1) + d2 * V(2)    j_avv = d1 * v1 + d2 * v2
    return {        "d1": d1,        "d2": d2,        "J_1": v1,        "J_avV": j_avv,        "J_avR": j_avr,    }

def simulate_empirical_average_reward(p: float, steps: int = 100_000) -> float:    """Empirical Monte Carlo rollout verifying continuous average reward r(pi)."""    import random
    random.seed(3)    state = 1    total_reward = 0.0
    for _ in range(steps):        if state == 1:            action = 1 if random.random() < p else 0            if action == 0:                reward = 1.0                state = 1            else:                reward = 3.0                state = 2        else:  # state == 2            reward = 0.0            state = 1        total_reward += reward
    return total_reward / steps

if __name__ == "__main__":    test_policies = [0.0, 0.5, 1.0]
    print("Policy (p) | d(s) = [d1, d2]   | J_1 (Start) | J_avV (Avg Val) | J_avR (Avg Reward)")    print("-" * 76)
    for p in test_policies:        res = evaluate_policy_objectives(p, gamma=0.8)        print(            f"  p = {p:3.1f}  | [{res['d1']:.4f}, {res['d2']:.4f}] |   {res['J_1']:7.4f}   |"            f"     {res['J_avV']:7.4f}     |     {res['J_avR']:7.4f}"        )
    # Validate empirical simulation against analytical average reward for p=0.5    empirical_r = simulate_empirical_average_reward(p=0.5, steps=100_000)    print(f"\nEmpirical Monte Carlo Average Reward (p=0.5): {empirical_r:.4f}")    print(f"Analytical Continuous Average Reward (p=0.5): 1.3333")
# Expected Output:# Policy (p) | d(s) = [d1, d2]   | J_1 (Start) | J_avV (Avg Val) | J_avR (Avg Reward)# ----------------------------------------------------------------------------#   p = 0.0  | [1.0000, 0.0000] |    5.0000   |      5.0000     |      1.0000#   p = 0.5  | [0.6667, 0.3333] |    7.1429   |      6.6667     |      1.3333#   p = 1.0  | [0.5000, 0.5000] |    8.3333   |      7.5000     |      1.5000# # Empirical Monte Carlo Average Reward (p=0.5): 1.3333# Analytical Continuous Average Reward (p=0.5): 1.3333

Watch Out For

The Continuing Task Discounting Paradox

A pervasive trap in reinforcement learning is applying the discounted episodic objective (J(θ)=E[∑t=0∞γtRt+1]J(\theta) = \mathbb{E}[\sum_{t=0}^\infty \gamma^t R_{t+1}]) to continuing, infinite-horizon tasks that lack an absorbing terminal state.

As Sutton & Barto (2018, Section 10.4) emphasize, using discounting in non-terminating tasks creates an artificial horizon bias:

  1. Time-Preference Distortion: The agent prioritizes immediate short-term payoffs over long-term sustainable throughput, even when the task runs forever.
  2. Distribution Mismatch: The discounted state distribution ∑tγtP(St=s∣s0)\sum_t \gamma^t P(S_t = s \mid s_0) depends heavily on the arbitrary choice of s0s_0 and does not match the true stationary distribution dπ(s)d^\pi(s).

The Fix: For true continuing, non-terminating tasks, use the Average Reward formulation (J(θ)=r(π)J(\theta) = r(\pi)) with undiscounted differential returns:

Gt≐∑k=0∞(Rt+k+1−r(π))G_t \doteq \sum_{k=0}^\infty \left( R_{t+k+1} - r(\pi) \right)

Differential returns measure performance relative to the steady-state baseline reward rate, eliminating artificial discounting bias while maintaining mathematical convergence.

The Quick Version

  • Three Canonical Metrics: Performance is formalized via Episodic Start-State Value (J1=V(s0)J_1 = V(s_0)), Continuing Average Value (JavV=∑dπ(s)V(s)J_{\text{avV}} = \sum d^\pi(s) V(s)), or Continuous Average Reward (JavR=r(π)J_{\text{avR}} = r(\pi)).
  • The Policy Gradient Miracle: The Policy Gradient Theorem proves that the gradient of the state distribution ∇θdπθ(s)\nabla_\theta d^{\pi_\theta}(s) cancels out across all three formulations, enabling model-free optimization.
  • Match Horizon to Objective: Use episodic objectives for tasks with designated start and termination states; use average reward formulations for continuing infinite-horizon industrial processes.
  • Avoid False Discounting: Discounting in continuing tasks introduces artificial horizon truncation; true non-terminating control requires differential returns anchored to the average reward rate.