Skip to content
AI360Xpert
Beta

Policy in Reinforcement Learning

A policy is an agent's decision-making strategy, defining how it selects actions in response to observed states to maximize long-term cumulative reward.

A reinforcement learning policy maps environmental states to action decisions, either directly via deterministic functions or as probability distributions across possible actions.
A reinforcement learning policy maps environmental states to action decisions, either directly via deterministic functions or as probability distributions across possible actions.

Why Does This Exist?

In reinforcement learning, an agent does not follow hardcoded rules or step-by-step procedures. Instead, it must decide what action to take in response to whatever situation it encounters. The policy (π\pi) is the mathematical formalization of that decision mechanism: the mapping from perceived states of the environment to actions.

Without a formally defined policy, reinforcement learning collapses:

  • No distinction between evaluation and execution: In value-based methods, an agent evaluates states using value functions (V(s)V(s) or Q(s,a)Q(s, a)), but an evaluation is not an action. A value function tells the agent how good a situation is, whereas a policy specifies what to actually do.
  • No framework for exploration: If an agent only possessed fixed actions, it could never systematically balance exploring untried behaviors against exploiting known successes. A stochastic policy allows controlled randomness that can be tuned as learning progresses.
  • No target for policy optimization: Direct policy methods like Policy Gradient Methods optimize the parameters of a policy function directly using gradient ascent. Without a formal policy representation, policy gradient algorithms cannot be derived.

A policy is the central entity in Markov Decision Processes. While models predict how the environment responds and value functions predict the accumulated return, the policy is the actor itself.

Think of It Like This

The NFL quarterback's situational playbook

Imagine an NFL quarterback walking up to the line of scrimmage before the snap. The defensive formation he sees—the depth of the safeties, the spacing of the cornerbacks, the presence of blitzers on the edge—represents the current state ss. The quarterback's playbook is his policy.

In a rigid, deterministic offensive scheme, every defensive look dictates one exact audible: if the defense shows a Cover-0 all-out blitz, the quarterback must check to a quick slant. There is no randomness or deliberation; one specific observed state triggers one exact action: a=μ(s)a = \mu(s).

In a modern read-and-react scheme, the policy is stochastic: the quarterback recognizes the defensive formation and assigns situational probabilities to different plays. Against a balanced Cover-2 zone, he might assign an 80% likelihood to handing the ball off to the running back, a 15% likelihood to a short pass over the middle, and a 5% likelihood to a deep shot down the sideline. Mixing plays prevents the opposing defensive coordinator from anticipating the call before the snap.

Where the analogy stops: A quarterback learns his playbook from coaches, film study, and prior football experience before game day. A reinforcement learning agent begins with zero football knowledge, initial random weights, and no playbook. It constructs its policy entirely through trial-and-error interaction with the environment, guided purely by reward signals.

How It Actually Works

Deterministic vs. stochastic policy formulations

A policy can take two fundamental mathematical forms depending on whether action selection is certain or probabilistic.

Deterministic Policy:    State s ──────► [ μ(s) ] ──────► Single Action aStochastic Policy:       State s ──────► [ π(a|s) ] ─────► Probability Distribution P(A=a|S=s)

1. Deterministic policies (μ\mu)

A deterministic policy is a direct mapping from the state space S\mathcal{S} to the action space A\mathcal{A}:

μ:S→A,a=μ(s)\mu: \mathcal{S} \to \mathcal{A}, \quad a = \mu(s)

Given state ss, the agent always selects the identical action aa. Deterministic policies are common in continuous control tasks (such as robotic joint torque control in algorithms like DDPG and TD3) and during test-time evaluation when exploration is disabled.

2. Stochastic policies (π\pi)

A stochastic policy defines a conditional probability distribution over actions given the current state:

π:S×A→[0,1],π(a∣s)=P(At=a∣St=s)\pi: \mathcal{S} \times \mathcal{A} \to [0, 1], \quad \pi(a \mid s) = \mathbb{P}(A_t = a \mid S_t = s)

For any valid stochastic policy over a discrete action space, two axioms must hold:

  1. Non-negativity: π(a∣s)≥0\pi(a \mid s) \ge 0 for all a∈Aa \in \mathcal{A} and s∈Ss \in \mathcal{S}.
  2. Total probability: ∑a∈Aπ(a∣s)=1\sum_{a \in \mathcal{A}} \pi(a \mid s) = 1 for all s∈Ss \in \mathcal{S}.

In continuous action spaces where A⊆Rd\mathcal{A} \subseteq \mathbb{R}^d, π(a∣s)\pi(a \mid s) is a probability density function satisfying ∫Aπ(a∣s) da=1\int_{\mathcal{A}} \pi(a \mid s)\, da = 1. It is commonly parameterized as a Gaussian distribution N(μθ(s),Σθ(s))\mathcal{N}(\mu_\theta(s), \Sigma_\theta(s)) whose mean and covariance are generated by neural networks.

3. Deriving policies from value functions

In value-based reinforcement learning, agents learn state-action values Q(s,a)Q(s, a) and derive policies from them:

  • Greedy policy: Selects the action with the highest estimated value: πgreedy(s)=arg⁡max⁡a∈AQ(s,a)\pi_{\text{greedy}}(s) = \arg\max_{a \in \mathcal{A}} Q(s, a)

  • ϵ\epsilon-greedy policy: Balances exploitation with uniform random exploration: π(a∣s)={1−ϵ+ϵ∣A∣if a=arg⁡max⁡a′Q(s,a′)ϵ∣A∣otherwise\pi(a \mid s) = \begin{cases} 1 - \epsilon + \frac{\epsilon}{|\mathcal{A}|} & \text{if } a = \arg\max_{a'} Q(s, a') \\ \frac{\epsilon}{|\mathcal{A}|} & \text{otherwise} \end{cases}

  • Softmax (Boltzmann) policy: Converts action-values into probabilities proportional to their expected return using a temperature parameter τ>0\tau > 0: π(a∣s)=exp⁡(Q(s,a)/τ)∑b∈Aexp⁡(Q(s,b)/τ)\pi(a \mid s) = \frac{\exp(Q(s, a) / \tau)}{\sum_{b \in \mathcal{A}} \exp(Q(s, b) / \tau)}

When τ→∞\tau \to \infty, the distribution becomes uniform random (maximum exploration). When τ→0+\tau \to 0^+, the distribution collapses to the deterministic greedy action (pure exploitation).

4. The optimal policy (π∗\pi^*)

A policy π∗\pi^* is defined as optimal if its expected return from any state is greater than or equal to that of any other policy:

Vπ∗(s)≥Vπ(s)∀s∈S,  ∀πV^{\pi^*}(s) \ge V^\pi(s) \quad \forall s \in \mathcal{S}, \; \forall \pi

A foundational theorem of Markov Decision Processes states that for any MDP, there exists at least one stationary, deterministic policy that is globally optimal (π∗=μ∗\pi^* = \mu^*). Even though optimal policies in fully observable MDPs can be deterministic, stochastic policies remain essential during training to drive exploration and discover that optimum.

Worked numerical example

Consider an autonomous vehicle approaching a signalized intersection where the traffic light turns yellow.

  • Observed state ss: Distance to stop line d=30 md = 30\text{ m}, velocity v=45 km/hv = 45\text{ km/h}, yellow duration elapsed t=1.2 st = 1.2\text{ s}.
  • Available discrete actions:
    • a1a_1: Maintain current speed
    • a2a_2: Brake firmly
    • a3a_3: Accelerate through the yellow light
  • Learned action values (Q(s,a)Q(s, a)):
    • Q(s,a1)=2.40Q(s, a_1) = 2.40
    • Q(s,a2)=1.10Q(s, a_2) = 1.10
    • Q(s,a3)=3.20Q(s, a_3) = 3.20

Step 1: Deterministic greedy action selection

Under a deterministic greedy policy μ(s)=arg⁡max⁡aQ(s,a)\mu(s) = \arg\max_a Q(s, a):

μ(s)=arg⁡max⁡{Q(s,a1)=2.40,  Q(s,a2)=1.10,  Q(s,a3)=3.20}=a3\mu(s) = \arg\max \{Q(s, a_1)=2.40, \; Q(s, a_2)=1.10, \; Q(s, a_3)=3.20\} = a_3

The vehicle deterministically selects a3a_3 ("Accelerate").

Step 2: Stochastic Softmax policy calculation (τ=1.0\tau = 1.0)

To enable exploration, we compute the Boltzmann distribution at temperature τ=1.0\tau = 1.0:

  1. Compute exponents: exp⁡(Q(s,a1)τ)=exp⁡(2.40)≈11.0232\exp\left(\frac{Q(s, a_1)}{\tau}\right) = \exp(2.40) \approx 11.0232 exp⁡(Q(s,a2)τ)=exp⁡(1.10)≈3.0042\exp\left(\frac{Q(s, a_2)}{\tau}\right) = \exp(1.10) \approx 3.0042 exp⁡(Q(s,a3)τ)=exp⁡(3.20)≈24.5325\exp\left(\frac{Q(s, a_3)}{\tau}\right) = \exp(3.20) \approx 24.5325

  2. Sum the exponents to obtain the partition function ZZ: Z=11.0232+3.0042+24.5325=38.5599Z = 11.0232 + 3.0042 + 24.5325 = 38.5599

  3. Calculate normalized action probabilities π(a∣s)\pi(a \mid s): π(a1∣s)=11.023238.5599≈0.2859(28.59%)\pi(a_1 \mid s) = \frac{11.0232}{38.5599} \approx 0.2859 \quad (28.59\%) π(a2∣s)=3.004238.5599≈0.0779(7.79%)\pi(a_2 \mid s) = \frac{3.0042}{38.5599} \approx 0.0779 \quad (7.79\%) π(a3∣s)=24.532538.5599≈0.6362(63.62%)\pi(a_3 \mid s) = \frac{24.5325}{38.5599} \approx 0.6362 \quad (63.62\%)

  4. Verification of the total probability axiom: 0.2859+0.0779+0.6362=1.0000(100.00%)0.2859 + 0.0779 + 0.6362 = 1.0000 \quad (100.00\%)

Step 3: Temperature sensitivity comparison

The temperature parameter τ\tau modulates the agent's exploration profile:

ActionQ(s,a)Q(s, a)Cool (τ=0.3\tau = 0.3)Standard (τ=1.0\tau = 1.0)Hot (τ=4.0\tau = 4.0)
a1a_1: Maintain Speed2.406.48%6.48\%28.59%28.59\%32.22%32.22\%
a2a_2: Brake firmly1.100.06%0.06\%7.79%7.79\%23.33%23.33\%
a3a_3: Accelerate3.2093.46%93.46\%63.62%63.62\%44.45%44.45\%

As τ→0\tau \to 0, action a3a_3 dominates (93.46%→100%93.46\% \to 100\%). As τ\tau increases, the probability mass flattens toward a uniform distribution (≈33.33%\approx 33.33\% each), encouraging active exploration of lower-value actions.

Code

from typing import List, Sequence, Tupleimport numpy as np
class PolicyEvaluator:    """Demonstrates deterministic greedy and stochastic softmax policies."""
    def __init__(self, action_names: Sequence[str]) -> None:        self.action_names = list(action_names)        self.num_actions = len(action_names)
    def deterministic_greedy(self, q_values: np.ndarray) -> Tuple[int, str]:        """Selects the action with the maximum Q-value.
        Args:            q_values: 1D array of shape (num_actions,) containing expected returns.
        Returns:            Tuple of (action_index, action_name).        """        action_idx = int(np.argmax(q_values))        return action_idx, self.action_names[action_idx]
    def softmax_probabilities(        self, q_values: np.ndarray, temperature: float = 1.0    ) -> np.ndarray:        """Computes Boltzmann action probabilities with numerical stability.
        Args:            q_values: 1D array of action values.            temperature: Exploration parameter tau > 0.
        Returns:            Normalized 1D probability distribution summing to 1.0.        """        if temperature <= 0:            raise ValueError("Temperature tau must be strictly positive.")
        # Subtract max for numerical stability to prevent float overflow        scaled_q = (q_values - np.max(q_values)) / temperature        exp_q = np.exp(scaled_q)        probabilities = exp_q / np.sum(exp_q)        return probabilities
    def sample_stochastic(        self, probabilities: np.ndarray, rng: np.random.Generator    ) -> Tuple[int, str]:        """Samples an action according to the probability distribution pi(a|s)."""        action_idx = int(rng.choice(self.num_actions, p=probabilities))        return action_idx, self.action_names[action_idx]

# Define intersection yellow light scenarioactions = ["Maintain Speed", "Brake firmly", "Accelerate"]q_values = np.array([2.40, 1.10, 3.20], dtype=np.float64)
evaluator = PolicyEvaluator(actions)rng = np.random.default_rng(seed=42)
# 1. Evaluate deterministic policydet_idx, det_name = evaluator.deterministic_greedy(q_values)print(f"Deterministic Greedy Policy:")print(f"  Selected: {det_name} (Index {det_idx}) with Q = {q_values[det_idx]:.2f}\n")
# 2. Evaluate stochastic Softmax policy at tau = 1.0probs_standard = evaluator.softmax_probabilities(q_values, temperature=1.0)print("Stochastic Policy Probabilities (tau = 1.0):")for name, prob in zip(actions, probs_standard):    print(f"  {name:<15}: {prob * 100:6.2f}%")print(f"  Total Probability Sum: {np.sum(probs_standard):.4f}\n")
# 3. Simulate 10,000 empirical action samplesnum_samples = 10_000samples = [evaluator.sample_stochastic(probs_standard, rng)[0] for _ in range(num_samples)]counts = np.bincount(samples, minlength=len(actions))
print(f"Empirical Frequencies from {num_samples:,} Samples:")for name, count in zip(actions, counts):    empirical_prob = (count / num_samples) * 100    print(f"  {name:<15}: {count:5d} draws ({empirical_prob:5.2f}%)")
Deterministic Greedy Policy:  Selected: Accelerate (Index 2) with Q = 3.20
Stochastic Policy Probabilities (tau = 1.0):  Maintain Speed :  28.59%  Brake firmly   :   7.79%  Accelerate     :  63.62%  Total Probability Sum: 1.0000
Empirical Frequencies from 10,000 Samples:  Maintain Speed :  2855 draws (28.55%)  Brake firmly   :   781 draws ( 7.81%)  Accelerate     :  6364 draws (63.64%)

Watch Out For

Deploying deterministic policies in partially observable or competitive settings

While fully observable Markov decision processes guarantee that an optimal deterministic policy exists, real-world tasks often violate full observability (POMDPs) or involve adversarial agents (game theory).

In Rock-Paper-Scissors or Poker, a deterministic policy is easily exploited: an opponent observing that an agent always plays "Rock" will counter with "Paper" 100% of the time, resulting in guaranteed defeat. Similarly, in a robot navigating a symmetrical corridor with limited sensors, a deterministic policy can cause the robot to oscillate indefinitely between identical states.

The fix: For competitive and partially observable environments, train and maintain a stochastic policy that can represent mixed strategies (such as playing Rock, Paper, and Scissors with equal 1/31/3 probabilities), or augment the state representation with an recurrent memory (such as an LSTM or Transformer) to track historical state sequences.

Premature policy collapse during training

In parameterized stochastic policies, an early stroke of luck can cause the agent to discover a modest local reward. If gradients aggressively boost the logits for that action, the Softmax probabilities can rapidly sharpen toward 1.0 for the favored action and 0.0 for all others.

Once π(aalternative∣s)≈0\pi(a_{\text{alternative}} \mid s) \approx 0, the agent stops sampling alternative actions entirely. As a result, it never collects the transition data needed to compute gradient updates for those actions, permanently trapping the policy in a suboptimal local maximum.

The fix: Incorporate an entropy regularization bonus into the policy optimization objective:

J(θ)=E[∑tγtRt]+βH(πθ(⋅∣St))J(\theta) = \mathbb{E}\left[\sum_{t} \gamma^t R_t\right] + \beta \mathcal{H}(\pi_\theta(\cdot \mid S_t))

where H(π)=−∑aπ(a∣s)log⁡π(a∣s)\mathcal{H}(\pi) = -\sum_a \pi(a \mid s) \log \pi(a \mid s) measures the entropy of the policy and β>0\beta > 0 penalizes premature certainty. Algorithms like Soft Actor-Critic (SAC) and Proximal Policy Optimization (PPO) use entropy bonuses to maintain healthy exploration throughout training.

The Quick Version

  • A policy (π\pi) is the decision-making rule of an RL agent, mapping environmental states to actions to maximize cumulative reward.
  • Deterministic policies (a=μ(s)a = \mu(s)) output a single designated action per state; stochastic policies (π(a∣s)\pi(a \mid s)) output a normalized probability distribution across the action space.
  • While fully observable MDPs guarantee the existence of an optimal deterministic policy (π∗\pi^*), stochastic policies are required during training to drive exploration and in POMDPs/games to avoid exploitable predictability.
  • Value-based agents derive policies using greedy selection (arg⁡max⁡aQ(s,a)\arg\max_a Q(s, a)), ϵ\epsilon-greedy exploration, or Boltzmann/Softmax distributions controlled by a temperature parameter τ\tau.