Skip to content
AI360Xpert
Beta

Q-Learning vs SARSA

Q-Learning learns the value of the optimal action regardless of what the agent actually does, while SARSA learns the value of the policy the agent actually executes, accounting for its own exploratory mistakes.

Backup diagrams contrasting Q-learning's off-policy max-operator target with SARSA's on-policy executed-action target.
Backup diagrams contrasting Q-learning's off-policy max-operator target with SARSA's on-policy executed-action target.

Why Does This Exist?

Temporal difference (TD) control methods learn action values directly from raw experience without requiring an environment transition model. However, an essential architectural fork emerges when calculating the TD target: should the update assume the agent will pick the theoretically optimal action on the next step, or should it account for the action the agent's current exploratory policy actually executes?

This distinction separates Q-learning (off-policy) from SARSA (on-policy). The famous "Cliff Walking" benchmark illustrates why this matters: an agent walking along the edge of a lethal cliff can reach the goal in fewer steps along the brink. Q-learning learns the value of this optimal edge path because its backup assumes zero exploratory errors (max⁡a′Q\max_{a'} Q). But during actual training under an ϵ\epsilon-greedy policy, the agent occasionally slips and falls off the cliff, incurring catastrophic negative rewards.

SARSA, in contrast, evaluates the policy being run—including the fact that the agent explores randomly with probability ϵ\epsilon. Because SARSA backs up the value of the action actually taken (Q(S′,A′)Q(S', A')), it recognizes that walking near the cliff is hazardous under imperfect execution and learns a safer, longer detour.

Think of It Like This

A race car driver in practice vs a defensive commuter

Consider two drivers learning routes through icy mountain roads.

The Q-learning driver calculates travel times under the assumption of absolute perfection: "If I take this sharp hairpin turn at 90 mph, I will shave off three minutes." The Q-learning math backs up the value of the absolute optimal move (max⁡\max), completely ignoring that the driver is still practicing and has a 10% chance of twitching the steering wheel and careening into the ravine.

The SARSA driver plans realistically: "I have a 10% chance of making an erratic move while practicing. If I drive along the canyon edge, that 10% mistake is fatal." Because SARSA backs up the move the driver actually attempts next, it penalizes the canyon road and steers toward the wider, salted highway. Q-learning learns how the track could be driven by an infallible expert; SARSA learns how the track is currently being driven by the student.

How It Actually Works

On-Policy vs Off-Policy Target Backup Equations

Both algorithms operate on tabular or function-approximated action values Q(s,a)Q(s, a), updating parameters using a 1-step temporal difference error.

Q-Learning (Watkins, 1989)

Q-learning is an off-policy TD control algorithm. The target policy is greedy with respect to current Q-values, while the behavior policy generating transitions is typically ϵ\epsilon-greedy.

The Q-learning update rule is:

Q(St,At)←Q(St,At)+α[Rt+1+γmax⁡a∈AQ(St+1,a)−Q(St,At)]Q(S_t, A_t) \leftarrow Q(S_t, A_t) + \alpha \left[ R_{t+1} + \gamma \max_{a \in \mathcal{A}} Q(S_{t+1}, a) - Q(S_t, A_t) \right]

Notice that the next action selected in the environment, At+1A_{t+1}, plays no role in the target calculation. The target strictly evaluates the maximum available Q-value in St+1S_{t+1}.

SARSA (Rummery & Niranjan, 1994)

SARSA takes its name from the complete 5-tuple sequence of experience it uses: (St,At,Rt+1,St+1,At+1)(S_t, A_t, R_{t+1}, S_{t+1}, A_{t+1}). It is an on-policy TD control algorithm: the policy evaluated is the exact policy used to generate actions.

The SARSA update rule is:

Q(St,At)←Q(St,At)+α[Rt+1+γQ(St+1,At+1)−Q(St,At)]Q(S_t, A_t) \leftarrow Q(S_t, A_t) + \alpha \left[ R_{t+1} + \gamma Q(S_{t+1}, A_{t+1}) - Q(S_t, A_t) \right]

Here, At+1A_{t+1} is sampled from the behavior policy π(St+1)\pi(S_{t+1}) (such as ϵ\epsilon-greedy). If ϵ\epsilon-greedy selects an exploratory, suboptimal action, SARSA propagates that suboptimal value directly into Q(St,At)Q(S_t, A_t).

Worked Example

An agent is in state S0S_0 and takes action AedgeA_{\text{edge}}. It transitions to state S1S_1 and receives immediate reward R=−1.0R = -1.0.

In state S1S_1, there are two actions:

  • a1a_1 (Stay on path): Q(S1,a1)=10.0Q(S_1, a_1) = 10.0
  • a2a_2 (Fall off cliff): Q(S1,a2)=−100.0Q(S_1, a_2) = -100.0

Let learning rate α=0.5\alpha = 0.5, discount factor γ=0.9\gamma = 0.9, and initial Q(S0,Aedge)=0.0Q(S_0, A_{\text{edge}}) = 0.0. During exploration, the behavior policy samples an exploratory mistake, selecting At+1=a2A_{t+1} = a_2 (falling off the cliff).

Q-Learning Update:

  1. Target computes max⁡\max: max⁡a′Q(S1,a′)=max⁡(10.0,−100.0)=10.0\max_{a'} Q(S_1, a') = \max(10.0, -100.0) = 10.0
  2. TD Target: yQ=−1.0+0.9×10.0=−1.0+9.0=8.0y_{\text{Q}} = -1.0 + 0.9 \times 10.0 = -1.0 + 9.0 = 8.0
  3. Value update: Q(S0,Aedge)←0.0+0.5×(8.0−0.0)=4.0Q(S_0, A_{\text{edge}}) \leftarrow 0.0 + 0.5 \times (8.0 - 0.0) = 4.0

SARSA Update:

  1. Target uses the executed action At+1=a2A_{t+1} = a_2: Q(S1,At+1)=Q(S1,a2)=−100.0Q(S_1, A_{t+1}) = Q(S_1, a_2) = -100.0
  2. TD Target: ySARSA=−1.0+0.9×(−100.0)=−1.0−90.0=−91.0y_{\text{SARSA}} = -1.0 + 0.9 \times (-100.0) = -1.0 - 90.0 = -91.0
  3. Value update: Q(S0,Aedge)←0.0+0.5×(−91.0−0.0)=−45.5Q(S_0, A_{\text{edge}}) \leftarrow 0.0 + 0.5 \times (-91.0 - 0.0) = -45.5

Q-learning rewarded taking AedgeA_{\text{edge}} (+4.0+4.0) because it assumed optimal future play. SARSA severely penalized AedgeA_{\text{edge}} (−45.5-45.5) because it observed that the exploratory policy actually chose the fatal action a2a_2.

Code

from typing import Dict, Tupleimport numpy as np
def step_q_learning(    q_table: Dict[str, np.ndarray],    s: str,    a_idx: int,    r: float,    s_next: str,    alpha: float = 0.5,    gamma: float = 0.9,) -> float:    """Updates Q-table using Q-Learning off-policy target."""    best_next_q = float(np.max(q_table[s_next]))    target = r + gamma * best_next_q    td_error = target - q_table[s][a_idx]    q_table[s][a_idx] += alpha * td_error    return q_table[s][a_idx]
def step_sarsa(    q_table: Dict[str, np.ndarray],    s: str,    a_idx: int,    r: float,    s_next: str,    a_next_idx: int,    alpha: float = 0.5,    gamma: float = 0.9,) -> float:    """Updates Q-table using SARSA on-policy target."""    actual_next_q = float(q_table[s_next][a_next_idx])    target = r + gamma * actual_next_q    td_error = target - q_table[s][a_idx]    q_table[s][a_idx] += alpha * td_error    return q_table[s][a_idx]
# Identical initial environments: S1 has [Stay=10.0, Fall=-100.0]q_ql = {"S0": np.array([0.0]), "S1": np.array([10.0, -100.0])}q_sarsa = {"S0": np.array([0.0]), "S1": np.array([10.0, -100.0])}
# The agent explores and accidentally executes action index 1 (Fall)res_ql = step_q_learning(q_ql, s="S0", a_idx=0, r=-1.0, s_next="S1")res_sarsa = step_sarsa(q_sarsa, s="S0", a_idx=0, r=-1.0, s_next="S1", a_next_idx=1)
print(f"Q-Learning Q(S0, edge): {res_ql:.1f}")# -> Q-Learning Q(S0, edge): 4.0
print(f"SARSA Q(S0, edge): {res_sarsa:.1f}")# -> SARSA Q(S0, edge): -45.5

Watch Out For

Assuming Off-Policy Optimality Always Yields Superior Online Performance

Practitioners frequently default to Q-learning assuming that learning the optimal policy Q∗Q^* is universally superior to learning the behavioral value QπQ^\pi. In simulation or offline replay buffers, off-policy methods are indeed more data-efficient. However, in real-world systems (robotics, autonomous navigation, industrial control) where learning happens online, Q-learning's indifference to its own exploration noise can cause real physical damage.

If an agent must train live in an environment where mistakes incur large penalties, use SARSA or decay the exploration parameter ϵ\epsilon aggressively. SARSA safely optimizes performance subject to exploration noise, converging to Q∗Q^* asymptotically only as ϵ→0\epsilon \to 0.

The Quick Version

  • Q-learning is off-policy: it targets R+γmax⁡a′Q(S′,a′)R + \gamma \max_{a'} Q(S', a'), learning optimal values while executing an exploratory policy.
  • SARSA is on-policy: it targets R+γQ(S′,A′)R + \gamma Q(S', A'), learning the value of the exploratory policy actually being executed.
  • In hazardous environments with online exploration, Q-learning learns risky optimal paths, while SARSA learns safer paths that buffer against exploration mistakes.