Q-Learning vs SARSA
Q-Learning learns the value of the optimal action regardless of what the agent actually does, while SARSA learns the value of the policy the agent actually executes, accounting for its own exploratory mistakes.
Why Does This Exist?
Temporal difference (TD) control methods learn action values directly from raw experience without requiring an environment transition model. However, an essential architectural fork emerges when calculating the TD target: should the update assume the agent will pick the theoretically optimal action on the next step, or should it account for the action the agent's current exploratory policy actually executes?
This distinction separates Q-learning (off-policy) from SARSA (on-policy). The famous "Cliff Walking" benchmark illustrates why this matters: an agent walking along the edge of a lethal cliff can reach the goal in fewer steps along the brink. Q-learning learns the value of this optimal edge path because its backup assumes zero exploratory errors (). But during actual training under an -greedy policy, the agent occasionally slips and falls off the cliff, incurring catastrophic negative rewards.
SARSA, in contrast, evaluates the policy being run—including the fact that the agent explores randomly with probability . Because SARSA backs up the value of the action actually taken (), it recognizes that walking near the cliff is hazardous under imperfect execution and learns a safer, longer detour.
Think of It Like This
A race car driver in practice vs a defensive commuter
Consider two drivers learning routes through icy mountain roads.
The Q-learning driver calculates travel times under the assumption of absolute perfection: "If I take this sharp hairpin turn at 90 mph, I will shave off three minutes." The Q-learning math backs up the value of the absolute optimal move (), completely ignoring that the driver is still practicing and has a 10% chance of twitching the steering wheel and careening into the ravine.
The SARSA driver plans realistically: "I have a 10% chance of making an erratic move while practicing. If I drive along the canyon edge, that 10% mistake is fatal." Because SARSA backs up the move the driver actually attempts next, it penalizes the canyon road and steers toward the wider, salted highway. Q-learning learns how the track could be driven by an infallible expert; SARSA learns how the track is currently being driven by the student.
How It Actually Works
On-Policy vs Off-Policy Target Backup Equations
Both algorithms operate on tabular or function-approximated action values , updating parameters using a 1-step temporal difference error.
Q-Learning (Watkins, 1989)
Q-learning is an off-policy TD control algorithm. The target policy is greedy with respect to current Q-values, while the behavior policy generating transitions is typically -greedy.
The Q-learning update rule is:
Notice that the next action selected in the environment, , plays no role in the target calculation. The target strictly evaluates the maximum available Q-value in .
SARSA (Rummery & Niranjan, 1994)
SARSA takes its name from the complete 5-tuple sequence of experience it uses: . It is an on-policy TD control algorithm: the policy evaluated is the exact policy used to generate actions.
The SARSA update rule is:
Here, is sampled from the behavior policy (such as -greedy). If -greedy selects an exploratory, suboptimal action, SARSA propagates that suboptimal value directly into .
Worked Example
An agent is in state and takes action . It transitions to state and receives immediate reward .
In state , there are two actions:
- (Stay on path):
- (Fall off cliff):
Let learning rate , discount factor , and initial . During exploration, the behavior policy samples an exploratory mistake, selecting (falling off the cliff).
Q-Learning Update:
- Target computes :
- TD Target:
- Value update:
SARSA Update:
- Target uses the executed action :
- TD Target:
- Value update:
Q-learning rewarded taking () because it assumed optimal future play. SARSA severely penalized () because it observed that the exploratory policy actually chose the fatal action .
Code
from typing import Dict, Tupleimport numpy as np
def step_q_learning( q_table: Dict[str, np.ndarray], s: str, a_idx: int, r: float, s_next: str, alpha: float = 0.5, gamma: float = 0.9,) -> float: """Updates Q-table using Q-Learning off-policy target.""" best_next_q = float(np.max(q_table[s_next])) target = r + gamma * best_next_q td_error = target - q_table[s][a_idx] q_table[s][a_idx] += alpha * td_error return q_table[s][a_idx]
def step_sarsa( q_table: Dict[str, np.ndarray], s: str, a_idx: int, r: float, s_next: str, a_next_idx: int, alpha: float = 0.5, gamma: float = 0.9,) -> float: """Updates Q-table using SARSA on-policy target.""" actual_next_q = float(q_table[s_next][a_next_idx]) target = r + gamma * actual_next_q td_error = target - q_table[s][a_idx] q_table[s][a_idx] += alpha * td_error return q_table[s][a_idx]
# Identical initial environments: S1 has [Stay=10.0, Fall=-100.0]q_ql = {"S0": np.array([0.0]), "S1": np.array([10.0, -100.0])}q_sarsa = {"S0": np.array([0.0]), "S1": np.array([10.0, -100.0])}
# The agent explores and accidentally executes action index 1 (Fall)res_ql = step_q_learning(q_ql, s="S0", a_idx=0, r=-1.0, s_next="S1")res_sarsa = step_sarsa(q_sarsa, s="S0", a_idx=0, r=-1.0, s_next="S1", a_next_idx=1)
print(f"Q-Learning Q(S0, edge): {res_ql:.1f}")# -> Q-Learning Q(S0, edge): 4.0
print(f"SARSA Q(S0, edge): {res_sarsa:.1f}")# -> SARSA Q(S0, edge): -45.5Watch Out For
Assuming Off-Policy Optimality Always Yields Superior Online Performance
Practitioners frequently default to Q-learning assuming that learning the optimal policy is universally superior to learning the behavioral value . In simulation or offline replay buffers, off-policy methods are indeed more data-efficient. However, in real-world systems (robotics, autonomous navigation, industrial control) where learning happens online, Q-learning's indifference to its own exploration noise can cause real physical damage.
If an agent must train live in an environment where mistakes incur large penalties, use SARSA or decay the exploration parameter aggressively. SARSA safely optimizes performance subject to exploration noise, converging to asymptotically only as .
The Quick Version
- Q-learning is off-policy: it targets , learning optimal values while executing an exploratory policy.
- SARSA is on-policy: it targets , learning the value of the exploratory policy actually being executed.
- In hazardous environments with online exploration, Q-learning learns risky optimal paths, while SARSA learns safer paths that buffer against exploration mistakes.