Fundamentals of Reinforcement Learning
Reinforcement learning is learning what to do by trial and error—an agent interacts with an environment, receiving rewards or penalties, to discover actions that maximize long-term return.
Why Does This Exist?
In supervised learning, an external teacher provides explicit ground-truth labels for every input. If you train a vision model to recognize traffic signs, each image comes pre-packaged with the correct category. But in real-world sequential decision-making—controlling an autonomous vehicle, balancing a power grid, or mastering chess—no supervisor stands by to tell the system the exact optimal motor voltage or steering angle at every microsecond.
When an agent acts sequentially, three distinct hurdles emerge that supervised and unsupervised methods cannot address:
- Evaluative feedback instead of instructive labels: The environment does not tell the agent what the best action would have been. It only provides a scalar score indicating how well the chosen action turned out.
- Delayed consequences: An action taken at the beginning of an episode may directly trigger a catastrophe or victory fifty time steps later. Optimizing solely for immediate gain leads to disastrous traps.
- Active data distribution: Unlike static dataset training where examples are independent and identically distributed (i.i.d.), the agent's current policy directly determines which future states it will visit. Poor early choices cascade into unfamiliar situations.
Reinforcement learning (RL) exists to solve this problem: it formalizes how an autonomous agent can learn goal-directed behavior purely from trial, error, and delayed reward signals.
Think of It Like This
Training a puppy with treats and timing
Imagine training an eight-week-old puppy to sit on command. You cannot hand the puppy an anatomy diagram or verbally explain the biomechanics of bending its hind legs.
Instead, the puppy is the agent, and your living room is the environment. The puppy's current posture and your vocal command constitute the state . The puppy experiments with various actions —it jumps, sniffs your shoe, wanders off, or bends its knees.
When its hindquarters finally touch the floor, you immediately deliver a tasty liver treat—the reward . The puppy does not know calculus, but over repeated trials, it updates its internal behavioral strategy (its policy ) to associate sitting with high cumulative reward.
Crucially, consider what happens if you delay the treat by twenty seconds. During those twenty seconds, the puppy stands back up, barks at the window, and chews on the carpet. If you hand over the treat then, the puppy assumes it is being rewarded for chewing the carpet. This is the temporal credit assignment problem: assigning reward to the specific prior action that actually earned it.
Where the analogy stops: Real puppies possess evolutionary instincts, fatigue, emotional attachment, and biological survival drives. A mathematical RL agent possesses none of these—it optimizes strictly and literally for the mathematical sum of the reward function you define, cheerfully exploiting every unintended loophole.
How It Actually Works
The agent-environment interaction loop
Reinforcement learning formalizes sequential decision making as a closed loop between an agent (the learner and decision-maker) and an environment (everything outside the agent).
The interaction unfolds across a sequence of discrete time steps :
- At time step , the agent senses the environment's current state , where is the set of all valid states.
- Based on this state, the agent selects an action according to its policy :
- The environment receives , updates its internal dynamics, and responds at step with two signals:
- A scalar reward .
- A subsequent state .
This transition is governed by the environment dynamics function:
This cycle produces an ongoing trajectory of experience:
The reward hypothesis and cumulative return
The core premise of reinforcement learning is encapsulated in Richard Sutton's Reward Hypothesis:
All of what we mean by goals and purposes can be well thought of as the maximization of the expected value of the cumulative sum of a received scalar signal (called reward).
The agent's objective is never to maximize just the immediate reward , but the expected cumulative return received from time step forward. For continuing tasks that do not naturally terminate, an unweighted infinite sum could diverge to infinity. RL introduces a discount factor to ensure mathematical convergence and reflect temporal preference:
The discount factor controls the agent's temporal horizon:
- When , the agent is purely myopic, optimizing solely for the immediate reward .
- As , the agent becomes increasingly far-sighted, giving substantial weight to rewards far into the future.
Returns obey a fundamental recursive relationship:
The temporal credit assignment problem
Because rewards can be delayed by dozens or hundreds of steps, the agent faces the challenge of temporal credit assignment: identifying which past decisions were responsible for a distant reward or penalty.
To solve this, RL algorithms estimate value functions that quantify the long-term return expected from a given situation:
- State-Value Function : Expected return starting from state under policy :
- Action-Value Function : Expected return starting from state , taking action , and subsequently following :
Value functions act as an internal scorekeeper, propagating distant rewards backwards so the agent can make forward-looking choices at every step.
Worked numerical example
Consider an agent navigating a short path with discount factor . The episode runs for 3 discrete time steps until reaching a terminal state:
- Time : Agent starts at . Chooses action .
Environment emits immediate reward and transitions to . - Time : Agent at . Chooses action .
Environment emits immediate penalty (energy cost) and transitions to . - Time : Agent at . Chooses action .
Environment emits terminal goal reward and transitions to terminal state .
We calculate the realized return at each step by working backwards from the end of the episode:
-
Return at :
-
Return at :
-
Return at :
We can verify by summing directly from the definition:
The myopic trap comparison
Now contrast this with a short-sighted agent that considers only immediate reward (). At , this agent chooses an alternative action offering immediate reward , but leading into a dead-end swamp state where subsequent rewards are and :
While the greedy agent seemed ahead after step 1 ( vs ), the long-term RL agent achieves a net return of versus . Optimizing cumulative return prevents the agent from walking into immediate-gratification traps.
Code
Below is a self-contained Python implementation of a 1D discrete environment that demonstrates the agent-environment step loop, transition tracking, and recursive discounted return calculation.
from dataclasses import dataclassfrom typing import List, Tuple
@dataclassclass Transition: """Represents a single step transition (S_t, A_t, R_{t+1}, S_{t+1}, done).""" state: int action: int reward: float next_state: int done: bool
class Simple1DGridWorld: """A 1D environment with 5 states: [0: Trap, 1, 2: Start, 3, 4: Goal]. Actions: 0 = Move Left (-1) 1 = Move Right (+1) """
def __init__(self, size: int = 5, goal: int = 4, trap: int = 0) -> None: self.size = size self.goal = goal self.trap = trap self.state = 2 # default center start
def reset(self, start_state: int = 2) -> int: """Resets the environment to initial state S_0.""" self.state = start_state return self.state
def step(self, action: int) -> Tuple[int, float, bool]: """Executes action A_t, returning (S_{t+1}, R_{t+1}, done).""" step_delta = -1 if action == 0 else 1 self.state = max(0, min(self.size - 1, self.state + step_delta))
if self.state == self.goal: return self.state, 10.0, True elif self.state == self.trap: return self.state, -5.0, True else: return self.state, -0.1, False
def compute_discounted_returns( transitions: List[Transition], gamma: float = 0.90) -> List[float]: """Computes G_t backwards using the recursive relation: G_t = R_{t+1} + gamma * G_{t+1}.""" returns: List[float] = [0.0] * len(transitions) running_return = 0.0
for t in reversed(range(len(transitions))): running_return = transitions[t].reward + gamma * running_return returns[t] = round(running_return, 4)
return returns
if __name__ == "__main__": env = Simple1DGridWorld() gamma = 0.90
# Execute a forward policy: move right twice from state 2 to reach goal 4 current_state = env.reset(start_state=2) actions = [1, 1] # Right, Right episode_history: List[Transition] = []
print(f"Initial State S_0: {current_state}") print("-" * 55)
for step_idx, action in enumerate(actions): next_state, reward, done = env.step(action) transition = Transition( state=current_state, action=action, reward=reward, next_state=next_state, done=done, ) episode_history.append(transition) current_state = next_state if done: break
# Compute long-term returns for each time step returns = compute_discounted_returns(episode_history, gamma=gamma)
action_labels = {0: "Left", 1: "Right"} for t, (trans, G_t) in enumerate(zip(episode_history, returns)): print( f"Step t={t}: S_{t}={trans.state} | " f"Action={action_labels[trans.action]} -> " f"R_{t+1}={trans.reward:+.1f}, S_{t+1}={trans.next_state} | " f"Return G_{t}={G_t:.2f}" )Expected Output
Initial State S_0: 2-------------------------------------------------------Step t=0: S_0=2 | Action=Right -> R_1=-0.1, S_1=3 | Return G_0=8.90Step t=1: S_1=3 | Action=Right -> R_2=+10.0, S_2=4 | Return G_1=10.00Watch Out For
Treating reinforcement learning like supervised learning
A common trap for machine learning engineers transitioning into RL is treating the collected dataset as static, independent, and identically distributed (i.i.d.) observations with immediate ground-truth targets.
The Failure Mode:
- Covariate Shift: If you record human actions and train a policy with supervised regression (behavioral cloning), the agent never learns error correction. When test-time execution drifts by just 2% from the demonstration trajectory, the agent lands in an unseen state, predicts inaccurate actions, and compounding drift causes catastrophic failure.
- Greedy Myopia: Supervising an agent on immediate one-step rewards produces short-sighted policies that grab small immediate payoffs while walking directly into terminal penalty traps.
- Exploration Starvation: In supervised learning, the dataset is fixed. In RL, if the agent does not actively explore unfamiliar actions, it can never discover that an initially low-reward path leads to a massive jackpot later.
The Fix:
- Maintain an active closed-loop interaction where the policy generates its own data distribution.
- Formulate the training objective around cumulative discounted return or value functions () rather than instantaneous one-step reward regression.
- Explicitly balance exploration and exploitation (e.g., using -greedy schedules, entropy bonuses, or curiosity-driven objectives) so the agent visits unmapped state-action spaces.
The Quick Version
- The Agent-Environment Loop: RL models decision-making as a discrete-time interaction cycle where the agent observes state , takes action , and receives reward and next state .
- The Reward Hypothesis: All learning objectives are represented as maximizing the expected cumulative discounted return , not merely instant gratification.
- Temporal Credit Assignment: Because rewards are often delayed, value functions and serve to estimate future return and attribute credit back to pivotal earlier decisions.
- Trial, Error, and Active Exploration: Unlike supervised learning with static labels, an RL agent must actively explore state-action space to collect its own non-i.i.d. training experience.