Basic Concepts in Reinforcement Learning
Reinforcement learning is trial-and-error learning where an agent discovers how to act in an environment by receiving rewards and aiming to maximize cumulative gain over time.
Why Does This Exist?
In supervised learning, an external supervisor gives an explicit ground-truth label for every input. The model is told immediately what output it should have produced. In unsupervised learning, an algorithm searches for latent structure without any task feedback or external score.
Real-world autonomous intelligence does not fit either regime. When a robotic arm learns manipulation, an autonomous vehicle navigates rush-hour traffic, or an agent plays chess, there is no supervisor standing by to score every micro-action at step 37 of a thousand-step episode. The agent only receives evaluative feedback: occasional positive or negative rewards indicating whether intermediate states and final outcomes were beneficial.
Without the formal framework of Reinforcement Learning (RL), learning by trial and error breaks down in three critical ways:
- Delayed credit assignment: An action taken early in an episode may produce consequences dozens or hundreds of steps later. Naive optimization cannot tell which past decision caused a delayed success or failure.
- Distributional shift through action: In static machine learning, the dataset is fixed. In sequential decision-making, every action taken alters the future distribution of states the agent visits. A single mistake moves the agent into unfamiliar regions of the state space.
- Evaluative vs. instructive feedback: Evaluative feedback tells an agent how well it performed, but not what the optimal action was. The agent must actively explore alternative actions to discover superior strategies.
Reinforcement learning formalizes this interaction into a rigorous mathematical loop where agents learn goal-directed behavior by maximizing long-term cumulative reward.
Think of It Like This
Anatomy of a board game championship
Think of reinforcement learning as mastering a high-stakes board game:
- State (): The exact arrangement of pieces on the board at your turn. It captures everything you need to know about the current situation to decide your next move.
- Action (): Any legal move you can play from the current board position.
- Policy (): Your personal playbook or strategy—a systematic rulebook that dictates which move you select given any board configuration.
- Environment Dynamics / Transition (): The game rules and your opponent's reaction. Once you move a piece, the board updates according to the rules and your opponent responds, producing the new board state.
- Reward (): Immediate points awarded or lost (e.g., capturing an opponent piece awards ; losing a piece costs ).
- Return (): Your total score across the entire game, where points scored soon are valued more predictably than hypothetical captures far into the future.
- Where the analogy stops: In a board game, turns are discrete, the entire board is usually visible, and rules are deterministic. In real-world reinforcement learning, states can be noisy continuous sensor feeds, transitions are often stochastic with unknown physical dynamics, and tasks may run infinitely without rounds or turns.
How It Actually Works
The Five Pillars and the Agent-Environment Loop
At each discrete time step , the interaction between the agent and the environment proceeds through five interconnected pillars:
| Pillar | Symbol | Formal Definition | Role in the RL Loop |
|---|---|---|---|
| State | Complete description of the environment at step | Sensory representation observed by the agent | |
| Action | Choice selected from available action space | Decision emitted by the agent to influence the world | |
| Transition | World dynamics governing state evolution | ||
| Reward | Scalar feedback signal emitted by the environment | Immediate evaluative score for the transition | |
| Policy | or | Decision function mapping states to actions |
The cycle repeats indefinitely or until a terminal state is reached:
- The agent observes the current state .
- The agent selects an action according to its policy .
- The environment processes , transitions to a new state , and emits a scalar reward .
- The agent receives and , updating its internal policy or value estimates.
Trajectories, Returns, and the Discount Factor
A sequence of interactions produces an experience trajectory :
Tasks fall into two primary structures:
- Episodic Tasks: Interaction naturally breaks into distinct episodes that terminate at time step upon reaching an absorbing terminal state (e.g., checkmate, game over, reaching a maze exit).
- Continuing Tasks: Interaction proceeds infinitely without termination (), such as an automated thermostat, industrial process control, or continuous inventory management.
To evaluate an agent's performance, we do not optimize isolated immediate rewards. Instead, we optimize the cumulative return , defined from time step onward with a discount factor :
This definition yields the fundamental recursive relationship used throughout reinforcement learning:
The discount factor plays two vital roles:
- Mathematical Convergence: When tasks are continuing (), summing unweighted rewards bounded by would diverge to infinity. When , the geometric series bounds the maximum possible return:
- Behavioral Horizon: The discount factor tunes the agent's effective planning horizon . When , the agent is purely myopic, maximizing only the immediate reward . As , the agent becomes farsighted, weighting distant consequences heavily.
Value Functions and the Optimization Target
Because future state transitions and policies can be stochastic, the return is a random variable. An agent evaluates states using expected returns:
- State-Value Function : The expected return starting from state under policy :
- Action-Value Function : The expected return of taking action in state and thereafter following :
The overarching objective of reinforcement learning is to discover an optimal policy that maximizes the expected return from the start:
Worked numerical example
Consider a 4-step episodic trajectory generated by an agent navigating an obstacle course:
At step 4, the agent reaches the terminal goal state . Because is terminal, no subsequent rewards exist, so .
Let the discount factor be . We compute the return at each step by propagating backwards using :
- Terminal step :
- Step : Transition to terminal goal yielding reward :
- Step : Transition to intermediate state yielding reward :
- Step : Step across rough terrain to yielding penalty :
- Initial step : Initial transition from to yielding reward :
We verify by expanding the full power series directly:
Comparing the return at under different values of illustrates the impact of discounting:
- (myopic): (ignores both the penalty and the goal)
- (short-sighted):
- (farsighted):
- (undiscounted sum):
Code
from dataclasses import dataclassfrom typing import List, Sequence
@dataclass(frozen=True)class TransitionStep: state: str action: str reward: float next_state: str
def compute_discounted_returns( rewards: Sequence[float], gamma: float) -> List[float]: """Compute backwards cumulative discounted return G_t for each step. Uses the recursive formulation: G_t = R_{t+1} + gamma * G_{t+1}. Runs in O(N) time and O(N) auxiliary space. """ if not (0.0 <= gamma <= 1.0): raise ValueError(f"Gamma must be in [0.0, 1.0], got {gamma}") n = len(rewards) returns: List[float] = [0.0] * n running_return = 0.0 # Backwards accumulation from terminal step T to step 0 for t in reversed(range(n)): running_return = rewards[t] + gamma * running_return returns[t] = running_return return returns
# Sample episodic trajectory: 4 steps leading to a goal statetrajectory: List[TransitionStep] = [ TransitionStep(state="S0", action="East", reward=2.0, next_state="S1"), TransitionStep(state="S1", action="South", reward=-1.0, next_state="S2"), TransitionStep(state="S2", action="East", reward=0.0, next_state="S3"), TransitionStep(state="S3", action="North", reward=10.0, next_state="S4_goal"),]
rewards = [step.reward for step in trajectory]gammas = [0.0, 0.5, 0.9, 1.0]
print("Step-by-step Discounted Returns G_t across Gammas:")header = f"{'Step':<6} | {'Reward':<8} | " + " | ".join([f"gamma={g:<4}" for g in gammas])print(header)print("-" * len(header))
all_returns = {g: compute_discounted_returns(rewards, g) for g in gammas}
for t, step in enumerate(trajectory): ret_str = " | ".join([f"{all_returns[g][t]:<6.2f}" for g in gammas]) print(f"t={t:<4} | R_{t+1}={step.reward:<4.1f} | {ret_str}")Expected output:
Step-by-step Discounted Returns G_t across Gammas:Step | Reward | gamma=0.0 | gamma=0.5 | gamma=0.9 | gamma=1.0 ---------------------------------------------------------------------t=0 | R_1=2.0 | 2.00 | 2.75 | 8.39 | 11.00 t=1 | R_2=-1.0 | -1.00 | 1.50 | 7.10 | 9.00 t=2 | R_3=0.0 | 0.00 | 5.00 | 9.00 | 10.00 t=3 | R_4=10.0 | 10.00 | 10.00 | 10.00 | 10.00 Watch Out For
Confusing immediate reward with expected return
Failure mode: Designing an agent or choosing actions greedily based on immediate reward rather than cumulative return or value function .
Symptom: The agent exhibits severe myopic failure. It falls into obvious traps by grabbing small instant rewards (e.g., picking up an isolated coin in front of a pit), refuses to accept short-term costs necessary for long-term gains (e.g., refusing to pay an engine acceleration cost or sacrifice a chess piece), or circles repeatedly in loops of minor positive rewards.
Concrete fix: Never optimize individual step rewards in isolation. Formulate all policy decisions through value functions or that estimate the expectation of total discounted future return . Always set the discount factor high enough to cover the task's natural credit assignment horizon ().
Setting discount factor gamma = 1.0 in continuing tasks
Failure mode: Setting in continuing tasks where no terminal state exists.
Symptom: When episodes do not terminate, the cumulative return diverges to or . Temporal difference errors explode, neural network value targets blow up to numerical overflow, and gradient descent destabilizes.
Concrete fix: For infinite-horizon continuing environments, strictly set (commonly between and ) so returns remain geometrically bounded, or reformulate the optimization objective using the average-reward MDP formulation.
The Quick Version
- Reinforcement learning models sequential decision-making through an agent-environment interaction loop governed by States, Actions, and Rewards.
- The policy determines the agent's behavior, while environment transition dynamics dictate physics, transitions, and feedback.
- The core objective is maximizing expected discounted return , rather than picking greedy immediate rewards.
- The discount factor balances short-term survival against long-term planning, and guarantees finite returns in infinite-horizon continuing tasks.