Skip to content
AI360Xpert
Beta

Fundamentals of Reinforcement Learning

Reinforcement learning is learning what to do by trial and error—an agent interacts with an environment, receiving rewards or penalties, to discover actions that maximize long-term return.

The agent-environment loop formalizes trial-and-error learning, where an agent observes states, executes actions, and receives delayed rewards to maximize cumulative return.
The agent-environment loop formalizes trial-and-error learning, where an agent observes states, executes actions, and receives delayed rewards to maximize cumulative return.

Why Does This Exist?

In supervised learning, an external teacher provides explicit ground-truth labels for every input. If you train a vision model to recognize traffic signs, each image comes pre-packaged with the correct category. But in real-world sequential decision-making—controlling an autonomous vehicle, balancing a power grid, or mastering chess—no supervisor stands by to tell the system the exact optimal motor voltage or steering angle at every microsecond.

When an agent acts sequentially, three distinct hurdles emerge that supervised and unsupervised methods cannot address:

  1. Evaluative feedback instead of instructive labels: The environment does not tell the agent what the best action would have been. It only provides a scalar score indicating how well the chosen action turned out.
  2. Delayed consequences: An action taken at the beginning of an episode may directly trigger a catastrophe or victory fifty time steps later. Optimizing solely for immediate gain leads to disastrous traps.
  3. Active data distribution: Unlike static dataset training where examples are independent and identically distributed (i.i.d.), the agent's current policy directly determines which future states it will visit. Poor early choices cascade into unfamiliar situations.

Reinforcement learning (RL) exists to solve this problem: it formalizes how an autonomous agent can learn goal-directed behavior purely from trial, error, and delayed reward signals.

Think of It Like This

Training a puppy with treats and timing

Imagine training an eight-week-old puppy to sit on command. You cannot hand the puppy an anatomy diagram or verbally explain the biomechanics of bending its hind legs.

Instead, the puppy is the agent, and your living room is the environment. The puppy's current posture and your vocal command constitute the state StS_t. The puppy experiments with various actions AtA_t—it jumps, sniffs your shoe, wanders off, or bends its knees.

When its hindquarters finally touch the floor, you immediately deliver a tasty liver treat—the reward Rt+1R_{t+1}. The puppy does not know calculus, but over repeated trials, it updates its internal behavioral strategy (its policy π\pi) to associate sitting with high cumulative reward.

Crucially, consider what happens if you delay the treat by twenty seconds. During those twenty seconds, the puppy stands back up, barks at the window, and chews on the carpet. If you hand over the treat then, the puppy assumes it is being rewarded for chewing the carpet. This is the temporal credit assignment problem: assigning reward to the specific prior action that actually earned it.

Where the analogy stops: Real puppies possess evolutionary instincts, fatigue, emotional attachment, and biological survival drives. A mathematical RL agent possesses none of these—it optimizes strictly and literally for the mathematical sum of the reward function you define, cheerfully exploiting every unintended loophole.

How It Actually Works

The agent-environment interaction loop

Reinforcement learning formalizes sequential decision making as a closed loop between an agent (the learner and decision-maker) and an environment (everything outside the agent).

The interaction unfolds across a sequence of discrete time steps t=0,1,2,…t = 0, 1, 2, \dots:

  1. At time step tt, the agent senses the environment's current state St∈SS_t \in \mathcal{S}, where S\mathcal{S} is the set of all valid states.
  2. Based on this state, the agent selects an action At∈A(St)A_t \in \mathcal{A}(S_t) according to its policy π\pi: π(a∣s)=P(At=a∣St=s)\pi(a \mid s) = P(A_t = a \mid S_t = s)
  3. The environment receives AtA_t, updates its internal dynamics, and responds at step t+1t+1 with two signals:
    • A scalar reward Rt+1∈RR_{t+1} \in \mathbb{R}.
    • A subsequent state St+1∈SS_{t+1} \in \mathcal{S}.

This transition is governed by the environment dynamics function: p(s′,r∣s,a)=P(St+1=s′,Rt+1=r∣St=s,At=a)p(s', r \mid s, a) = P(S_{t+1} = s', R_{t+1} = r \mid S_t = s, A_t = a)

This cycle produces an ongoing trajectory of experience: τ=(S0,A0,R1,S1,A1,R2,S2,… )\tau = (S_0, A_0, R_1, S_1, A_1, R_2, S_2, \dots)

The reward hypothesis and cumulative return

The core premise of reinforcement learning is encapsulated in Richard Sutton's Reward Hypothesis:

All of what we mean by goals and purposes can be well thought of as the maximization of the expected value of the cumulative sum of a received scalar signal (called reward).

The agent's objective is never to maximize just the immediate reward Rt+1R_{t+1}, but the expected cumulative return GtG_t received from time step tt forward. For continuing tasks that do not naturally terminate, an unweighted infinite sum could diverge to infinity. RL introduces a discount factor γ∈[0,1)\gamma \in [0, 1) to ensure mathematical convergence and reflect temporal preference:

Gt=∑k=0∞γkRt+k+1=Rt+1+γRt+2+γ2Rt+3+…G_t = \sum_{k=0}^{\infty} \gamma^k R_{t+k+1} = R_{t+1} + \gamma R_{t+2} + \gamma^2 R_{t+3} + \dots

The discount factor controls the agent's temporal horizon:

  • When γ=0\gamma = 0, the agent is purely myopic, optimizing solely for the immediate reward Rt+1R_{t+1}.
  • As γ→1\gamma \to 1, the agent becomes increasingly far-sighted, giving substantial weight to rewards far into the future.

Returns obey a fundamental recursive relationship: Gt=Rt+1+γGt+1G_t = R_{t+1} + \gamma G_{t+1}

The temporal credit assignment problem

Because rewards can be delayed by dozens or hundreds of steps, the agent faces the challenge of temporal credit assignment: identifying which past decisions were responsible for a distant reward or penalty.

To solve this, RL algorithms estimate value functions that quantify the long-term return expected from a given situation:

  • State-Value Function Vπ(s)V^\pi(s): Expected return starting from state ss under policy π\pi: Vπ(s)=Eπ[Gt∣St=s]V^\pi(s) = \mathbb{E}_\pi \left[ G_t \mid S_t = s \right]
  • Action-Value Function Qπ(s,a)Q^\pi(s, a): Expected return starting from state ss, taking action aa, and subsequently following π\pi: Qπ(s,a)=Eπ[Gt∣St=s,At=a]Q^\pi(s, a) = \mathbb{E}_\pi \left[ G_t \mid S_t = s, A_t = a \right]

Value functions act as an internal scorekeeper, propagating distant rewards backwards so the agent can make forward-looking choices at every step.

Worked numerical example

Consider an agent navigating a short path with discount factor γ=0.90\gamma = 0.90. The episode runs for 3 discrete time steps until reaching a terminal state:

  • Time t=0t=0: Agent starts at S0=StartS_0 = \text{Start}. Chooses action A0=AdvanceA_0 = \text{Advance}.
    Environment emits immediate reward R1=0.0R_1 = 0.0 and transitions to S1=CorridorS_1 = \text{Corridor}.
  • Time t=1t=1: Agent at S1S_1. Chooses action A1=Cross Rough TerrainA_1 = \text{Cross Rough Terrain}.
    Environment emits immediate penalty R2=−2.0R_2 = -2.0 (energy cost) and transitions to S2=Goal ApproachS_2 = \text{Goal Approach}.
  • Time t=2t=2: Agent at S2S_2. Chooses action A2=Enter GoalA_2 = \text{Enter Goal}.
    Environment emits terminal goal reward R3=+10.0R_3 = +10.0 and transitions to terminal state S3=GoalS_3 = \text{Goal}.

We calculate the realized return GtG_t at each step by working backwards from the end of the episode:

  1. Return at t=2t=2: G2=R3=10.00G_2 = R_3 = 10.00

  2. Return at t=1t=1: G1=R2+γG2=−2.00+0.90×(10.00)=−2.00+9.00=+7.00G_1 = R_2 + \gamma G_2 = -2.00 + 0.90 \times (10.00) = -2.00 + 9.00 = +7.00

  3. Return at t=0t=0: G0=R1+γG1=0.00+0.90×(7.00)=0.00+6.30=+6.30G_0 = R_1 + \gamma G_1 = 0.00 + 0.90 \times (7.00) = 0.00 + 6.30 = +6.30

We can verify G0G_0 by summing directly from the definition: G0=R1+γR2+γ2R3=0.0+0.90×(−2.0)+(0.90)2×(10.0)=−1.80+0.81×10.0=+6.30G_0 = R_1 + \gamma R_2 + \gamma^2 R_3 = 0.0 + 0.90 \times (-2.0) + (0.90)^2 \times (10.0) = -1.80 + 0.81 \times 10.0 = +6.30

The myopic trap comparison

Now contrast this with a short-sighted agent that considers only immediate reward (γ=0\gamma = 0). At t=0t=0, this agent chooses an alternative action A0′=Take SnackA_0' = \text{Take Snack} offering immediate reward R1′=+1.50R_1' = +1.50, but leading into a dead-end swamp state where subsequent rewards are R2′=−5.00R_2' = -5.00 and R3′=0.00R_3' = 0.00:

G0trap=R1′+γR2′+γ2R3′=1.50+0.90(−5.00)+(0.90)2(0.00)=1.50−4.50=−3.00G_0^{\text{trap}} = R_1' + \gamma R_2' + \gamma^2 R_3' = 1.50 + 0.90(-5.00) + (0.90)^2(0.00) = 1.50 - 4.50 = -3.00

While the greedy agent seemed ahead after step 1 (+1.50+1.50 vs 0.000.00), the long-term RL agent achieves a net return of +6.30+6.30 versus −3.00-3.00. Optimizing cumulative return GtG_t prevents the agent from walking into immediate-gratification traps.

Code

Below is a self-contained Python implementation of a 1D discrete environment that demonstrates the agent-environment step loop, transition tracking, and recursive discounted return calculation.

from dataclasses import dataclassfrom typing import List, Tuple

@dataclassclass Transition:    """Represents a single step transition (S_t, A_t, R_{t+1}, S_{t+1}, done)."""    state: int    action: int    reward: float    next_state: int    done: bool

class Simple1DGridWorld:    """A 1D environment with 5 states: [0: Trap, 1, 2: Start, 3, 4: Goal].        Actions:      0 = Move Left (-1)      1 = Move Right (+1)    """
    def __init__(self, size: int = 5, goal: int = 4, trap: int = 0) -> None:        self.size = size        self.goal = goal        self.trap = trap        self.state = 2  # default center start
    def reset(self, start_state: int = 2) -> int:        """Resets the environment to initial state S_0."""        self.state = start_state        return self.state
    def step(self, action: int) -> Tuple[int, float, bool]:        """Executes action A_t, returning (S_{t+1}, R_{t+1}, done)."""        step_delta = -1 if action == 0 else 1        self.state = max(0, min(self.size - 1, self.state + step_delta))
        if self.state == self.goal:            return self.state, 10.0, True        elif self.state == self.trap:            return self.state, -5.0, True        else:            return self.state, -0.1, False

def compute_discounted_returns(    transitions: List[Transition], gamma: float = 0.90) -> List[float]:    """Computes G_t backwards using the recursive relation: G_t = R_{t+1} + gamma * G_{t+1}."""    returns: List[float] = [0.0] * len(transitions)    running_return = 0.0
    for t in reversed(range(len(transitions))):        running_return = transitions[t].reward + gamma * running_return        returns[t] = round(running_return, 4)
    return returns

if __name__ == "__main__":    env = Simple1DGridWorld()    gamma = 0.90
    # Execute a forward policy: move right twice from state 2 to reach goal 4    current_state = env.reset(start_state=2)    actions = [1, 1]  # Right, Right    episode_history: List[Transition] = []
    print(f"Initial State S_0: {current_state}")    print("-" * 55)
    for step_idx, action in enumerate(actions):        next_state, reward, done = env.step(action)        transition = Transition(            state=current_state,            action=action,            reward=reward,            next_state=next_state,            done=done,        )        episode_history.append(transition)        current_state = next_state        if done:            break
    # Compute long-term returns for each time step    returns = compute_discounted_returns(episode_history, gamma=gamma)
    action_labels = {0: "Left", 1: "Right"}    for t, (trans, G_t) in enumerate(zip(episode_history, returns)):        print(            f"Step t={t}: S_{t}={trans.state} | "            f"Action={action_labels[trans.action]} -> "            f"R_{t+1}={trans.reward:+.1f}, S_{t+1}={trans.next_state} | "            f"Return G_{t}={G_t:.2f}"        )

Expected Output

Initial State S_0: 2-------------------------------------------------------Step t=0: S_0=2 | Action=Right -> R_1=-0.1, S_1=3 | Return G_0=8.90Step t=1: S_1=3 | Action=Right -> R_2=+10.0, S_2=4 | Return G_1=10.00

Watch Out For

Treating reinforcement learning like supervised learning

A common trap for machine learning engineers transitioning into RL is treating the collected dataset as static, independent, and identically distributed (i.i.d.) observations with immediate ground-truth targets.

The Failure Mode:

  1. Covariate Shift: If you record human actions and train a policy with supervised regression (behavioral cloning), the agent never learns error correction. When test-time execution drifts by just 2% from the demonstration trajectory, the agent lands in an unseen state, predicts inaccurate actions, and compounding drift causes catastrophic failure.
  2. Greedy Myopia: Supervising an agent on immediate one-step rewards produces short-sighted policies that grab small immediate payoffs while walking directly into terminal penalty traps.
  3. Exploration Starvation: In supervised learning, the dataset is fixed. In RL, if the agent does not actively explore unfamiliar actions, it can never discover that an initially low-reward path leads to a massive jackpot later.

The Fix:

  • Maintain an active closed-loop interaction where the policy generates its own data distribution.
  • Formulate the training objective around cumulative discounted return GtG_t or value functions (V,QV, Q) rather than instantaneous one-step reward regression.
  • Explicitly balance exploration and exploitation (e.g., using ϵ\epsilon-greedy schedules, entropy bonuses, or curiosity-driven objectives) so the agent visits unmapped state-action spaces.

The Quick Version

  • The Agent-Environment Loop: RL models decision-making as a discrete-time interaction cycle where the agent observes state StS_t, takes action AtA_t, and receives reward Rt+1R_{t+1} and next state St+1S_{t+1}.
  • The Reward Hypothesis: All learning objectives are represented as maximizing the expected cumulative discounted return Gt=∑k=0∞γkRt+k+1G_t = \sum_{k=0}^{\infty} \gamma^k R_{t+k+1}, not merely instant gratification.
  • Temporal Credit Assignment: Because rewards are often delayed, value functions V(s)V(s) and Q(s,a)Q(s, a) serve to estimate future return and attribute credit back to pivotal earlier decisions.
  • Trial, Error, and Active Exploration: Unlike supervised learning with static labels, an RL agent must actively explore state-action space to collect its own non-i.i.d. training experience.