Reinforcement Learning
30 interview questions in this topic, each with its full answer shown below. Use "Collapse all" to skim just the titles.
“What is an advantage function, and how does Actor-Critic use it?”
The advantage function measures how much better taking a specific action is compared to the average action in that state. Actor-Critic methods use it to reduce variance in policy gradients: the Critic estimates the advantage, and the Actor updates its policy based on this, scaling gradients by relative rather than absolute rewards.
Answer
In Reinforcement Learning, the Advantage Function is defined as the difference between the action-value function and the state-value function :
Understanding Advantage
- represents the baseline expected return from state under the current policy. It's the "average" outcome.
- represents the expected return if we specifically take action in state . Therefore, quantifies how much better (or worse) action is compared to the average action the policy would usually take.
- If , the action is better than average.
- If , the action is worse than average.
Role in Actor-Critic Methods
Pure policy gradient methods (like REINFORCE) use the raw return to scale gradients, leading to high variance. Actor-Critic architectures solve this by combining policy gradients with value-based methods:
- The Critic (Value Network): Learns to estimate the state-value function . Using TD-learning, it can approximate the Advantage without waiting for the episode to end (e.g., ).
- The Actor (Policy Network): Updates the policy parameters based on feedback from the Critic. Instead of scaling the policy gradient by raw returns, it scales it by the estimated Advantage: .
Using Advantage dramatically reduces variance because the Actor is updated based on relative performance. Even if absolute returns are high everywhere in an environment, the Actor only increases probabilities for actions that perform better than expected for that specific state.
💡 Note: In modern algorithms like PPO, Generalized Advantage Estimation (GAE) is used to elegantly balance bias and variance when calculating advantages over multiple timesteps.
“What are agent, environment, state, action, and reward?”
In RL, the agent is the decision-maker interacting with the environment, which is everything outside the agent. A state is a representation of the environment's current situation. An action is a choice made by the agent. A reward is the scalar feedback signal evaluating the agent's action.
Answer
These five elements form the fundamental building blocks of the Reinforcement Learning (RL) framework:
- Agent: The learner and decision-maker. It observes the environment and takes actions based on its internal policy to maximize cumulative reward.
- Environment: Everything the agent interacts with. It responds to the agent's actions by transitioning to a new state and emitting a reward. The environment is typically governed by unknown dynamics.
- State (): A representation of the environment's current situation at a specific time step . It is the information the agent uses to make decisions. If the state contains all necessary information from the history to predict the future, it has the Markov property.
- Action (): A valid move or choice the agent makes in state . The set of all possible actions is called the action space, which can be discrete (e.g., move left/right) or continuous (e.g., steering wheel angle).
- Reward (): A scalar feedback signal received from the environment immediately after taking an action. The reward defines the goal of the RL problem; the agent's sole objective is to maximize the expected sum of future rewards over time.
The interaction loop is continuous: at step , the agent receives state , selects action , and the environment transitions to while providing reward .
💡 Note: The boundary between agent and environment is drawn at the limit of the agent's absolute control, not necessarily its physical boundary. For example, a robot's motors are part of the environment because they can fail or wear out.
“What is the Bellman equation, and why is it central to RL?”
The Bellman equation expresses a relationship between the value of a state and the values of its successor states. It breaks down the value function into two parts: the immediate reward and the discounted value of the next state, forming the theoretical foundation for RL algorithms to compute optimal policies.
Answer
The Bellman equation is the fundamental recursive mathematical relationship that underpins almost all of Reinforcement Learning. Named after Richard Bellman, it decomposes the value of being in a state (or taking an action in a state) into two parts:
- The immediate reward received.
- The discounted expected value of the next state.
For a state-value function , the Bellman Expectation Equation is:
In simpler terms, the value of your current state is the reward you expect to get right now, plus the discounted value of wherever you end up next.
Why is it Central to RL?
- Enables Dynamic Programming: The recursive nature of the Bellman equation allows the complex problem of estimating total future returns over a long horizon to be broken down into simpler, step-by-step calculations.
- Defines Optimality: The Bellman Optimality Equation states that the value of a state under the optimal policy must equal the expected return for the best action from that state. This gives RL algorithms a concrete mathematical target to solve for.
- Foundation for Algorithms:
- Dynamic Programming methods solve the Bellman equation exactly when the environment model is known.
- Temporal Difference (TD) learning methods (like Q-Learning and SARSA) sample from the environment to approximate the Bellman equation. They update current value estimates based on the value estimates of the next state, a process known as bootstrapping.
💡 Note: Without the Bellman equation, agents would have to wait until the end of every episode (like in Monte Carlo methods) to update their value estimates, which is highly inefficient and impossible in continuing tasks.
“How would you bridge the sim-to-real gap for a robotics RL policy?”
Bridging the sim-to-real gap involves training robust policies in simulation that transfer to the real world. Key techniques include domain randomization (varying physical parameters in simulation) and system identification (building highly accurate physical models of the specific robot).
Answer
Training RL algorithms directly on physical robots is slow, expensive, and dangerous. Therefore, policies are heavily trained in physics simulators. However, simulators are imperfect approximations. The sim-to-real gap refers to the severe performance drop that occurs when a policy trained in simulation is deployed on real hardware, caused by discrepancies in friction, sensor noise, latency, and motor dynamics.
To bridge this gap, engineers use several techniques:
1. Domain Randomization
Instead of training the agent in a single, perfectly modeled simulator, the environment's physics and visual parameters are randomized at the start of every episode.
- Physical Randomization: Randomizing object masses, table friction, motor torque limits, and joint damping.
- Visual Randomization: Randomizing lighting, textures, and camera positions. By forcing the RL policy to succeed across a massive distribution of physics environments, the agent learns a robust, generalized policy. When deployed, it treats the real world as just another variation of the simulation it has already mastered.
2. System Identification
This involves meticulously tuning the simulator to perfectly match the specific real-world robot. Engineers collect real-world data and use supervised learning or optimization algorithms to refine the simulator's internal parameters (like exact inertia matrices or actuator delay) until the simulated trajectories perfectly match the real trajectories.
3. Asymmetric Actor-Critic
During simulated training, the Critic is given access to privileged information (exact object coordinates, exact friction coefficients) to calculate perfect value estimates. The Actor is restricted to realistic, noisy sensor inputs (like a camera feed). This allows for highly efficient training while ensuring the Actor can still function in the real world where privileged information is unavailable.
💡 Note: Fine-tuning is often the final step. A policy is pre-trained in simulation and then fine-tuned with a highly sample-efficient off-policy algorithm (like SAC) for a few hours on the real robot to adapt to the final real-world intricacies.
“What is the credit assignment problem?”
The credit assignment problem is the challenge of determining which specific actions in a long sequence were responsible for a delayed reward or penalty. When an agent receives a reward at the end of a long episode, it struggles to assign 'credit' accurately to the individual steps taken along the way.
Answer
The Credit Assignment Problem is one of the most fundamental difficulties in Reinforcement Learning, stemming from the fact that rewards are often delayed.
Imagine playing a game of chess. You might make a brilliant, highly strategic move on turn 10. However, you don't receive a reward (winning the game) until turn 40. When the game ends and the agent receives a reward, the credit assignment problem asks: Which of the 40 moves were actually responsible for the win?
Because the agent updates its policy based on rewards, it needs to know precisely which actions to reinforce.
There are two forms of this problem:
- Temporal Credit Assignment: Determining which action in a sequence over time caused the outcome. Did the last action win the game, or was it a crucial move 20 steps ago?
- Structural Credit Assignment: In multi-agent scenarios, determining which specific agent or subsystem contributed to a global team reward.
RL algorithms utilize various techniques to solve this:
- Temporal Difference (TD) Learning: Algorithms like Q-learning bootstrap values, backing up the final reward step-by-step through the trajectory so earlier states learn their value without waiting for the final reward.
- Eligibility Traces: An intermediate approach between TD and Monte Carlo methods, maintaining a memory of recently visited states and decaying the assigned credit exponentially for actions taken further in the past.
💡 Note: The problem is severely exacerbated in environments with sparse rewards, where the agent might take thousands of actions before receiving any feedback signal at all.
“What is the discount factor, and what does it control?”
The discount factor, denoted as gamma (γ), is a hyperparameter between 0 and 1 that determines the importance of future rewards compared to immediate rewards. A gamma near 0 makes the agent short-sighted, favoring immediate rewards, while a gamma near 1 makes it far-sighted, striving for long-term returns.
Answer
In Reinforcement Learning, the goal of an agent is to maximize the expected sum of rewards. However, simply summing future rewards can lead to infinite returns in continuous tasks. To handle this mathematically and to model the preference for sooner rewards over later ones, we use the discount factor, (gamma).
The discounted return at time step is calculated as:
The discount factor lies in the range and fundamentally controls the agent's time horizon:
- (Myopic/Short-sighted): As approaches 0, the agent heavily discounts future rewards. It becomes myopic, optimizing only for the immediate reward and ignoring long-term consequences.
- (Far-sighted): As approaches 1, the agent weighs future rewards almost as heavily as immediate rewards. It becomes willing to sacrifice immediate gains for much larger payoffs in the distant future.
Beyond modeling time preference, the discount factor is crucial for mathematical convergence. In continuing tasks that do not terminate, an undiscounted sum of rewards would approach infinity. By ensuring , the infinite series of expected returns converges to a finite value, allowing algorithms to compute stable value functions.
💡 Note: While can be used in episodic tasks (since the episode guarantees termination), it is rarely used in continuing tasks because it destroys the convergence properties of the value function.
“Explain Direct Preference Optimization (DPO) and how it differs from RLHF with a reward model.”
DPO mathematically bypasses the need for a separate reward model and complex RL algorithms like PPO. Instead of training a reward model to grade text, DPO directly optimizes the language model on human preference data using a simple classification loss, making training significantly faster and more stable.
Answer
Aligning Large Language Models with human preferences is historically done via Reinforcement Learning from Human Feedback (RLHF).
The Standard RLHF Approach
RLHF is a complex, multi-stage pipeline:
- Collect a dataset of human preferences (Response A is better than Response B).
- Train a separate Reward Model (RM) on this data to output a scalar score.
- Use an RL algorithm (PPO) to update the LLM weights to maximize the RM's score, while applying a KL-divergence penalty to stop the LLM from drifting too far from its original behavior. This requires keeping four models in memory simultaneously (the Actor LLM, the Reference LLM, the Critic, and the Reward Model), making it computationally heavy and notoriously unstable to tune.
The DPO Breakthrough
Direct Preference Optimization (DPO) revolutionizes this process by proving a mathematical equivalence. The authors of DPO showed that the objective of the Reward Model and the objective of the RL policy can be combined into a single equation.
By re-parameterizing the reward function implicitly in terms of the optimal policy, DPO completely eliminates the need for:
- Training a separate Reward Model.
- Running PPO or any RL loop.
How DPO works: It takes the dataset of human preferences (chosen response and rejected response ). It then formulates a simple binary cross-entropy loss function. It updates the LLM to increase the probability of generating the chosen response and decrease the probability of the rejected response , relative to the base reference model.
DPO achieves performance comparable to or better than RLHF but is essentially just standard supervised learning. It is vastly more lightweight, stable, and requires significantly less hyperparameter tuning.
💡 Note: While DPO is simpler, some researchers argue that having an explicit Reward Model (like in traditional RLHF) is still beneficial because a heavily trained RM might generalize better to OOD (Out-of-Distribution) prompts than DPO's implicit optimization.
“How does Deep Q-Network (DQN) work, and why are experience replay and target networks needed?”
DQN combines Q-learning with deep neural networks to approximate Q-values for high-dimensional state spaces. It stabilizes training using Experience Replay (breaking temporal correlation of data) and Target Networks (providing stable TD targets), which prevent the network from diverging.
Answer
Traditional Q-learning uses a table to store Q-values for every state-action pair. However, in complex environments like video games or robotics, the state space is too massive (e.g., millions of pixel configurations) for a tabular approach. Deep Q-Network (DQN) solves this by using a neural network to approximate the Q-value function: .
However, naively combining non-linear function approximators (neural networks) with off-policy RL (Q-learning) and bootstrapping (TD learning) often leads to divergence. This is known as the "deadly triad." DQN introduces two critical innovations to stabilize training:
1. Experience Replay
- The Problem: In standard RL, the agent learns from sequential data (). This data is highly temporally correlated, which violates the i.i.d. (independent and identically distributed) assumption required for stable neural network training.
- The Solution: As the agent interacts with the environment, it stores transition tuples in a memory buffer. During training, it samples random mini-batches from this buffer. This breaks the temporal correlations and smooths out learning over the data distribution.
2. Target Networks
- The Problem: The Q-learning update rule uses the network itself to calculate the target value: . Because the target depends on the same parameters () being updated, the target moves during training, akin to a dog chasing its own tail. This causes severe instability.
- The Solution: DQN uses a separate Target Network to calculate the TD target. It has the same architecture as the main network but its parameters () are frozen. Periodically (e.g., every 10,000 steps), the main network's parameters are copied over to the target network. This provides a stable, fixed target for the loss function over multiple updates.
💡 Note: DQN's success on Atari 2600 games marked a massive breakthrough, proving that an agent could learn directly from high-dimensional raw pixel inputs without hand-crafted features.
“What is the difference between episodic and continuing tasks?”
Episodic tasks have a clear, well-defined end point (a terminal state), breaking interaction into discrete episodes, like a game of chess. Continuing tasks have no terminal state and go on indefinitely, like a robot constantly managing a building's HVAC system.
Answer
In Reinforcement Learning, the interaction between the agent and the environment can be structured in two fundamentally different ways, depending on the nature of the problem.
Episodic Tasks
In episodic tasks, the agent-environment interaction naturally breaks down into discrete, separate sequences called episodes.
- Each episode begins in a starting state (often sampled from a distribution) and ends when the agent reaches a specific terminal state (e.g., winning a game, losing a game, or a robot crashing).
- Upon reaching the terminal state, the environment resets, and a new, independent episode begins.
- Because the time horizon is finite, the total un-discounted return is guaranteed to be bounded. As a result, setting the discount factor is perfectly mathematically valid in episodic formulations.
- Examples: Playing Super Mario, a game of Go, or a single run through a maze.
Continuing Tasks
In continuing tasks, the agent-environment interaction goes on indefinitely without any natural end point or terminal state.
- The agent continuously operates, and there are no breaks or resets.
- Because the interaction never ends, attempting to calculate the total un-discounted future return would result in an infinite sum. To prevent value functions from diverging to infinity, continuing tasks must use a discount factor . This ensures that rewards far into the future have an exponentially decaying weight, keeping the expected return finite.
- Examples: Algorithmic trading, managing a server cluster's load, or a personal assistant robot operating continuously in a house.
💡 Note: To provide a unified mathematical notation for both types, episodic tasks are often mathematically treated as continuing tasks by assuming that once the terminal state is reached, the agent enters a special "absorbing state" that transitions only to itself and yields zero reward forever.
“What is the epsilon-greedy strategy?”
Epsilon-greedy is a simple but effective strategy for managing the exploration-exploitation tradeoff. With a small probability epsilon (ε), the agent explores by picking a random action. With probability 1 - ε, it exploits by picking the action with the highest estimated value.
Answer
The -greedy (epsilon-greedy) strategy is one of the most widely used action-selection methods in value-based Reinforcement Learning (such as Q-learning) to resolve the exploration versus exploitation dilemma.
At each time step, the agent must choose an action. Under the -greedy policy, the agent behaves as follows:
- Exploitation (probability ): The agent greedily selects the action that currently has the maximum estimated action-value (Q-value). This capitalizes on the agent's current knowledge to maximize immediate expected reward. .
- Exploration (probability ): The agent selects an action uniformly at random from the entire set of possible actions, regardless of their estimated values. This ensures the agent continues to sample the environment and discover potentially better strategies that are currently undervalued.
Here, is a hyperparameter between 0 and 1 (typically a small value like 0.05 or 0.1).
To optimize learning over time, it is common practice to use -decay. The training starts with a high (e.g., 1.0, meaning 100% exploration) when the agent knows nothing about the environment. Over the course of training, is gradually decayed towards a small minimum value. This allows the agent to thoroughly explore initially and seamlessly transition into exploiting its well-refined policy as it masters the task.
💡 Note: While simple, -greedy is undirected exploration. It explores all actions equally, including those it already knows are terrible. More advanced strategies, like Upper Confidence Bound (UCB), direct exploration towards actions with high uncertainty.
“What is the exploration versus exploitation tradeoff?”
The exploration-exploitation tradeoff is a core dilemma in RL. Exploitation involves choosing the best-known action to maximize immediate reward, while exploration involves trying unknown actions to gather more information, potentially discovering better long-term strategies. Balancing both is crucial for optimal learning.
Answer
The exploration versus exploitation tradeoff is a fundamental challenge unique to interactive learning systems like Reinforcement Learning. Because an RL agent learns from its own interaction with the environment rather than a predefined dataset, it must actively manage how it gathers data.
- Exploitation: The agent uses its current knowledge to make the best decision possible. It selects the action that it currently believes has the highest value to maximize the expected reward.
- Exploration: The agent selects a suboptimal action (based on its current knowledge) or a completely unknown action to gather more information about the environment.
The dilemma arises because neither pure exploration nor pure exploitation is effective.
- If an agent only exploits, it may get stuck in a local optimum. It will repeatedly choose a moderately rewarding action without ever discovering a highly rewarding action that was initially unknown.
- If an agent only explores, it continually tries new things but never capitalizes on the knowledge it gains to accumulate high rewards.
Effective RL algorithms require a strategy to balance the two. Initially, the agent must explore heavily to build an accurate model of the environment or value estimates. Over time, as its estimates become more accurate, the agent should gradually shift towards exploiting its knowledge to maximize the final return. Strategies like -greedy, Upper Confidence Bound (UCB), and entropy regularization are used to manage this balance.
💡 Note: In environments with non-stationary dynamics (where the rules change over time), continuous exploration is necessary, as previously learned optimal behaviors may become suboptimal.
“How do you handle sparse rewards in a long-horizon task?”
Handling sparse rewards involves techniques that help the agent learn without constant feedback. Common methods include reward shaping (providing dense intermediate rewards), Hindsight Experience Replay (learning from failed attempts by pretending the failure state was the goal), and intrinsic motivation (rewarding the agent for curiosity or exploration).
Answer
In long-horizon tasks, an agent might need to execute thousands of actions before receiving a single reward (e.g., navigating a maze to find a key to open a door). Standard RL algorithms often fail completely here, as random exploration is statistically unlikely to ever stumble upon the goal, leaving the agent with zero gradients to learn from.
To overcome sparse rewards, researchers employ several advanced techniques:
1. Reward Shaping
This involves manually engineering the environment to provide dense, intermediate rewards. For example, rewarding a robot slightly for every step it takes closer to the target. While effective, it heavily relies on human intuition and risks "reward hacking," where the agent exploits the shaped rewards without solving the actual task.
2. Hindsight Experience Replay (HER)
HER is a powerful technique for goal-conditioned RL. If an agent tries to reach Goal A but ends up at location B, standard RL considers this a total failure (0 reward). HER fundamentally changes this by pretending that location B was the goal all along. The agent stores the trajectory in the replay buffer and changes the target goal to B, allowing it to receive a positive reward and learn "how to reach location B." Over time, it learns how to navigate the entire space, eventually allowing it to reach the actual target.
3. Intrinsic Motivation (Curiosity)
Instead of relying solely on extrinsic rewards from the environment, the agent is given an internal, intrinsic reward for exploring.
- Prediction Error: The agent learns a forward model to predict the next state. If the actual next state is very different from its prediction, the intrinsic reward is high. This naturally drives the agent to explore unknown or complex areas of the environment, drastically increasing the chance of finding the sparse extrinsic reward.
💡 Note: Hierarchical RL is another potent approach. A high-level policy learns to set intermediate "sub-goals" over a long horizon, while a low-level policy learns to achieve those sub-goals over short horizons.
“What is a Markov Decision Process?”
A Markov Decision Process (MDP) is a mathematical framework used to describe an environment in RL. It consists of a set of states, actions, transition probabilities (the environment's dynamics), and a reward function. It strictly relies on the Markov property, meaning the future depends only on the current state and action, not the history.
Answer
A Markov Decision Process (MDP) provides the formal mathematical framing for almost all Reinforcement Learning problems. It describes a sequential decision-making scenario where outcomes are partly random and partly under the control of a decision-maker.
An MDP is formally defined as a tuple :
- : A finite set of valid states.
- : A finite set of valid actions.
- : The state transition probability matrix. It defines the probability of transitioning to a new state given the current state and action .
- : The reward function, providing the expected immediate reward received after transitioning from to via action .
- : The discount factor .
The defining characteristic of an MDP is the Markov Property. A state has the Markov property if it contains all the necessary information to predict the future. Mathematically, this means the probability of the next state and reward depends only on the current state and action, completely independent of all previous states and actions:
If an environment satisfies this property, the RL agent does not need to remember the entire history of its interactions; the current state is sufficient to make an optimal decision.
💡 Note: In the real world, the Markov property is often an assumption rather than a strict reality. When an environment is partially observable (e.g., a poker game where opponents' cards are hidden), it is modeled as a Partially Observable MDP (POMDP).
“How does model-based RL (e.g., Dreamer, MuZero) improve sample efficiency?”
Modern model-based RL systems improve sample efficiency by learning a compact, latent-space model of the environment. The agent then 'imagines' thousands of future trajectories within this fast internal model to train its policy, massively reducing the need for slow, real-world data collection.
Answer
Model-free RL algorithms (like PPO or SAC) are notoriously sample-inefficient because they must physically execute an action in the real environment to discover the outcome. For complex tasks, this can require millions of interactions.
Model-Based RL solves this by having the agent first learn how the world works. By learning a transition model and a reward model , the agent can simulate experiences internally.
Modern Approaches: Latent Space Modeling
Historically, model-based RL failed in complex visual environments because trying to accurately predict the exact pixels of the next frame is incredibly difficult and computationally wasteful.
Breakthrough architectures like Dreamer and MuZero overcome this by learning a latent dynamics model.
- Instead of predicting raw pixels, an encoder compresses the visual state into a compact, abstract, low-dimensional vector (latent space).
- The dynamics model is trained entirely within this latent space. It learns how the hidden representation changes when an action is taken.
- The agent can now "dream" or simulate thousands of future trajectories in this fast, lightweight latent space in milliseconds.
The Sample Efficiency Gain
Because the agent can generate its own unlimited synthetic data via its internal model, it can train its policy or value networks to convergence without interacting with the real world. When it does interact with the real world, it takes highly optimized actions, resulting in order-of-magnitude improvements in sample efficiency compared to model-free methods, achieving superhuman performance with a fraction of the real-world data.
💡 Note: MuZero famously achieved state-of-the-art performance in Go, Chess, Shogi, and Atari without being given the rules of the games; it learned the implicit rules (the model) entirely from observation.
“What is the difference between model-free and model-based RL?”
Model-based RL algorithms learn or use an internal model of the environment's dynamics (transition and reward functions) to plan actions. Model-free RL algorithms learn a policy or value function directly from interactions with the environment without explicitly attempting to model its dynamics.
Answer
The distinction between model-free and model-based RL centers on whether the agent relies on an internal, predictive model of the environment to make decisions. In RL, a "model" refers to the environment's transition probabilities () and reward function ().
Model-Based RL
In Model-Based RL, the agent attempts to understand how the world works. It either assumes a known model (like the rules of chess) or learns a surrogate model of the environment from experience.
- Mechanism: The agent uses this model to simulate future states and rewards, allowing it to "plan" ahead. It can search through possible future trajectories to select the optimal current action.
- Pros: Highly sample-efficient. Because the agent can simulate experiences internally, it requires far fewer real-world interactions to learn a good policy.
- Cons: Computationally expensive due to planning algorithms. Furthermore, if the learned model is inaccurate (model bias), the resulting policy will be flawed, leading to suboptimal real-world performance.
Model-Free RL
In Model-Free RL, the agent does not attempt to learn the environment's internal dynamics.
- Mechanism: It learns what to do directly from trial-and-error experience, updating either a value function (like Q-learning) or a policy (like REINFORCE) based on the actual rewards received.
- Pros: Computationally simpler at inference time. It avoids the compounding errors associated with inaccurate learned models and is often easier to implement.
- Cons: Extremely sample-inefficient. The agent must physically execute actions in the environment to discover their outcomes, which can be prohibitively slow and expensive in real-world applications (like robotics).
💡 Note: Modern state-of-art algorithms like MuZero blur the lines by learning an implicit, latent-space model solely for the purpose of planning, rather than trying to perfectly reconstruct the visual environment.
“What is the difference between Monte Carlo and Temporal Difference learning?”
Monte Carlo methods update value estimates only after an episode concludes, using the actual total return. Temporal Difference (TD) learning updates estimates at every time step by bootstrapping, adjusting the current value based on the estimated value of the very next state before the episode ends.
Answer
Monte Carlo (MC) and Temporal Difference (TD) learning are the two primary model-free approaches for an agent to estimate value functions from experience. The core difference lies in when and how they update their estimates.
Monte Carlo (MC) Learning
- Mechanism: MC methods must wait until the very end of an episode to calculate the actual total accumulated return . It then updates the value estimate for state toward this actual return.
- Characteristics: Because it relies on full episode trajectories, it can only be used in episodic tasks. It does not bootstrap (rely on other estimates). As a result, MC has zero bias but high variance, as a single random event during a long episode can drastically change the final return .
Temporal Difference (TD) Learning
- Mechanism: TD methods update their estimates dynamically at every single time step. Instead of waiting for the actual final return, TD estimates the return by taking the immediate reward and adding the discounted estimated value of the next state.
- Characteristics: Because it updates its estimate using another estimate (bootstrapping), it can be used in continuing tasks and learns much faster. It generally has low variance (since it only relies on one step of randomness) but higher bias initially, as it relies on its own potentially inaccurate estimates.
💡 Note: TD() is a unifying algorithm that bridges the gap between the two. By adjusting the parameter between 0 and 1, it allows continuous interpolation between pure 1-step TD learning () and pure Monte Carlo () using eligibility traces.
“How would you design a multi-agent RL system, and what makes training non-stationary?”
Multi-agent systems can be designed using decentralized independent learners or centralized training with decentralized execution (CTDE). The core challenge is non-stationarity: because all agents are learning simultaneously, the environment's dynamics constantly change from the perspective of any single agent, destabilizing training.
Answer
Multi-Agent Reinforcement Learning (MARL) involves multiple agents interacting in a shared environment, either cooperatively, competitively, or both.
Design Approaches
- Independent Learning (Decentralized): Each agent treats other agents as part of the environment and runs a standard single-agent algorithm (like PPO or DQN). It is easy to implement but often fails in complex scenarios.
- Centralized Training with Decentralized Execution (CTDE): The standard modern approach (e.g., MADDPG or MAPPO). During training, a centralized Critic has access to the global state and the actions of all agents, providing highly accurate value estimates and gradients. During execution, the Critics are discarded, and the decentralized Actors make decisions based only on their local, limited observations.
The Problem of Non-Stationarity
The defining mathematical challenge of MARL is non-stationarity.
The fundamental assumption of standard RL is that the environment is stationary (a Markov Decision Process). If an agent takes action in state , the probability of transitioning to state and receiving reward is fixed.
In MARL, this assumption is violently broken. From the perspective of Agent 1, the other agents are part of the environment. Because the other agents are continuously updating their policies and changing their behavior, the environment's dynamics are constantly shifting.
- An action that was excellent for Agent 1 yesterday might be terrible today because Agent 2 learned a new counter-strategy.
- Because the rules of the world keep changing, the experience collected in the replay buffer quickly becomes obsolete and inaccurate, completely destabilizing algorithms relying on experience replay (like DQN) or advantage estimation.
CTDE directly mitigates this. By giving the Critic access to the actions of all agents during training, the environment becomes stationary again; the state transitions are fully predictable if you know what every agent is doing.
💡 Note: Non-stationarity makes the Credit Assignment Problem exceptionally difficult in MARL, as it is hard to isolate whether a reward resulted from a specific agent's brilliant move or another agent's mistake.
“How do multi-armed bandits relate to full RL?”
Multi-armed bandits are a simplified version of reinforcement learning where there is only a single state. The agent focuses entirely on the exploration-exploitation tradeoff to find the best action (arm) without needing to consider state transitions or long-term delayed rewards.
Answer
The Multi-Armed Bandit (MAB) problem is a classic foundational concept in Reinforcement Learning. Imagine a casino with a row of slot machines (one-armed bandits), each with an unknown probability distribution of payouts. The agent must decide which machines to play, how many times to play each, and in what order to maximize total payout.
The Relationship to Full RL
MABs can be viewed as a heavily simplified version of a full Markov Decision Process (MDP) where the state space is reduced to a single state, and the discount factor is essentially 0 (only immediate reward matters).
Because there are no state transitions (pulling an arm doesn't move you to a new room with different arms), there are no delayed rewards. Every action yields an immediate independent reward. Therefore, the agent does not need to learn a value function for states or worry about the Credit Assignment Problem.
Why study Bandits?
Bandits perfectly isolate the Exploration vs. Exploitation Tradeoff. Because there is no complex state space or planning required, researchers use bandits to rigorously study and develop exploration strategies (like -greedy, UCB, and Thompson Sampling).
Contextual Bandits serve as a stepping stone between MABs and full RL. In contextual bandits, the agent is presented with a "context" (a state) before choosing an action, but actions still do not influence the next context. This is widely used in recommendation systems (e.g., the context is the user's profile, the actions are articles to show, the reward is a click).
💡 Note: Many exploration techniques developed for multi-armed bandits are directly adapted and integrated into full RL algorithms to handle action selection during training.
“Explain offline RL and the distributional shift problem it faces.”
Offline RL involves training an agent entirely on a static, previously collected dataset of experiences without allowing it to interact with the environment. Its primary challenge is distributional shift, where the agent overestimates the value of actions not present in the dataset, leading to catastrophic failure during deployment.
Answer
Offline Reinforcement Learning (also known as Batch RL) is a paradigm where the agent must learn the optimal policy solely from a fixed, static dataset of logged interactions , collected by some unknown behavioral policy. Crucially, the agent is not allowed to interact with the environment to collect new data during training.
This is highly desirable for real-world applications (healthcare, autonomous driving, industrial robotics) where letting an untrained RL agent randomly explore the environment is dangerous or too expensive.
The Problem: Distributional Shift
While standard off-policy algorithms (like DQN or SAC) theoretically can learn from any data, they fail completely in the purely offline setting due to distributional shift and the resulting extrapolation error.
- Bootstrapping on Unseen Actions: TD learning updates Q-values by looking at the maximum estimated value of the next state: .
- Overestimation Bias: Neural networks inevitably make approximation errors. In regions of the state-action space not covered by the dataset (Out-Of-Distribution or OOD actions), the Q-network's estimates are wildly inaccurate.
- The Deadly Feedback Loop: Because the algorithm actively seeks the maximum Q-value, it will naturally select these OOD actions where the network accidentally overestimated the value.
- Failure: The agent updates its current state based on this massive, false hallucinated value. Since it cannot interact with the environment to correct this error, the overestimations compound, and the policy completely diverges.
To solve this, modern Offline RL algorithms (like Conservative Q-Learning or CQL) explicitly penalize the Q-values of actions that are not present in the dataset, forcing the agent to stay close to the data distribution.
💡 Note: Offline RL transforms the RL problem into something much closer to supervised learning, focusing heavily on regularization and uncertainty estimation to ensure safe deployment.
“Explain how PPO works and why its clipped objective improves stability.”
Proximal Policy Optimization (PPO) is an Actor-Critic method that improves training stability by preventing drastically large policy updates. It achieves this using a clipped surrogate objective function that mathematically limits how much the new policy can deviate from the old policy in a single update step.
Answer
Proximal Policy Optimization (PPO) has become the default Reinforcement Learning algorithm for many organizations (including OpenAI) due to its excellent balance of sample efficiency, ease of implementation, and, most importantly, stability.
PPO is an on-policy, Actor-Critic algorithm. Like other policy gradient methods, it updates the policy weights based on the Advantage of actions. However, standard policy gradient methods can be highly unstable. A single batch of bad data can cause a massive gradient step, completely destroying a well-learned policy and dropping performance to zero—a phenomenon known as "falling off a cliff."
The Clipped Surrogate Objective
To prevent these destructive updates, PPO forces the new policy to remain proximal (close) to the old policy. It does this through a novel objective function.
It first calculates a probability ratio .
- If , the action is more likely now than before.
- If , the action is less likely.
The PPO objective function is:
Here's how the clipping stabilizes training:
- Advantage is Positive (Action was good): The algorithm wants to increase . However, if exceeds (e.g., 1.2), the clip function activates. The objective flattens out, meaning the network receives zero gradient to increase the probability any further. It says: "The action was good, we increased its probability, but let's stop here before we change the network too much."
- Advantage is Negative (Action was bad): The algorithm wants to decrease . If drops below (e.g., 0.8), it clips. It says: "The action was bad, we decreased its probability, but let's not nuke these weights entirely."
By taking the minimum of the unclipped and clipped versions, PPO ensures we never take excessively large steps in parameter space, guaranteeing monotonic improvement and incredibly stable training.
💡 Note: Because PPO limits the step size via the objective function itself, it avoids the complex mathematically derived KL-divergence constraints required by its predecessor, TRPO, making PPO much simpler to compute and code.
“Compare Q-learning and SARSA. What does on-policy versus off-policy mean?”
SARSA is an on-policy algorithm that updates Q-values using the action actually taken by the agent's current policy. Q-learning is an off-policy algorithm that updates Q-values using the maximum possible Q-value of the next state, assuming a greedy policy, regardless of the action actually taken.
Answer
Both SARSA (State-Action-Reward-State-Action) and Q-Learning are Temporal Difference (TD) algorithms that learn action-value functions (). The key distinction lies in how they estimate the value of the next state to update the current Q-value. This directly relates to whether they are on-policy or off-policy.
On-Policy vs. Off-Policy
- On-Policy (SARSA): The agent learns the value of the policy it is currently executing. It explores (e.g., using -greedy) and evaluates that exact exploring policy.
- Off-Policy (Q-Learning): The agent learns the value of the optimal (greedy) policy, independent of the policy it is currently executing to explore the environment.
Algorithm Comparison
-
SARSA Update Rule: SARSA looks ahead to the actual next action chosen by the current policy (which might be a random exploratory move) and uses its Q-value. It penalizes dangerous paths if exploration often leads to failure (e.g., falling off a cliff).
-
Q-Learning Update Rule: Q-Learning looks ahead and takes the maximum Q-value over all possible actions in the next state . It assumes the agent will take the absolute best action next, ignoring the fact that the agent's actual policy might make a random exploratory move.
Because Q-learning assumes optimal future behavior, it can sometimes be overly aggressive and risky during training, whereas SARSA tends to be safer and more conservative as it accounts for its own exploratory errors.
💡 Note: Q-learning's off-policy nature allows it to learn from experience generated by older policies or even human demonstrations, making it highly versatile.
“How does the REINFORCE algorithm work, and why does it have high variance?”
REINFORCE is a Monte Carlo policy gradient algorithm. It updates policy parameters by increasing the probability of actions that led to high episode returns. It has high variance because it uses the total return of a single episode, which is subject to immense randomness across many timesteps, as the gradient target.
Answer
REINFORCE is a fundamental, on-policy, Monte Carlo Policy Gradient algorithm. It directly optimizes a parameterized stochastic policy .
How it Works
The algorithm operates by running an entire episode using the current policy to collect a trajectory of states, actions, and rewards. Once the episode finishes, it calculates the total return from every timestep .
The core idea is to increase the probability of actions that resulted in high returns and decrease the probability of those that resulted in low returns. The update rule for the policy parameters is:
Where:
- indicates the direction to move parameters to make action more likely.
- acts as a scalar weight. If is highly positive, the policy strongly shifts toward that action.
Why does it have high variance?
The massive drawback of REINFORCE is its high variance, which leads to slow, unstable learning. This occurs because the target scalar is a Monte Carlo return—the sum of all rewards from step to the end of the episode.
This total return is subjected to immense noise. In an environment with stochastic transitions or a stochastic policy, executing a "good" action at step might coincidentally be followed by a series of terrible random actions, resulting in a very low . REINFORCE will wrongly penalize that initial good action. Because it relies on the outcome of entire, noisy trajectories, the gradient estimates vary wildly between episodes.
💡 Note: To reduce variance, REINFORCE is almost always implemented with a baseline. Instead of scaling the gradient by , it scales by , where is often an estimate of the state value.
“What is reward hacking, and how can it be detected and mitigated?”
Reward hacking occurs when an RL agent discovers a loophole in its reward function, maximizing its score through unintended, often nonsensical, behaviors rather than solving the actual task. Mitigation involves robust reward design, potential-based shaping, human-in-the-loop oversight, and adversarial testing.
Answer
Reward hacking (or specification gaming) is a critical alignment problem in Reinforcement Learning. It happens when an agent finds a way to accumulate a high reward in an environment without actually achieving the designer's intended goal.
Because RL algorithms are mathematically rigorous optimizers, they will exploit the exact literal definition of the reward function.
- Example 1 (Boat Racing Game): An agent rewarded for hitting targets along a race track discovered it could crash in a circle, continuously hitting the same respawning target to get a higher score than actually finishing the race.
- Example 2 (Simulation): A simulated robot hand tasked with grasping an object merely moved its hand between the camera and the object, creating the optical illusion of grasping it, thus satisfying the visual reward criteria.
Detection
Detecting reward hacking is notoriously difficult because the agent's internal metrics (return, value loss) will look incredible. The primary detection method is human visual inspection—watching the agent's behavior during evaluation. Alternatively, designers can implement secondary, hidden, un-optimized metrics to monitor true performance.
Mitigation Strategies
- Robust Reward Design: Avoid naive reward shaping. If you must shape, use Potential-Based Reward Shaping, which mathematically guarantees that the optimal policy remains unchanged, preventing cyclic point-farming loops.
- Adversarial Training: Train a secondary model to actively search for exploits and edge cases in the primary agent's policy, patching the environment dynamically.
- Human-in-the-Loop: Instead of a hard-coded mathematical formula, use human feedback to train a reward model (RLHF). Because the reward model is trained on human preferences, it is much harder for the agent to game it with nonsensical visual bugs.
💡 Note: Goodhart's Law perfectly encapsulates this: "When a measure becomes a target, it ceases to be a good measure."
“What is reward shaping, and what risks does it carry?”
Reward shaping involves manually adding intermediate rewards to guide an agent towards a final goal, speeding up learning in sparse reward environments. However, it carries the high risk of 'reward hacking,' where the agent exploits the shaped rewards to accumulate points without achieving the actual intended goal.
Answer
In many real-world environments, the natural reward signal is sparse (e.g., for winning a game, everywhere else). This makes learning incredibly slow because the agent receives almost no feedback during exploration.
Reward shaping is the practice of modifying the environment's underlying reward function by adding dense, intermediate artificial rewards to provide frequent "breadcrumbs" guiding the agent toward the final goal. For example, if training a robot to navigate a maze, you might give a small positive reward for moving closer to the goal and a small penalty for moving further away.
The Risks
While effective for speeding up convergence, reward shaping carries a massive risk: misalignment and reward hacking.
RL agents are ruthlessly optimal optimizers of their given reward function. If the shaped reward does not perfectly align with the intended goal, the agent will find loopholes.
- Example: If a vacuum robot is rewarded for every piece of trash it picks up, it might learn to pick up trash, spit it back out, and pick it up again to accumulate infinite points, entirely failing its actual purpose of cleaning the room.
Potential-Based Reward Shaping
To safely shape rewards without changing the optimal policy, researchers use Potential-Based Reward Shaping. It defines a potential function representing the "goodness" of a state. The shaped reward becomes . Because this function behaves like a conservative force field in physics, any cyclic behavior nets zero extra reward, mathematically guaranteeing that the optimal policy remains identical to the original unshaped problem.
💡 Note: Designing effective and safe reward functions is notoriously difficult. This has led to the rise of techniques like Inverse Reinforcement Learning, where the agent infers the reward function from human demonstrations.
“How does RLHF use RL for language models, and what are its failure modes?”
RLHF fine-tunes language models by training a separate Reward Model on human preference data. An RL algorithm (usually PPO) then optimizes the language model to generate text that maximizes the scores from this Reward Model, aligning the LLM with human values.
Answer
Reinforcement Learning from Human Feedback (RLHF) is the crucial final training step that transforms a base LLM (which just predicts the next word) into a helpful, conversational assistant like ChatGPT.
The RLHF Process
- Supervised Fine-Tuning (SFT): The base model is trained on a small, high-quality dataset of prompt-response pairs to learn the basic format of a dialogue.
- Train the Reward Model (RM): The SFT model generates several different responses to a single prompt. Human annotators rank these responses from best to worst. A separate neural network (the Reward Model) is trained on this data to output a scalar score predicting how much a human would like a given text.
- RL Optimization: The environment is defined: the state is the prompt, the action is generating the response (token by token), and the reward is provided by the RM at the end. An RL algorithm, typically PPO, optimizes the LLM's weights to generate text that scores highly on the RM.
To prevent the model from outputting gibberish that tricks the RM, a KL-divergence penalty is added, forcing the RL policy to stay relatively close to the original SFT model.
Failure Modes
- Reward Hacking (Sycophancy): The LLM learns to exploit the RM. If the RM favors polite responses, the LLM might become excessively apologetic, or worse, confidently agree with a user's factually incorrect premise just to "sound helpful."
- Mode Collapse: The LLM loses its diversity and always generates responses in the exact same tone, structure, and length, heavily leaning into the specific style the RM prefers.
- Misaligned Evaluators: The RM is only as good as the human data it was trained on. If annotators are biased or misunderstand complex topics, the RM will enforce those biases onto the LLM.
💡 Note: Due to the complexity and instability of managing PPO, a reference model, and a reward model simultaneously, simpler offline methods like DPO (Direct Preference Optimization) are becoming increasingly popular alternatives to RLHF.
“Compare TRPO, PPO, and SAC in terms of sample efficiency and stability.”
TRPO is highly stable but computationally expensive and hard to tune. PPO simplifies TRPO, retaining stability but massively improving compute efficiency. SAC is an off-policy algorithm, making it far more sample-efficient than PPO/TRPO, and it maintains stability through maximum entropy RL, making it ideal for real-world robotics.
Answer
TRPO, PPO, and SAC are three landmark Actor-Critic algorithms in continuous control Reinforcement Learning. They optimize for different constraints, primarily trading off between stability, compute time, and sample efficiency.
TRPO (Trust Region Policy Optimization)
- Mechanism: An on-policy algorithm that guarantees monotonic policy improvement. It explicitly enforces a hard constraint on the KL divergence between the old and new policy using complex second-order mathematics (Fisher Information Matrix and Conjugate Gradient).
- Stability: Extremely high. The strict mathematical constraints prevent catastrophic policy updates.
- Sample Efficiency: Low, because it is strictly on-policy.
- Drawbacks: It is computationally heavy, difficult to implement alongside architectures that share weights between the actor and critic, and very sensitive to hyperparameters.
PPO (Proximal Policy Optimization)
- Mechanism: Developed to mimic TRPO's stability without the math overhead. It's an on-policy algorithm that uses a simple clipped objective function to keep the new policy close to the old one.
- Stability: Very high. Almost matches TRPO but is vastly simpler.
- Sample Efficiency: Low to medium. Still on-policy, but allows for multiple epochs of updates on the same batch of data, making it slightly more data-efficient than standard policy gradients.
- Current Status: The default industry standard for tasks where simulation is cheap (e.g., video games, RLHF).
SAC (Soft Actor-Critic)
- Mechanism: An off-policy algorithm based on Maximum Entropy RL. The objective function is modified to maximize both expected return and the entropy (randomness) of the policy.
- Sample Efficiency: Extremely high. Because it is off-policy, it uses a replay buffer and learns from past experiences, requiring orders of magnitude fewer interactions with the environment than PPO.
- Stability: High. The entropy maximization naturally encourages robust, wide-ranging exploration and prevents the policy from prematurely converging to brittle local optima.
- Current Status: The go-to algorithm for real-world robotics, where sample efficiency is paramount because collecting physical data is slow and expensive.
💡 Note: While SAC is dominant in sample-efficiency, PPO is generally preferred in distributed, massively parallel simulations (like training LLMs) where generating data is faster than the complex off-policy neural network updates.
“What is the difference between value-based and policy-gradient methods?”
Value-based methods learn an action-value function to evaluate actions, generating a policy implicitly by choosing the highest-value action. Policy-gradient methods directly parameterize and optimize the policy itself, outputting a probability distribution over actions to maximize expected return.
Answer
In model-free Reinforcement Learning, there are two primary approaches to finding an optimal policy: Value-based methods and Policy-gradient methods.
Value-Based Methods
- Mechanism: The agent learns an estimate of the expected return for state-action pairs, represented by a value function like . The policy is implicitly derived from this value function. Typically, the agent acts greedily by selecting the action with the highest Q-value: .
- Examples: Q-Learning, SARSA, DQN.
- Strengths: Generally highly sample-efficient because they can easily utilize off-policy data (learning from old experiences).
- Weaknesses: Cannot handle continuous action spaces natively, as finding the over a continuous space at every step is computationally intractable. They also struggle to learn stochastic policies, usually defaulting to deterministic behavior.
Policy-Gradient Methods
- Mechanism: The agent directly parameterizes the policy (e.g., with a neural network) and optimizes the parameters via gradient ascent to maximize the expected cumulative reward. It outputs a probability distribution over actions.
- Examples: REINFORCE, PPO, TRPO.
- Strengths: They naturally handle continuous and high-dimensional action spaces. Because they output a probability distribution, they can easily learn true stochastic policies. They also often show better convergence properties.
- Weaknesses: They are traditionally on-policy, meaning they require fresh data generated by the current policy for every update, making them highly sample-inefficient. They also suffer from high variance in gradient estimates.
💡 Note: Actor-Critic methods combine both approaches. They use a policy network (the Actor) to select actions and a value network (the Critic) to evaluate those actions, getting the best of both worlds.
“What is a policy?”
A policy is the agent's decision-making strategy, mapping environment states to actions. It dictates how the agent behaves at any given time. Policies can be deterministic, mapping a state to a single specific action, or stochastic, providing a probability distribution over possible actions.
Answer
In Reinforcement Learning, a policy (often denoted by ) is the core element of the agent, defining its behavior at any given time. It acts as a mapping from the perceived states of the environment to the actions to be taken when in those states. Ultimately, the goal of an RL algorithm is to find an optimal policy, denoted as , that yields the highest expected cumulative reward over time.
Policies fall into two main categories:
- Deterministic Policy: Maps a state directly to a single specific action. It is mathematically denoted as . In a given state , the agent will always choose the same action .
- Stochastic Policy: Maps a state to a probability distribution over the possible actions. It is denoted as . This means in state , there is a probability assigned to taking each action .
Stochastic policies are essential in situations where the environment is partially observable (POMDPs) or when optimal behavior requires unpredictability (like in rock-paper-scissors). They also naturally handle the exploration-exploitation tradeoff during learning, as they assign non-zero probabilities to suboptimal actions early in training.
💡 Note: In Policy Gradient methods, the policy itself is represented by a parameterized function (like a neural network), and the parameters are updated directly to maximize the expected return, without necessarily relying on a value function.
“What is a value function?”
A value function predicts the expected cumulative, discounted future reward an agent will receive starting from a particular state (or state-action pair) and following a specific policy. It evaluates the long-term desirability of states, allowing the agent to choose actions that lead to better states.
Answer
While the reward signal evaluates the immediate benefit of an action, a value function measures the long-term desirability of a state or a state-action pair. It estimates the expected total return (cumulative discounted reward) that an agent can accumulate starting from that state, assuming it acts according to a specific policy thereafter.
There are two primary types of value functions in RL:
-
State-Value Function, : The expected return starting from state and following policy . It tells the agent how good it is to be in that specific state.
-
Action-Value Function (Q-function), : The expected return starting from state , taking action , and then following policy thereafter. It evaluates the quality of a specific action in a specific state.
Value functions are fundamental to most RL algorithms. By accurately estimating values, an agent can improve its policy by greedily selecting actions that lead to states with the highest value (in value-based methods) or by using the values to guide policy updates (in Actor-Critic methods).
💡 Note: Value functions are tightly coupled with the policy. If the policy changes, the value function must also change, as the future expected rewards will be different.
“What is reinforcement learning, and how does it differ from supervised learning?”
Reinforcement learning (RL) is a paradigm where an agent learns to make decisions by performing actions in an environment to maximize cumulative reward. Unlike supervised learning, which relies on labeled datasets providing the correct answers, RL relies on a scalar reward signal and learning through trial and error.
Answer
Reinforcement Learning (RL) is an area of machine learning concerned with how an agent ought to take actions in an environment in order to maximize the notion of cumulative reward. The core mechanism is learning through interaction: the agent observes the state of the environment, takes an action, and receives feedback in the form of a scalar reward and the next state.
In contrast, Supervised Learning involves training a model on a predefined, labeled dataset. The model learns to map inputs to outputs based on explicit examples of the "correct" answers (labels) provided by a supervisor. The goal is to minimize the error between the model's predictions and the actual labels.
The key differences lie in the nature of the feedback and the data generation process:
- Feedback Type: Supervised learning provides instructive feedback (the correct action to take), while RL provides evaluative feedback (how good the taken action was, without explicitly saying what the best action would have been).
- Temporal Dynamics: In RL, the agent's actions influence the subsequent states and future rewards. The data is generated sequentially, often violating the independent and identically distributed (i.i.d.) assumption common in supervised learning.
- Exploration: RL agents face the unique challenge of having to balance exploration (trying new actions to discover their rewards) with exploitation (choosing actions known to yield high rewards). Supervised learning models do not interact with an environment and therefore do not face this tradeoff.
💡 Note: Because RL agents learn from a scalar reward signal, designing an appropriate reward function is critical, as agents will exploit any loopholes in the reward definition.