Hindsight Experience Replay (HER)
Even when an agent misses its intended objective, it successfully achieved something; by pretending in hindsight that what it reached was the target all along, it extracts valuable learning signals from failures in sparse-reward environments.
Why Does This Exist?
In multi-goal reinforcement learning tasks—such as a robotic arm learning to place a block at an arbitrary 3D coordinate—rewards are naturally binary and sparse. The agent receives a reward of when the object settles within a tight threshold of the target goal , and at every other time step. Under standard exploration, an untrained agent executing random motor babbling has a vanishingly small probability of reaching a distant target coordinate by chance.
Consequently, almost every rollout trajectory finishes with complete failure: the replay buffer fills with transitions where across all steps. Because the value landscape is completely flat, temporal-difference updates have no non-zero reward signals to bootstrap from. Q-learning stalls indefinitely because all actions appear equally futile.
The conventional workaround was reward shaping: replacing binary rewards with negative Euclidean distances . However, engineered distance rewards introduce severe pitfalls:
- They create unintended local optima (for example, hovering near an obstacle rather than reaching around it).
- They alter the optimal policy objective.
- They fail in non-Euclidean state spaces (such as maze navigation or complex manipulation).
Hindsight Experience Replay (HER), introduced by Marcin Andrychowicz et al. (2017), solves this bottleneck without modifying the true reward function. HER recognizes that every failed trajectory to target goal is simultaneously a perfect, successful trajectory to whichever state the agent actually managed to visit. By relabeling the transitions in hindsight with achieved states, the agent manufactures guaranteed success signals () out of complete failures.
Think of It Like This
The Archery Student and the Missed Target
Imagine a novice archer taking aim at a distant red bullseye on Target A. The archer draws the bow, releases, and the arrow flies wide, lodging deeply into the bark of an old oak tree thirty yards to the left.
Under strict tournament scoring, this shot is an absolute failure: zero points. If the archer discards the attempt as useless waste, they learn almost nothing about how their muscles, release velocity, and stance influenced arrow trajectory.
Now imagine a perceptive coach standing nearby who says: "Forget Target A for a moment. Pretend your intention all along was to hit that exact oak tree. Your stance, draw tension, and release angle were an absolute masterclass in striking that tree trunk! Let us record that technique into your muscle memory."
Once the student logs thousands of relabeled shots—learning how to deliberately strike oak trees, fence posts, rocks, and dirt mounds—they develop a comprehensive mental map of bow mechanics across the entire field. When asked once more to hit Target A, they simply consult their repertoire: Target A is merely another point in space they now know how to reach.
The analogy stops when considering physics: in real archery, an arrow's trajectory does not change based on what you wished to hit. Similarly, in reinforcement learning, the environment's transition dynamics are strictly independent of the goal vector . This physical invariance is precisely what allows an agent to retroactively change the goal label in its replay memory without invalidating transition physics.
How It Actually Works
Goal-Conditioned Formulation and Universal Value Functions
HER operates within the framework of Goal-Conditioned Markov Decision Processes (Universal Value Function Approximators, or UVFAs). The state space is paired with a goal space , and there is a mapping that extracts the achieved goal from an environmental state.
The policy and action-value function take both the current state and the desired goal as inputs. The environment provides a sparse goal-conditioned reward:
where is the next state and is a task tolerance threshold.
The Goal Relabeling Mechanism
During standard rollouts, an agent is assigned a desired goal and generates an episode trajectory:
Under standard replay storage, all transitions are stored in replay buffer using original goal . If was never reached, for all .
HER augments the buffer by creating additional hindsight copies of each transition . For each transition at step , HER selects a set of alternative goals from states visited during the episode, recomputes the reward under , and stores the augmented tuple:
Andrychowicz et al. evaluated four primary goal-sampling strategies:
final: Relabel using exclusively the final achieved state of the rollout: .future: For transition at time step , randomly sample alternative goals from states visited later in the same episode (). This is the empirical gold standard ().episode: Sample alternative goals randomly from any time step across the current episode.random: Sample alternative goals uniformly from arbitrary states encountered in prior episodes.
The future strategy outperforms all others because it preserves temporal coherence: it samples goals that were physically reachable within the remaining time horizon of that trajectory.
Why Off-Policy RL is Mandatory
HER is fundamentally incompatible with on-policy reinforcement learning algorithms like PPO or A2C. An on-policy algorithm requires that the actions were generated directly by the current policy conditioned on the exact evaluated goal: . Because the trajectory was collected under the original goal , conditioning post-hoc on constitutes off-policy distribution shift.
Off-policy algorithms (such as DDPG, TD3, or SAC) evaluate transitions using the Bellman optimality equation:
Because the physical environment transition is invariant to the goal argument, the transition remains completely valid under any goal . The target network simply evaluates the value of reaching from .
Worked numerical example
Consider a 1D continuous positioning task. An agent starts at coordinate and is tasked with reaching with tolerance . The discount factor is and the learning rate is .
The reward function is:
The agent executes constant velocity actions for steps:
- Step 0: . Distance .
- Step 1: . Distance .
- Step 2: . Distance .
1. Standard Experience Replay Failure
Every transition in the buffer stores . If prior values are uniformly initialized to , every Bellman target evaluates to . The agent receives zero directional signal explaining how to navigate towards .
2. HER Goal Relabeling
We apply HER using the final strategy: we relabel the rollout with achieved goal .
-
Relabeled Step 2 (, goal ):
- Check success condition: .
- State satisfies the goal, marking successful termination with respect to . Terminal value .
- Bellman target:
- Initial Q-estimate: .
- TD Error:
- Updated Q-value:
-
Relabeled Step 1 (, goal ):
- Check success condition: .
- Next state is not terminal. It bootstraps from the newly updated downstream Q-value:
- Initial Q-estimate: .
- TD Error:
- Updated Q-value:
Notice the transformation: the single hindsight success at immediately propagated a positive gradient to Step 2 and a gradient to Step 1. The agent has acquired concrete knowledge of how to reach coordinate . As multiple rollouts achieve varied endpoints, the Universal Value Function interpolates across the goal space, unlocking rapid convergence toward the true target .
Code
The following self-contained implementation demonstrates a complete HindsightExperienceReplayBuffer supporting both future and final relabeling strategies alongside original trajectory storage.
from dataclasses import dataclassfrom typing import Callable, List, Optional, Tupleimport numpy as np
@dataclassclass Transition: state: np.ndarray action: np.ndarray reward: float next_state: np.ndarray done: bool goal: np.ndarray
class HindsightExperienceReplayBuffer: """ Experience Replay Buffer implementing Hindsight Experience Replay (HER). Stores rollout episodes and creates hindsight-relabeled transitions. """
def __init__( self, capacity: int = 10_000, k_future: int = 4, strategy: str = "future", reward_fn: Optional[Callable[[np.ndarray, np.ndarray], float]] = None, ) -> None: self.capacity = capacity self.k_future = k_future self.strategy = strategy self.reward_fn = reward_fn or self._default_sparse_reward self.buffer: List[Transition] = [] self.position = 0
@staticmethod def _default_sparse_reward( achieved_state: np.ndarray, goal: np.ndarray, threshold: float = 0.5 ) -> float: """Sparse binary reward: 0.0 if within threshold distance, else -1.0.""" distance = float(np.linalg.norm(achieved_state - goal)) return 0.0 if distance <= threshold else -1.0
def add_episode( self, trajectory: List[Tuple[np.ndarray, np.ndarray, np.ndarray, bool, np.ndarray]], ) -> None: """ Stores an entire episode trajectory and augments it with HER relabeling.
trajectory: List of (state, action, next_state, done, original_goal) tuples. """ episode_length = len(trajectory) if episode_length == 0: return
for t in range(episode_length): state, action, next_state, done, original_goal = trajectory[t]
# 1. Store the original transition under original goal orig_reward = self.reward_fn(next_state, original_goal) self._store( Transition( state=state, action=action, reward=orig_reward, next_state=next_state, done=done, goal=original_goal, ) )
# 2. HER Relabeling if self.strategy == "future": # Sample k states observed after step t in the current episode candidate_indices = list(range(t, episode_length)) sampled_indices = np.random.choice( candidate_indices, size=self.k_future, replace=True ) for idx in sampled_indices: achieved_goal = trajectory[idx][2] # next_state of step idx relabelled_reward = self.reward_fn(next_state, achieved_goal) relabelled_done = bool(relabelled_reward == 0.0)
self._store( Transition( state=state, action=action, reward=relabelled_reward, next_state=next_state, done=relabelled_done, goal=achieved_goal, ) )
elif self.strategy == "final": # Relabel using strictly the final state of the episode final_achieved_goal = trajectory[-1][2] relabelled_reward = self.reward_fn(next_state, final_achieved_goal) relabelled_done = bool(relabelled_reward == 0.0)
self._store( Transition( state=state, action=action, reward=relabelled_reward, next_state=next_state, done=relabelled_done, goal=final_achieved_goal, ) )
def _store(self, transition: Transition) -> None: if len(self.buffer) < self.capacity: self.buffer.append(transition) else: self.buffer[self.position] = transition self.position = (self.position + 1) % self.capacity
def sample_batch(self, batch_size: int) -> List[Transition]: indices = np.random.choice(len(self.buffer), size=batch_size, replace=False) return [self.buffer[i] for i in indices]
def __len__(self) -> int: return len(self.buffer)
# Demonstrationif __name__ == "__main__": np.random.seed(42)
her_buffer = HindsightExperienceReplayBuffer( capacity=1000, k_future=4, strategy="future" )
# Simulated rollout attempting to reach distant goal [10.0] target_goal = np.array([10.0]) rollout = [ (np.array([0.0]), np.array([2.0]), np.array([2.0]), False, target_goal), (np.array([2.0]), np.array([2.0]), np.array([4.0]), False, target_goal), (np.array([4.0]), np.array([2.0]), np.array([6.0]), True, target_goal), ]
her_buffer.add_episode(rollout)
orig_successes = sum( 1 for tr in her_buffer.buffer if np.array_equal(tr.goal, target_goal) and tr.reward == 0.0 ) total_successes = sum(1 for tr in her_buffer.buffer if tr.reward == 0.0)
print(f"Total stored transitions: {len(her_buffer)}") print(f"Successes under target goal [10.0]: {orig_successes}") print(f"Hindsight relabeled successes (r=0.0): {total_successes}")
sample = her_buffer.sample_batch(batch_size=3) for i, tr in enumerate(sample): print( f"Sample {i}: s={tr.state[0]:.1f} -> s'={tr.next_state[0]:.1f} | " f"goal={tr.goal[0]:.1f} | reward={tr.reward:.1f}" )Output:
Total stored transitions: 15Successes under target goal [10.0]: 0Hindsight relabeled successes (r=0.0): 8Sample 0: s=0.0 -> s'=2.0 | goal=10.0 | reward=-1.0Sample 1: s=4.0 -> s'=6.0 | goal=6.0 | reward=0.0Sample 2: s=4.0 -> s'=6.0 | goal=6.0 | reward=0.0Watch Out For
Sampling Over-Saturation: The k-Ratio Imbalance
A frequent trap when implementing HER is setting the hindsight replay parameter excessively high (for example, or ) in hopes of maximizing positive reinforcement.
When is too large, the replay buffer becomes overwhelmingly saturated with one-step transitions where the relabeled goal was reached almost instantly (). The policy rapidly learns to stabilize locally around trivial nearby states but suffers from goal myopia: it never learns the extended temporal sequences necessary to cross long distances toward the original training goals. Conversely, setting provides too sparse a hindsight signal, slowing down exploration.
The Fix: Stick to the empirical ratio established in the seminal paper: for standard manipulation tasks (yielding an approximate ratio of 4 hindsight transitions for every 1 original transition, or an 80/20 mix during mini-batch sampling). If training on very long-horizon tasks, combine HER with prioritized experience replay (PER) so transitions with high temporal-difference error are sampled preferentially over trivial one-step arrivals.
The Quick Version
- Core Intuition: In sparse-reward multi-goal RL, failed trajectories to target goal are successful trajectories to the states actually reached; relabeling them in hindsight extracts rich learning signals from failures.
- Goal Relabeling Modes: The
futurestrategy (sampling states visited later in the same episode) outperformsfinal,episode, andrandombecause it respects the reachable time horizon. - Off-Policy Synergy: HER strictly requires off-policy reinforcement learning (such as DDPG, TD3, or SAC) because changing the goal vector post-rollout introduces an off-policy distribution shift that invalidates on-policy methods like PPO.
- Reward Shaping Alternative: HER solves the sparse-reward exploration plateau without the distortion, local optima, and task-specific engineering traps of hand-crafted distance rewards.