Skip to content
AI360Xpert
Beta

Hindsight Experience Replay (HER)

Even when an agent misses its intended objective, it successfully achieved something; by pretending in hindsight that what it reached was the target all along, it extracts valuable learning signals from failures in sparse-reward environments.

Hindsight Experience Replay converts unrewarded trajectories into successful demonstrations by substituting the original goal with actually achieved states.
Hindsight Experience Replay converts unrewarded trajectories into successful demonstrations by substituting the original goal with actually achieved states.

Why Does This Exist?

In multi-goal reinforcement learning tasks—such as a robotic arm learning to place a block at an arbitrary 3D coordinate—rewards are naturally binary and sparse. The agent receives a reward of 00 when the object settles within a tight threshold ϵ\epsilon of the target goal gg, and −1-1 at every other time step. Under standard exploration, an untrained agent executing random motor babbling has a vanishingly small probability of reaching a distant target coordinate by chance.

Consequently, almost every rollout trajectory finishes with complete failure: the replay buffer fills with transitions where r=−1r = -1 across all steps. Because the value landscape is completely flat, temporal-difference updates have no non-zero reward signals to bootstrap from. Q-learning stalls indefinitely because all actions appear equally futile.

The conventional workaround was reward shaping: replacing binary rewards with negative Euclidean distances ∥s−g∥\|s - g\|. However, engineered distance rewards introduce severe pitfalls:

  1. They create unintended local optima (for example, hovering near an obstacle rather than reaching around it).
  2. They alter the optimal policy objective.
  3. They fail in non-Euclidean state spaces (such as maze navigation or complex manipulation).

Hindsight Experience Replay (HER), introduced by Marcin Andrychowicz et al. (2017), solves this bottleneck without modifying the true reward function. HER recognizes that every failed trajectory to target goal gg is simultaneously a perfect, successful trajectory to whichever state sTs_T the agent actually managed to visit. By relabeling the transitions in hindsight with achieved states, the agent manufactures guaranteed success signals (r=0r = 0) out of complete failures.

Think of It Like This

The Archery Student and the Missed Target

Imagine a novice archer taking aim at a distant red bullseye on Target A. The archer draws the bow, releases, and the arrow flies wide, lodging deeply into the bark of an old oak tree thirty yards to the left.

Under strict tournament scoring, this shot is an absolute failure: zero points. If the archer discards the attempt as useless waste, they learn almost nothing about how their muscles, release velocity, and stance influenced arrow trajectory.

Now imagine a perceptive coach standing nearby who says: "Forget Target A for a moment. Pretend your intention all along was to hit that exact oak tree. Your stance, draw tension, and release angle were an absolute masterclass in striking that tree trunk! Let us record that technique into your muscle memory."

Once the student logs thousands of relabeled shots—learning how to deliberately strike oak trees, fence posts, rocks, and dirt mounds—they develop a comprehensive mental map of bow mechanics across the entire field. When asked once more to hit Target A, they simply consult their repertoire: Target A is merely another point in space they now know how to reach.

The analogy stops when considering physics: in real archery, an arrow's trajectory does not change based on what you wished to hit. Similarly, in reinforcement learning, the environment's transition dynamics P(st+1∣st,at)P(s_{t+1} \mid s_t, a_t) are strictly independent of the goal vector gg. This physical invariance is precisely what allows an agent to retroactively change the goal label in its replay memory without invalidating transition physics.

How It Actually Works

Goal-Conditioned Formulation and Universal Value Functions

HER operates within the framework of Goal-Conditioned Markov Decision Processes (Universal Value Function Approximators, or UVFAs). The state space S\mathcal{S} is paired with a goal space G\mathcal{G}, and there is a mapping m:S→Gm: \mathcal{S} \to \mathcal{G} that extracts the achieved goal from an environmental state.

The policy π(a∣s,g)\pi(a \mid s, g) and action-value function Q(s,a,g)Q(s, a, g) take both the current state ss and the desired goal gg as inputs. The environment provides a sparse goal-conditioned reward:

r(s,a,g)={0if ∥m(s′)−g∥≤ϵ−1otherwiser(s, a, g) = \begin{cases} 0 & \text{if } \|m(s') - g\| \le \epsilon \\ -1 & \text{otherwise} \end{cases}

where s′s' is the next state and ϵ\epsilon is a task tolerance threshold.

The Goal Relabeling Mechanism

During standard rollouts, an agent is assigned a desired goal gg and generates an episode trajectory:

τ=((s0,a0,r0,s1,g),(s1,a1,r1,s2,g),…,(sT−1,aT−1,rT−1,sT,g))\tau = \left( (s_0, a_0, r_0, s_1, g), (s_1, a_1, r_1, s_2, g), \dots, (s_{T-1}, a_{T-1}, r_{T-1}, s_T, g) \right)

Under standard replay storage, all TT transitions are stored in replay buffer D\mathcal{D} using original goal gg. If gg was never reached, rt=−1r_t = -1 for all t∈[0,T−1]t \in [0, T-1].

HER augments the buffer by creating additional hindsight copies of each transition (st,at,st+1)(s_t, a_t, s_{t+1}). For each transition at step tt, HER selects a set of alternative goals G′G' from states visited during the episode, recomputes the reward under g′∈G′g' \in G', and stores the augmented tuple:

(st,at,r′,st+1,g′)wherer′=r(st+1,at,g′)(s_t, a_t, r', s_{t+1}, g') \quad \text{where} \quad r' = r(s_{t+1}, a_t, g')

Andrychowicz et al. evaluated four primary goal-sampling strategies:

  1. final: Relabel using exclusively the final achieved state of the rollout: g′=m(sT)g' = m(s_T).
  2. future: For transition at time step tt, randomly sample kk alternative goals from states visited later in the same episode (t′∈[t+1,T]t' \in [t+1, T]). This is the empirical gold standard (k≈4k \approx 4).
  3. episode: Sample kk alternative goals randomly from any time step across the current episode.
  4. random: Sample kk alternative goals uniformly from arbitrary states encountered in prior episodes.

The future strategy outperforms all others because it preserves temporal coherence: it samples goals that were physically reachable within the remaining time horizon of that trajectory.

Why Off-Policy RL is Mandatory

HER is fundamentally incompatible with on-policy reinforcement learning algorithms like PPO or A2C. An on-policy algorithm requires that the actions ata_t were generated directly by the current policy conditioned on the exact evaluated goal: at∼π(⋅∣st,g′)a_t \sim \pi(\cdot \mid s_t, g'). Because the trajectory was collected under the original goal gg, conditioning post-hoc on g′g' constitutes off-policy distribution shift.

Off-policy algorithms (such as DDPG, TD3, or SAC) evaluate transitions using the Bellman optimality equation:

yt=r′+γQtarget(st+1,π(st+1,g′),g′)y_t = r' + \gamma Q_{\text{target}}\left(s_{t+1}, \pi(s_{t+1}, g'), g'\right)

Because the physical environment transition P(st+1∣st,at)P(s_{t+1} \mid s_t, a_t) is invariant to the goal argument, the transition (st,at,st+1)(s_t, a_t, s_{t+1}) remains completely valid under any goal g′g'. The target network simply evaluates the value of reaching g′g' from st+1s_{t+1}.

Worked numerical example

Consider a 1D continuous positioning task. An agent starts at coordinate s0=0.0s_0 = 0.0 and is tasked with reaching gtarget=10.0g_{\text{target}} = 10.0 with tolerance ϵ=0.5\epsilon = 0.5. The discount factor is γ=0.90\gamma = 0.90 and the learning rate is α=0.20\alpha = 0.20.

The reward function is:

r(s′,g)={0.0if ∣s′−g∣≤0.5−1.0otherwiser(s', g) = \begin{cases} 0.0 & \text{if } |s' - g| \le 0.5 \\ -1.0 & \text{otherwise} \end{cases}

The agent executes constant velocity actions at=+2.0a_t = +2.0 for T=3T = 3 steps:

  • Step 0: s0=0.0→a0=2.0s1=2.0s_0 = 0.0 \xrightarrow{a_0=2.0} s_1 = 2.0. Distance ∣2.0−10.0∣=8.0>0.5  ⟹  r0=−1.0|2.0 - 10.0| = 8.0 > 0.5 \implies r_0 = -1.0.
  • Step 1: s1=2.0→a1=2.0s2=4.0s_1 = 2.0 \xrightarrow{a_1=2.0} s_2 = 4.0. Distance ∣4.0−10.0∣=6.0>0.5  ⟹  r1=−1.0|4.0 - 10.0| = 6.0 > 0.5 \implies r_1 = -1.0.
  • Step 2: s2=4.0→a2=2.0s3=6.0s_2 = 4.0 \xrightarrow{a_2=2.0} s_3 = 6.0. Distance ∣6.0−10.0∣=4.0>0.5  ⟹  r2=−1.0|6.0 - 10.0| = 4.0 > 0.5 \implies r_2 = -1.0.

1. Standard Experience Replay Failure

Every transition in the buffer stores r=−1.0r = -1.0. If prior values are uniformly initialized to −5.0-5.0, every Bellman target evaluates to −1.0+0.90(−5.0)=−5.50-1.0 + 0.90(-5.0) = -5.50. The agent receives zero directional signal explaining how to navigate towards 10.010.0.

2. HER Goal Relabeling

We apply HER using the final strategy: we relabel the rollout with achieved goal g′=s3=6.0g' = s_3 = 6.0.

  • Relabeled Step 2 (s2=4.0→a2=2.0s3=6.0s_2 = 4.0 \xrightarrow{a_2=2.0} s_3 = 6.0, goal g′=6.0g'=6.0):

    • Check success condition: ∣s3−g′∣=∣6.0−6.0∣=0.0≤0.5  ⟹  r2′=0.0|s_3 - g'| = |6.0 - 6.0| = 0.0 \le 0.5 \implies r'_2 = 0.0.
    • State s3s_3 satisfies the goal, marking successful termination with respect to g′g'. Terminal value V(s3,g′)=0.0V(s_3, g') = 0.0.
    • Bellman target: y2=r2′+γ⋅0.0=0.0+0.0=0.0y_2 = r'_2 + \gamma \cdot 0.0 = 0.0 + 0.0 = 0.0
    • Initial Q-estimate: Q(s2=4.0,a2=2.0,g′=6.0)=−2.50Q(s_2=4.0, a_2=2.0, g'=6.0) = -2.50.
    • TD Error: δ2=y2−Q=0.0−(−2.50)=+2.50\delta_2 = y_2 - Q = 0.0 - (-2.50) = +2.50
    • Updated Q-value: Qnew(s2,a2,g′)=−2.50+0.20×(+2.50)=−2.00Q_{\text{new}}(s_2, a_2, g') = -2.50 + 0.20 \times (+2.50) = -2.00
  • Relabeled Step 1 (s1=2.0→a1=2.0s2=4.0s_1 = 2.0 \xrightarrow{a_1=2.0} s_2 = 4.0, goal g′=6.0g'=6.0):

    • Check success condition: ∣s2−g′∣=∣4.0−6.0∣=2.0>0.5  ⟹  r1′=−1.0|s_2 - g'| = |4.0 - 6.0| = 2.0 > 0.5 \implies r'_1 = -1.0.
    • Next state s2s_2 is not terminal. It bootstraps from the newly updated downstream Q-value: y1=r1′+γQnew(s2,a2,g′)=−1.0+0.90×(−2.00)=−1.0−1.80=−2.80y_1 = r'_1 + \gamma Q_{\text{new}}(s_2, a_2, g') = -1.0 + 0.90 \times (-2.00) = -1.0 - 1.80 = -2.80
    • Initial Q-estimate: Q(s1=2.0,a1=2.0,g′=6.0)=−4.00Q(s_1=2.0, a_1=2.0, g'=6.0) = -4.00.
    • TD Error: δ1=−2.80−(−4.00)=+1.20\delta_1 = -2.80 - (-4.00) = +1.20
    • Updated Q-value: Qnew(s1,a1,g′)=−4.00+0.20×(+1.20)=−3.76Q_{\text{new}}(s_1, a_1, g') = -4.00 + 0.20 \times (+1.20) = -3.76

Notice the transformation: the single hindsight success at s3s_3 immediately propagated a positive +2.50+2.50 gradient to Step 2 and a +1.20+1.20 gradient to Step 1. The agent has acquired concrete knowledge of how to reach coordinate 6.06.0. As multiple rollouts achieve varied endpoints, the Universal Value Function interpolates across the goal space, unlocking rapid convergence toward the true target 10.010.0.

Code

The following self-contained implementation demonstrates a complete HindsightExperienceReplayBuffer supporting both future and final relabeling strategies alongside original trajectory storage.

from dataclasses import dataclassfrom typing import Callable, List, Optional, Tupleimport numpy as np

@dataclassclass Transition:    state: np.ndarray    action: np.ndarray    reward: float    next_state: np.ndarray    done: bool    goal: np.ndarray

class HindsightExperienceReplayBuffer:    """    Experience Replay Buffer implementing Hindsight Experience Replay (HER).    Stores rollout episodes and creates hindsight-relabeled transitions.    """
    def __init__(        self,        capacity: int = 10_000,        k_future: int = 4,        strategy: str = "future",        reward_fn: Optional[Callable[[np.ndarray, np.ndarray], float]] = None,    ) -> None:        self.capacity = capacity        self.k_future = k_future        self.strategy = strategy        self.reward_fn = reward_fn or self._default_sparse_reward        self.buffer: List[Transition] = []        self.position = 0
    @staticmethod    def _default_sparse_reward(        achieved_state: np.ndarray, goal: np.ndarray, threshold: float = 0.5    ) -> float:        """Sparse binary reward: 0.0 if within threshold distance, else -1.0."""        distance = float(np.linalg.norm(achieved_state - goal))        return 0.0 if distance <= threshold else -1.0
    def add_episode(        self,        trajectory: List[Tuple[np.ndarray, np.ndarray, np.ndarray, bool, np.ndarray]],    ) -> None:        """        Stores an entire episode trajectory and augments it with HER relabeling.
        trajectory: List of (state, action, next_state, done, original_goal) tuples.        """        episode_length = len(trajectory)        if episode_length == 0:            return
        for t in range(episode_length):            state, action, next_state, done, original_goal = trajectory[t]
            # 1. Store the original transition under original goal            orig_reward = self.reward_fn(next_state, original_goal)            self._store(                Transition(                    state=state,                    action=action,                    reward=orig_reward,                    next_state=next_state,                    done=done,                    goal=original_goal,                )            )
            # 2. HER Relabeling            if self.strategy == "future":                # Sample k states observed after step t in the current episode                candidate_indices = list(range(t, episode_length))                sampled_indices = np.random.choice(                    candidate_indices, size=self.k_future, replace=True                )                for idx in sampled_indices:                    achieved_goal = trajectory[idx][2]  # next_state of step idx                    relabelled_reward = self.reward_fn(next_state, achieved_goal)                    relabelled_done = bool(relabelled_reward == 0.0)
                    self._store(                        Transition(                            state=state,                            action=action,                            reward=relabelled_reward,                            next_state=next_state,                            done=relabelled_done,                            goal=achieved_goal,                        )                    )
            elif self.strategy == "final":                # Relabel using strictly the final state of the episode                final_achieved_goal = trajectory[-1][2]                relabelled_reward = self.reward_fn(next_state, final_achieved_goal)                relabelled_done = bool(relabelled_reward == 0.0)
                self._store(                    Transition(                        state=state,                        action=action,                        reward=relabelled_reward,                        next_state=next_state,                        done=relabelled_done,                        goal=final_achieved_goal,                    )                )
    def _store(self, transition: Transition) -> None:        if len(self.buffer) < self.capacity:            self.buffer.append(transition)        else:            self.buffer[self.position] = transition        self.position = (self.position + 1) % self.capacity
    def sample_batch(self, batch_size: int) -> List[Transition]:        indices = np.random.choice(len(self.buffer), size=batch_size, replace=False)        return [self.buffer[i] for i in indices]
    def __len__(self) -> int:        return len(self.buffer)

# Demonstrationif __name__ == "__main__":    np.random.seed(42)
    her_buffer = HindsightExperienceReplayBuffer(        capacity=1000, k_future=4, strategy="future"    )
    # Simulated rollout attempting to reach distant goal [10.0]    target_goal = np.array([10.0])    rollout = [        (np.array([0.0]), np.array([2.0]), np.array([2.0]), False, target_goal),        (np.array([2.0]), np.array([2.0]), np.array([4.0]), False, target_goal),        (np.array([4.0]), np.array([2.0]), np.array([6.0]), True, target_goal),    ]
    her_buffer.add_episode(rollout)
    orig_successes = sum(        1 for tr in her_buffer.buffer if np.array_equal(tr.goal, target_goal) and tr.reward == 0.0    )    total_successes = sum(1 for tr in her_buffer.buffer if tr.reward == 0.0)
    print(f"Total stored transitions: {len(her_buffer)}")    print(f"Successes under target goal [10.0]: {orig_successes}")    print(f"Hindsight relabeled successes (r=0.0): {total_successes}")
    sample = her_buffer.sample_batch(batch_size=3)    for i, tr in enumerate(sample):        print(            f"Sample {i}: s={tr.state[0]:.1f} -> s'={tr.next_state[0]:.1f} | "            f"goal={tr.goal[0]:.1f} | reward={tr.reward:.1f}"        )

Output:

Total stored transitions: 15Successes under target goal [10.0]: 0Hindsight relabeled successes (r=0.0): 8Sample 0: s=0.0 -> s'=2.0 | goal=10.0 | reward=-1.0Sample 1: s=4.0 -> s'=6.0 | goal=6.0 | reward=0.0Sample 2: s=4.0 -> s'=6.0 | goal=6.0 | reward=0.0

Watch Out For

Sampling Over-Saturation: The k-Ratio Imbalance

A frequent trap when implementing HER is setting the hindsight replay parameter kk excessively high (for example, k=16k = 16 or k=32k = 32) in hopes of maximizing positive reinforcement.

When kk is too large, the replay buffer becomes overwhelmingly saturated with one-step transitions where the relabeled goal was reached almost instantly (st+1≈g′s_{t+1} \approx g'). The policy rapidly learns to stabilize locally around trivial nearby states but suffers from goal myopia: it never learns the extended temporal sequences necessary to cross long distances toward the original training goals. Conversely, setting k=1k = 1 provides too sparse a hindsight signal, slowing down exploration.

The Fix: Stick to the empirical ratio established in the seminal paper: k≈4k \approx 4 for standard manipulation tasks (yielding an approximate ratio of 4 hindsight transitions for every 1 original transition, or an 80/20 mix during mini-batch sampling). If training on very long-horizon tasks, combine HER with prioritized experience replay (PER) so transitions with high temporal-difference error are sampled preferentially over trivial one-step arrivals.

The Quick Version

  • Core Intuition: In sparse-reward multi-goal RL, failed trajectories to target goal gg are successful trajectories to the states actually reached; relabeling them in hindsight extracts rich learning signals from failures.
  • Goal Relabeling Modes: The future strategy (sampling k≈4k \approx 4 states visited later in the same episode) outperforms final, episode, and random because it respects the reachable time horizon.
  • Off-Policy Synergy: HER strictly requires off-policy reinforcement learning (such as DDPG, TD3, or SAC) because changing the goal vector post-rollout introduces an off-policy distribution shift that invalidates on-policy methods like PPO.
  • Reward Shaping Alternative: HER solves the sparse-reward exploration plateau without the distortion, local optima, and task-specific engineering traps of hand-crafted distance rewards.