Reinforcement Learning2017intermediate10 min read

Hindsight Experience Replay

إعادة تشغيل الخبرة بأثر رجعي

Andrychowicz, M. · Wolski, F. · Ray, A. · Schneider, J. · Fong, R. · Welinder, P. · McGrew, B. · Tobin, J. · Abbeel, P. · Zaremba, W. — NeurIPS

The problem

In many real-world RL tasks — especially robotic manipulation — the is binary and sparse: you get +1 only when you fully achieve the goal, and 0 everywhere else. With random , the probability of accidentally achieving a complex goal is near zero, so the receives no positive reward signal and cannot learn. Traditional is brittle, domain-specific, and often introduces .

The contribution

Hindsight (HER): after each , re-store the same transitions in the but with the original goal replaced by a goal that was actually achieved during that episode. This converts every failed into a successful demonstration for an alternative goal. HER works with any (DDPG, DQN, SAC), requires no reward engineering, and acts as an implicit curriculum — easy goals are learned first, enabling progressive mastery of harder goals.

The impact

HER became a foundational technique in goal-conditioned and robotic manipulation. It demonstrated that policies trained purely in simulation with binary rewards could transfer to physical robots. Its principle — learn from failure by redefining what counts as success — influenced all subsequent work on multi-goal RL, , and sim-to-real transfer in robotics.

Imagine a child learning to throw darts. She aims for the bullseye and misses — the dart lands in the lower-left corner. A binary-reward teacher would say: "Wrong. No points." She throws a hundred more times, never hits the bullseye, and learns nothing.

An HER teacher walks up to the board after each throw and draws a new bullseye around wherever the dart actually landed. "See? You hit this target perfectly! Here's what your arm did to get there." Now every throw teaches something: the child builds a mental map of how arm angles map to landing positions, and gradually steers toward the real bullseye.

The problem: sparse rewards make learning nearly impossible

In goal-conditioned reinforcement learning, the agent must reach a specific target gg from a starting state. The simplest and most natural reward is binary:

r(s,a,g)=−[fg(s′)=0]r(s, a, g) = -[f_g(s') = 0]

The agent receives 00 when the goal is achieved and −1-1 otherwise. This avoids reward shaping entirely — no handcrafted distance terms, no domain heuristics, no reward hacking. But there's a devastating catch: with high-dimensional continuous actions (like a robotic arm with 7 joints), the chance of randomly stumbling into the goal state is astronomically small.

The agent could explore for millions of steps and never receive a single positive reward. No positive reward means no signal, and no gradient signal means no learning. The agent is stuck in a desert with no oasis in sight.

Open in Lab
Toggle between sparse and dense rewards. Notice how the agent's learning signal vanishes under sparse rewards — the loss curve flatlines because no positive example is ever seen.
The demo wakes as you arrive…

Why not just shape the reward?

A natural instinct is to add a distance-based reward: give the agent partial credit for getting closer to the goal. This is called reward shaping, and while it sometimes works, it introduces serious problems:

  • Domain expertise required. You need to know what "closer" means for every task — and for a 7-joint arm, Euclidean distance to the goal may not correlate with actual reachability.
  • Reward hacking. The agent may find shortcuts that maximize the shaped reward without achieving the goal — vibrating near the target without grasping, or exploiting simulator quirks.
  • Brittleness. A shaped reward that works for pushing may fail for pick-and-place. Each new task needs re-engineering.

The authors show experimentally that shaped rewards actually hurt performance when combined with HER, confirming that binary rewards are not just simpler but better in this framework.

The idea: redefine success after the fact

The key insight of HER is borrowed from human cognition: when we fail at a task, we don't discard the experience — we ask "what did I achieve?" and file that knowledge for later.

Concretely, consider a π(a∣s,g)\pi(a \mid s, g) trained with an off- algorithm like DDPG. After the agent completes an episode pursuing goal gg, HER stores the transitions twice:

  • Once with the original goal gg (and typically reward −1-1 for failure).
  • Again with a substitute goal g′g' chosen from the states actually visited during that episode (and now reward 00 because g′g' was achieved).

Because the algorithm is off-policy, it can legally learn from transitions generated under a different goal than the one currently being optimized. The policy and are conditioned on the goal, so swapping the goal simply re-contextualizes the same trajectory.

Open in Lab
Watch goal relabeling in action. The agent fails to reach the green goal, but HER creates a new training signal by pretending the achieved state was the goal all along.
The demo wakes as you arrive…

Four strategies for choosing substitute goals

When HER re-stores a transition, it needs to pick a substitute goal g′g'. The paper proposes four strategies — different answers to "which achieved state should I pretend was my goal?"

  • Final — use the last state of the episode. Simple: one extra transition per step.
  • Future — sample kk states visited after the current transition in the same episode. This is the best performer: it creates a natural curriculum where early-step transitions are paired with nearby goals, and later steps with the episode's outcome.
  • Episode — sample kk states from anywhere in the episode. Slightly weaker than future because it includes past states that the agent already passed through.
  • Random — sample kk states from the entire replay buffer. Weakest: the substitute goals may be completely unrelated to the trajectory.

The authors find that future with k=4k=4 works best across all tested environments. This makes intuitive sense: telling the agent "you reached states X₁, X₂, X₃, X₄ after this action" gives the densest local learning signal while maintaining trajectory coherence.

Open in Lab
Select a strategy to see which states become substitute goals. Notice how "future" creates the most informative and coherent relabeling.
The demo wakes as you arrive…

The algorithm step by step

HER wraps around any off-policy RL algorithm. Here is the complete procedure, assuming DDPG as the base learner:

Open in Lab
Click each step to see what happens at that stage of the HER pipeline.
The demo wakes as you arrive…

The critical detail: the relabeled transitions and the original transitions are stored in the same replay buffer. When the algorithm samples a for , it draws from both kinds without distinction. This means the training data naturally contains a mix of "real failures" (original goal, r=−1r=-1) and "virtual successes" (substitute goal, r=0r=0), giving the value function both positive and negative examples to learn from.

Formal framework: Universal Value Function Approximators

HER operates within the Universal Value Function Approximator (UVFA) framework. Instead of a standard value function Q(s,a)Q(s, a), the critic takes the goal as input:

Q(s,a,g)=E[∑t=0Tγtr(st,at,g)∣s0=s,a0=a]Q(s, a, g) = \mathbb{E}\left[\sum_{t=0}^{T} \gamma^t r(s_t, a_t, g) \mid s_0 = s, a_0 = a\right]
Goal-conditioned Q-function (UVFA) — The value function now asks: "how good is action a in state s *for reaching goal g*?" This conditioning is what makes goal relabeling mathematically valid — the same trajectory is consistent with different goals.

The policy is similarly conditioned: π(a∣s,g)\pi(a \mid s, g). Both QQ and π\pi concatenate the state with the goal as input. During relabeling, when we swap gg for g′g', we recompute:

r′=r(st,at,g′)=−[fg′(st+1)=0]r' = r(s_t, a_t, g') = -[f_{g'}(s_{t+1}) = 0]

If the agent's next state st+1s_{t+1} matches goal g′g' (i.e. fg′(st+1)=1f_{g'}(s_{t+1}) = 1), the reward becomes 00 (success). This is the only change — state, action, and next state remain identical.

HER in code

Hindsight Experience Replay — the core relabeling looppython

Simplified to show the idea — not the real implementation.

def her_relabel(episode, k=4, strategy='future'):
    """Relabel transitions from one episode using HER.

    episode: list of (state, action, reward, next_state, goal, achieved_goal)
    k: number of substitute goals per transition
    strategy: 'future', 'final', 'episode', or 'random'
    """
    relabeled = []
    T = len(episode)

    for t, (s, a, r, s_next, g, ag) in enumerate(episode):
        # 1) Always store the original transition
        relabeled.append((s, a, r, s_next, g))

        # 2) Pick substitute goals based on strategy
        if strategy == 'future':
            # Sample k goals from states visited AFTER this step
            future_indices = np.random.randint(t + 1, T, size=k)  # exclusive
            new_goals = [episode[i][5] for i in future_indices]  # achieved_goal
        elif strategy == 'final':
            new_goals = [episode[-1][5]]  # last achieved state
        elif strategy == 'episode':
            indices = np.random.randint(0, T, size=k)
            new_goals = [episode[i][5] for i in indices]

        # 3) For each substitute goal, recompute reward and store
        for g_prime in new_goals:
            r_new = 0.0 if goal_achieved(s_next, g_prime) else -1.0
            relabeled.append((s, a, r_new, s_next, g_prime))

    return relabeled  # → feed these into the replay buffer

Why it works: implicit curriculum and universal generalization

HER's effectiveness rests on two pillars:

Implicit curriculum. Early in training, the agent's policy is essentially random, so it only reaches states near the start. When HER relabels these with the "future" strategy, the substitute goals are close and easy. As the policy improves and reaches farther states, the relabeled goals become harder — a natural difficulty progression that no one designed.

Universal . Because the Q-function is conditioned on the goal, learning "how to reach state X" simultaneously teaches the agent about reaching nearby states. A single trajectory, once relabeled with multiple goals, trains the value function across a region of the goal space, not just a single point. This is dramatically more sample-efficient than learning each goal independently.

Experiments: robotic manipulation with binary rewards

The paper evaluates HER on three simulated robotic manipulation tasks using a 7-DOF Fetch robotic arm in MuJoCo, each with binary reward (−1-1 or 00):

  • Push — slide a puck to a target position on a table.
  • Slide — strike a puck so it slides to a distant target (requires precise force).
  • Pick-and-place — grasp an object, lift it, and place it at a 3D target position.

Without HER, DDPG with sparse rewards learns nothing on any of these tasks — the success rate stays at 0% across all training. With HER using the "future" strategy (k=4k=4), the agent achieves near-perfect success rates on push and slide, and strong performance on the much harder pick-and-place task.

Open in Lab
Success rate over training epochs. DDPG alone (no HER) flatlines at 0%. HER+DDPG with the "future" strategy achieves near-100% on push and slide.
The demo wakes as you arrive…

From simulation to real robot

A remarkable result: policies trained entirely in MuJoCo simulation were deployed directly on a physical Fetch robot with no . The push task transferred successfully to the real world. This demonstrated that HER doesn't just solve the sparse-reward problem in simulation — it produces policies robust enough for real-world deployment.

The sim-to-real transfer worked because the policy learned genuine manipulation skills from the relabeled experience, not simulator-specific tricks. When the only signal is "did you reach this state or not?", there's no room for reward hacking — the policy must actually learn to control the arm.

Impact and legacy

  1. 2017

    HER published at NeurIPS

    Andrychowicz et al. demonstrate that binary-reward robotic manipulation is solvable with goal relabeling. Policies transfer from simulation to a physical Fetch robot.

  2. 2018

    Multi-goal environments standardized

    OpenAI Gym introduces GoalEnv interface with achieved_goal / desired_goal fields, directly inspired by HER's requirements.

  3. 2018

    Demonstrations + HER

    Nair et al. combine HER with human demonstrations for long-horizon tasks like block stacking, extending HER's reach to more complex manipulation.

  4. 2018

    Curiosity-driven exploration + HER

    Researchers combine intrinsic motivation with HER, further improving exploration in environments where even random actions rarely visit diverse states.

  5. 2019

    HER integrated into Stable Baselines

    HER becomes a standard component in major RL libraries (Stable Baselines, RLlib), establishing it as infrastructure rather than a research prototype.

  6. 2020

    Goal-conditioned RL boom

    HER's principle inspires a wave of goal-conditioned methods including goal-conditioned supervised learning, planning with hindsight, and energy-based hindsight models.

HER's deepest legacy is a shift in mindset: failure is not wasted experience — it's experience for a different goal. This principle extends far beyond robotics. Recent work applies the same hindsight idea to language agent trajectories, curriculum generation, and under uncertainty.

CitationAndrychowicz, Wolski, Ray, Schneider, Fong, Welinder, McGrew, Tobin, Abbeel, Zaremba. Hindsight Experience Replay. NeurIPS, 2017.

Terms in this paper