Reinforcement Learning2017intermediate10 min read
Hindsight Experience Replay
إعادة تشغيل الخبرة بأثر رجعي
Andrychowicz, M. · Wolski, F. · Ray, A. · Schneider, J. · Fong, R. · Welinder, P. · McGrew, B. · Tobin, J. · Abbeel, P. · Zaremba, W. — NeurIPS
The problem
In many real-world RL tasks — especially robotic manipulation — the is binary and sparse: you get +1 only when you fully achieve the goal, and 0 everywhere else. With random , the probability of accidentally achieving a complex goal is near zero, so the receives no positive reward signal and cannot learn. Traditional is brittle, domain-specific, and often introduces .
The contribution
Hindsight (HER): after each , re-store the same transitions in the but with the original goal replaced by a goal that was actually achieved during that episode. This converts every failed into a successful demonstration for an alternative goal. HER works with any (DDPG, DQN, SAC), requires no reward engineering, and acts as an implicit curriculum — easy goals are learned first, enabling progressive mastery of harder goals.
The impact
HER became a foundational technique in goal-conditioned and robotic manipulation. It demonstrated that policies trained purely in simulation with binary rewards could transfer to physical robots. Its principle — learn from failure by redefining what counts as success — influenced all subsequent work on multi-goal RL, , and sim-to-real transfer in robotics.
Imagine a child learning to throw darts. She aims for the bullseye and misses — the dart lands in the lower-left corner. A binary-reward teacher would say: "Wrong. No points." She throws a hundred more times, never hits the bullseye, and learns nothing.
An HER teacher walks up to the board after each throw and draws a new bullseye around wherever the dart actually landed. "See? You hit this target perfectly! Here's what your arm did to get there." Now every throw teaches something: the child builds a mental map of how arm angles map to landing positions, and gradually steers toward the real bullseye.
The problem: sparse rewards make learning nearly impossible
In goal-conditioned reinforcement learning, the agent must reach a specific target from a starting state. The simplest and most natural reward is binary:
The agent receives when the goal is achieved and otherwise. This avoids reward shaping entirely — no handcrafted distance terms, no domain heuristics, no reward hacking. But there's a devastating catch: with high-dimensional continuous actions (like a robotic arm with 7 joints), the chance of randomly stumbling into the goal state is astronomically small.
The agent could explore for millions of steps and never receive a single positive reward. No positive reward means no signal, and no gradient signal means no learning. The agent is stuck in a desert with no oasis in sight.
Why not just shape the reward?
A natural instinct is to add a distance-based reward: give the agent partial credit for getting closer to the goal. This is called reward shaping, and while it sometimes works, it introduces serious problems:
- Domain expertise required. You need to know what "closer" means for every task — and for a 7-joint arm, Euclidean distance to the goal may not correlate with actual reachability.
- Reward hacking. The agent may find shortcuts that maximize the shaped reward without achieving the goal — vibrating near the target without grasping, or exploiting simulator quirks.
- Brittleness. A shaped reward that works for pushing may fail for pick-and-place. Each new task needs re-engineering.
The authors show experimentally that shaped rewards actually hurt performance when combined with HER, confirming that binary rewards are not just simpler but better in this framework.
The idea: redefine success after the fact
The key insight of HER is borrowed from human cognition: when we fail at a task, we don't discard the experience — we ask "what did I achieve?" and file that knowledge for later.
Concretely, consider a trained with an off- algorithm like DDPG. After the agent completes an episode pursuing goal , HER stores the transitions twice:
- Once with the original goal (and typically reward for failure).
- Again with a substitute goal chosen from the states actually visited during that episode (and now reward because was achieved).
Because the algorithm is off-policy, it can legally learn from transitions generated under a different goal than the one currently being optimized. The policy and are conditioned on the goal, so swapping the goal simply re-contextualizes the same trajectory.
Four strategies for choosing substitute goals
When HER re-stores a transition, it needs to pick a substitute goal . The paper proposes four strategies — different answers to "which achieved state should I pretend was my goal?"
- Final — use the last state of the episode. Simple: one extra transition per step.
- Future — sample states visited after the current transition in the same episode. This is the best performer: it creates a natural curriculum where early-step transitions are paired with nearby goals, and later steps with the episode's outcome.
- Episode — sample states from anywhere in the episode. Slightly weaker than future because it includes past states that the agent already passed through.
- Random — sample states from the entire replay buffer. Weakest: the substitute goals may be completely unrelated to the trajectory.
The authors find that future with works best across all tested environments. This makes intuitive sense: telling the agent "you reached states X₁, X₂, X₃, X₄ after this action" gives the densest local learning signal while maintaining trajectory coherence.
The algorithm step by step
HER wraps around any off-policy RL algorithm. Here is the complete procedure, assuming DDPG as the base learner:
The critical detail: the relabeled transitions and the original transitions are stored in the same replay buffer. When the algorithm samples a for , it draws from both kinds without distinction. This means the training data naturally contains a mix of "real failures" (original goal, ) and "virtual successes" (substitute goal, ), giving the value function both positive and negative examples to learn from.
Formal framework: Universal Value Function Approximators
HER operates within the Universal Value Function Approximator (UVFA) framework. Instead of a standard value function , the critic takes the goal as input:
The policy is similarly conditioned: . Both and concatenate the state with the goal as input. During relabeling, when we swap for , we recompute:
If the agent's next state matches goal (i.e. ), the reward becomes (success). This is the only change — state, action, and next state remain identical.
HER in code
Simplified to show the idea — not the real implementation.
def her_relabel(episode, k=4, strategy='future'):
"""Relabel transitions from one episode using HER.
episode: list of (state, action, reward, next_state, goal, achieved_goal)
k: number of substitute goals per transition
strategy: 'future', 'final', 'episode', or 'random'
"""
relabeled = []
T = len(episode)
for t, (s, a, r, s_next, g, ag) in enumerate(episode):
# 1) Always store the original transition
relabeled.append((s, a, r, s_next, g))
# 2) Pick substitute goals based on strategy
if strategy == 'future':
# Sample k goals from states visited AFTER this step
future_indices = np.random.randint(t + 1, T, size=k) # exclusive
new_goals = [episode[i][5] for i in future_indices] # achieved_goal
elif strategy == 'final':
new_goals = [episode[-1][5]] # last achieved state
elif strategy == 'episode':
indices = np.random.randint(0, T, size=k)
new_goals = [episode[i][5] for i in indices]
# 3) For each substitute goal, recompute reward and store
for g_prime in new_goals:
r_new = 0.0 if goal_achieved(s_next, g_prime) else -1.0
relabeled.append((s, a, r_new, s_next, g_prime))
return relabeled # → feed these into the replay bufferWhy it works: implicit curriculum and universal generalization
HER's effectiveness rests on two pillars:
Implicit curriculum. Early in training, the agent's policy is essentially random, so it only reaches states near the start. When HER relabels these with the "future" strategy, the substitute goals are close and easy. As the policy improves and reaches farther states, the relabeled goals become harder — a natural difficulty progression that no one designed.
Universal . Because the Q-function is conditioned on the goal, learning "how to reach state X" simultaneously teaches the agent about reaching nearby states. A single trajectory, once relabeled with multiple goals, trains the value function across a region of the goal space, not just a single point. This is dramatically more sample-efficient than learning each goal independently.
Experiments: robotic manipulation with binary rewards
The paper evaluates HER on three simulated robotic manipulation tasks using a 7-DOF Fetch robotic arm in MuJoCo, each with binary reward ( or ):
- Push — slide a puck to a target position on a table.
- Slide — strike a puck so it slides to a distant target (requires precise force).
- Pick-and-place — grasp an object, lift it, and place it at a 3D target position.
Without HER, DDPG with sparse rewards learns nothing on any of these tasks — the success rate stays at 0% across all training. With HER using the "future" strategy (), the agent achieves near-perfect success rates on push and slide, and strong performance on the much harder pick-and-place task.
From simulation to real robot
A remarkable result: policies trained entirely in MuJoCo simulation were deployed directly on a physical Fetch robot with no . The push task transferred successfully to the real world. This demonstrated that HER doesn't just solve the sparse-reward problem in simulation — it produces policies robust enough for real-world deployment.
The sim-to-real transfer worked because the policy learned genuine manipulation skills from the relabeled experience, not simulator-specific tricks. When the only signal is "did you reach this state or not?", there's no room for reward hacking — the policy must actually learn to control the arm.
Impact and legacy
2017
HER published at NeurIPS
Andrychowicz et al. demonstrate that binary-reward robotic manipulation is solvable with goal relabeling. Policies transfer from simulation to a physical Fetch robot.
2018
Multi-goal environments standardized
OpenAI Gym introduces GoalEnv interface with achieved_goal / desired_goal fields, directly inspired by HER's requirements.
2018
Demonstrations + HER
Nair et al. combine HER with human demonstrations for long-horizon tasks like block stacking, extending HER's reach to more complex manipulation.
2018
Curiosity-driven exploration + HER
Researchers combine intrinsic motivation with HER, further improving exploration in environments where even random actions rarely visit diverse states.
2019
HER integrated into Stable Baselines
HER becomes a standard component in major RL libraries (Stable Baselines, RLlib), establishing it as infrastructure rather than a research prototype.
2020
Goal-conditioned RL boom
HER's principle inspires a wave of goal-conditioned methods including goal-conditioned supervised learning, planning with hindsight, and energy-based hindsight models.
HER's deepest legacy is a shift in mindset: failure is not wasted experience — it's experience for a different goal. This principle extends far beyond robotics. Recent work applies the same hindsight idea to language agent trajectories, curriculum generation, and under uncertainty.
CitationAndrychowicz, Wolski, Ray, Schneider, Fong, Welinder, McGrew, Tobin, Abbeel, Zaremba. Hindsight Experience Replay. NeurIPS, 2017.
Terms in this paper
- Experience Replayإعادة تشغيل التجارب
- Sparse Rewardالمكافأة الشحيحة
- Goal-Conditioned Policyالسياسة المشروطة بالأهداف
- Reward Functionدالة صياغة المكافآت
- Sample Efficiencyكفاءة استخدام العيّنات
- Off-Policyخوارزمية التعلم خارج السياسة الحالية
- Explorationالاستكشاف (تجربة أفعال جديدة)
- Replay Bufferذاكرة التجارب
- Reward Shapingصياغة وهندسة دالة المكافأة