Reinforcement Learning2016intermediate11 min read

High-Dimensional Continuous Control Using Generalized Advantage Estimation

التحكم المستمر في الأبعاد العالية باستخدام تقدير الأفضلية المعمَّم

Schulman, J. · Moritz, P. · Levine, S. · Jordan, M. · Abbeel, P. — ICLR

The problem

methods directly optimize what we want — cumulative — and work naturally with neural networks. But they face two stubborn challenges: they need enormous amounts of data (high in estimates), and is unstable because the data distribution shifts as the improves (nonstationarity). Reducing variance by using a introduces bias; the question is how to control this tradeoff.

The contribution

: an exponentially-weighted sum of temporal-difference residuals that estimates the function. A single parameter λ ∈ [0,1] smoothly interpolates between the one-step TD estimate (low variance, high bias) and the estimate (high variance, low bias), analogous to TD(λ). Combined with optimization for both policy and value function, GAE enabled neural-network agents to learn complex 3D locomotion gaits — bipedal running, quadrupedal galloping, and standing up from the ground — for the first time.

The impact

GAE became a standard building block in modern deep . , the most widely used policy gradient algorithm today, relies on GAE for advantage estimation. From game-playing agents to robotic control to RLHF for large language models, GAE's bias-variance knob is quietly running behind the scenes whenever a policy gradient is computed.

Imagine you're coaching a soccer player. After each game, you could review the whole match to decide what worked — accurate but overwhelming. Or you could judge each individual kick — fast but misleading, because a great pass might lead to a goal three moves later.

GAE is like reviewing weighted highlights: recent plays get the most attention, but you still glance at what happened next. A dial called λ controls how far ahead you look — turn it up for the full picture (noisy), down for instant feedback (potentially misleading), or somewhere in between for the best of both worlds.

The problem: policy gradients are noisy

In reinforcement learning, the collects trajectories — sequences of states, actions, and rewards — and uses them to improve its policy. The policy gradient tells us which direction to adjust the policy parameters so that future trajectories yield higher cumulative reward.

The catch: the raw policy gradient uses the total return following each action, and that return is the sum of many future rewards, each subject to the randomness of the and the policy. Like measuring the height of a wave in a storm, each measurement is valid but wildly variable. This high variance means you need millions of samples before the gradient signal rises above the noise.

A natural fix is to subtract a baseline — typically the value function V(s)V(s) — from the return. What remains is the advantage function A(s,a)=Q(s,a)−V(s)A(s,a) = Q(s,a) - V(s): how much better (or worse) this specific action was compared to the average action from this . The advantage is a much cleaner signal because the common baseline soaks up most of the randomness.

Open in Lab
Drag λ between 0 and 1 to see how GAE trades off bias and variance. The scales tip toward bias at λ=0 and toward variance at λ=1.
The demo wakes as you arrive…

Building block: the TD residual

Before we can build GAE, we need its fundamental brick: the temporal-difference (TD) residual. After taking action ata_t in state sts_t, the agent observes reward rtr_t and lands in state st+1s_{t+1}. The TD residual measures the surprise: how much better the actual one-step outcome was compared to what the value function predicted.

Think of it like checking your bank balance. You expected V(st)V(s_t) in your account. After one transaction, you received rtr_t and now your expected future balance is V(st+1)V(s_{t+1}). The residual δt\delta_t is the discrepancy — the unexpected gain or loss.

δt=rt+γV(st+1)−V(st)\delta_t = r_t + \gamma V(s_{t+1}) - V(s_t)
TD residual — the one-step surprise — r_t = immediate reward · γV(s_{t+1}) = discounted prediction of everything to come · V(s_t) = what we expected · the difference = how much better or worse reality was

If the value function VV were perfect, δt\delta_t would be an unbiased estimate of the advantage. In practice VV is approximate (a ), so δt\delta_t carries bias — but it's a single number per timestep, so its variance is low. This is the low-variance-high-bias extreme.

At the other extreme, you could ignore VV entirely and use the full Monte Carlo return — summing all actual rewards to the end of the . That's unbiased but has enormous variance because it depends on every future random event.

GAE gives us a smooth path between these two extremes.

The GAE formula: a λ-weighted blend

The key insight is to take an exponentially-weighted average of multi-step advantage estimates. A one-step estimate uses one TD residual. A two-step estimate chains two residuals together. A k-step estimate chains k residuals. GAE blends all of them with weights that decay exponentially by the factor γλ\gamma\lambda:

A^tGAE(γ,λ)=∑l=0∞(γλ)l δt+l\hat{A}_t^{GAE(\gamma,\lambda)} = \sum_{l=0}^{\infty} (\gamma\lambda)^l \, \delta_{t+l}
Generalized Advantage Estimation — the core formula — Each future TD residual δ is weighted by (γλ)ˡ — close residuals count most, distant ones fade away · λ=0 keeps only δ_t (one-step, low variance, high bias) · λ=1 sums all residuals (Monte Carlo, high variance, low bias) · λ in between interpolates smoothly

This is strikingly analogous to TD(λ) from learning. In TD(λ), eligibility traces blend one-step and multi-step value updates. GAE applies the same exponential blending idea, but to the advantage function used in policy gradients rather than to value updates.

The elegance is in the simplicity: one parameter controls the entire bias-variance spectrum. In practice, λ≈0.95\lambda \approx 0.95–0.970.97 tends to work well — mostly Monte Carlo–like (rich multi-step signal) but with just enough decay to tame the variance.

Open in Lab
Adjust λ to see how it changes the weights on each future TD residual. Notice how λ=0 uses only the first residual while λ=1 weights all residuals equally.
The demo wakes as you arrive…

Two knobs: γ vs λ

GAE has two parameters that both affect the , but they play different roles:

  • γ (gamma) — the . It downweights rewards that are far in the future. Lowering γ makes the agent more short-sighted, reducing variance (fewer future rewards matter) but introducing bias (the agent ignores delayed consequences). Think of γ as the agent's time horizon for caring about outcomes.

  • λ (lambda) — the trace decay. It downweights TD residuals that are far in the future. Lowering λ makes the advantage estimate rely more on the value function's predictions (low variance, high bias). Think of λ as how much you trust your value function versus raw experience.

In practice, both are set between 0.9 and 0.99. A common recipe: γ=0.99\gamma = 0.99, λ=0.95\lambda = 0.95.

Open in Lab
Explore how different (γ, λ) combinations affect the effective weight each future step receives. Bright = high weight, dark = low weight.
The demo wakes as you arrive…

Special cases: the two extremes

The beauty of GAE is that its two boundary values recover two well-known estimators:

When λ = 0: the infinite sum collapses to just δt\delta_t, the one-step TD residual. This is the same as the simple advantage estimate rt+γV(st+1)−V(st)r_t + \gamma V(s_{t+1}) - V(s_t). It has low variance because it depends on a single transition, but high bias if the value function is imperfect.

When λ = 1: the sum becomes ∑l=0∞γlδt+l\sum_{l=0}^{\infty} \gamma^l \delta_{t+l}, which telescopes into ∑l=0∞γlrt+l−V(st)\sum_{l=0}^{\infty} \gamma^l r_{t+l} - V(s_t) — the discounted Monte Carlo return minus the baseline. Zero bias (given the correct discount), but high variance because it uses every future reward.

Any λ\lambda between 0 and 1 gives you a point on the continuous spectrum between these extremes.

Open in Lab
Move λ from 0 to 1 and watch the GAE estimate morph from the one-step TD advantage to the full Monte Carlo advantage. The shaded area shows the variance envelope.
The demo wakes as you arrive…

Stable training: trust regions for both networks

Having a good advantage estimator is only half the battle. The paper's second contribution is applying trust region optimization to both the policy network and the value function network.

For the policy, TRPO constrains each update so the new policy stays close to the old one (measured by ). This prevents catastrophic policy collapses where one bad update ruins everything.

For the value function, the paper uses a similar trust-region constraint. This is important because the value function appears inside the advantage estimate: if the value function changes too wildly between iterations, the advantage signal becomes unreliable. Think of it as calibrating your measuring instrument before every measurement — if the instrument drifts, your readings are meaningless.

A subtle point: the value function must be fitted after computing the policy gradient, using the old value function for the advantage. If you updated VV first and it overfitted, the TD residuals δt\delta_t would all become near-zero, and the policy gradient would vanish.

The complete policy gradient with GAE

The policy gradient tells us how to nudge the policy parameters to increase expected reward. With GAE plugged in as the advantage estimator, the gradient becomes:

g≈E[∑t=0∞∇θlog⁡πθ(at∣st)  A^tGAE(γ,λ)]g \approx \mathbb{E}\left[\sum_{t=0}^{\infty} \nabla_\theta \log \pi_\theta(a_t|s_t) \;\hat{A}_t^{GAE(\gamma,\lambda)}\right]
Policy gradient with GAE — ∇log π = direction that makes action a_t more likely · Â^GAE = how much better than average that action was · positive  → make more likely, negative  → make less likely

Read this formula as a simple instruction to the : for each action the agent took, if the GAE advantage says it was better than average, increase its probability; if worse, decrease it. The magnitude of the advantage determines how aggressive the update is. Because GAE controls variance, these updates are stable enough to make steady progress on complex tasks.

GAE in code

Compute GAE advantages — the backward pass trickpython

Simplified to show the idea — not the real implementation.

import numpy as np

def compute_gae(rewards, values, gamma=0.99, lam=0.95):
    """Compute GAE advantages for one trajectory.

    rewards: array of r_t        (length T)
    values:  array of V(s_t)      (length T+1, includes terminal)
    gamma:   discount factor
    lam:     GAE lambda (bias-variance knob)
    """
    T = len(rewards)
    advantages = np.zeros(T)
    gae = 0.0  # running sum, built from the end

    for t in reversed(range(T)):
        delta = rewards[t] + gamma * values[t+1] - values[t]  # TD residual
        gae = delta + gamma * lam * gae  # accumulate weighted residuals
        advantages[t] = gae

    returns = advantages + values[:T]  # GAE advantage + baseline = target
    return advantages, returns

# The trick: iterate backward so each step reuses the sum from t+1.
# This makes the whole computation O(T) — linear in trajectory length.

Experiments: 3D locomotion from scratch

The paper demonstrated GAE on some of the hardest continuous control tasks of its time, using the MuJoCo physics simulator:

  • Bipedal walker — a simulated humanoid learning to walk and run. With 33 state dimensions and 10 action dimensions, the agent must coordinate hip, knee, and ankle joints simultaneously.

  • Quadrupedal runner — a four-legged robot learning a galloping gait. Even more joints and contacts to coordinate.

  • Standing up — a humanoid starting on the ground and learning to stand. This requires a multi-phase strategy that pure one-step methods struggle with.

The key finding: λ = 0 (one-step TD) performed poorly on all tasks due to excessive bias, while λ between 0.9 and 0.99 with appropriate γ produced the best learning curves. This validated that the intermediate bias-variance point is crucial for hard problems.

Open in Lab
Simulated learning curves for different λ values on a locomotion task. Notice how λ=0 plateaus early while λ=0.96 reaches the highest performance.
The demo wakes as you arrive…

The full algorithm at a glance

Putting all the pieces together, the GAE-based policy optimization algorithm repeats the following loop:

Open in Lab
Step through the GAE + TRPO training loop: collect data, compute advantages, update policy, fit value function.
The demo wakes as you arrive…

Legacy: the engine inside PPO and beyond

GAE's most visible legacy is Proximal Policy Optimization (PPO), published by the same lead author one year later. PPO replaced TRPO's complex second-order optimization with a simple clipped , but kept GAE as its advantage estimator unchanged. Today, virtually every PPO implementation — from game-playing agents to robotic manipulation to RLHF for language models — computes advantages using the exact formula from this paper.

The influence extends further:

  • RLHF pipelines for Claude, GPT, and other language models use PPO with GAE to compute advantages when from human preference data.

  • Robotics — real-world robot learning systems routinely use GAE when training locomotion and manipulation policies.

  • Game AI — OpenAI Five (Dota 2) and other complex game agents relied on GAE for stable training across billions of environment steps.

  1. 2015

    TRPO published

    Schulman introduces Trust Region Policy Optimization — monotonic improvement guarantees via constrained KL updates. GAE builds on this foundation.

  2. 2016

    GAE published (ICLR)

    The GAE paper shows how λ-weighted TD residuals dramatically reduce variance in policy gradient estimation, enabling complex 3D locomotion.

  3. 2017

    PPO published

    Proximal Policy Optimization simplifies TRPO with a clipped objective while keeping GAE as the standard advantage estimator.

  4. 2019

    OpenAI Five defeats Dota 2 world champions

    PPO+GAE at massive scale — 45,000 years of gameplay experience, with GAE providing stable advantage estimates throughout.

  5. 2022

    InstructGPT and ChatGPT

    RLHF training of language models uses PPO+GAE to align model outputs with human preferences, bringing GAE into the LLM era.

CitationSchulman, Moritz, Levine, Jordan, Abbeel. High-Dimensional Continuous Control Using Generalized Advantage Estimation. ICLR, 2016.

Terms in this paper