Reinforcement Learning2017intermediate10 min read

Proximal Policy Optimization Algorithms

خوارزمية تحسين السياسة القريبة (PPO)

Schulman, J. · Wolski, F. · Dhariwal, P. · Radford, A. · Klimov, O. — arXiv

The problem

methods in are powerful but fragile. Standard REINFORCE uses each data sample once then discards it — terribly sample-inefficient. TRPO guarantees monotonic improvement through a hard KL-divergence constraint, but its second-order optimization is complex, hard to implement with shared architectures, and incompatible with noise or . Practitioners needed an algorithm with TRPO's stability but the simplicity of first-order methods like .

The contribution

PPO introduces a clipped that prevents destructively large updates without any second-order optimization. The r(θ) = π_θ(a|s) / π_old(a|s) is clipped to [1−ε, 1+ε], creating a flat ceiling that stops the optimizer once the policy has changed "enough." Combined with multiple epochs of minibatch updates on the same data, PPO achieves TRPO-level stability with the simplicity of a few lines of code. An adaptive KL penalty variant is also proposed.

The impact

PPO became the default RL algorithm of the deep learning era. OpenAI used it to train ChatGPT via RLHF, DeepSeek-R1 used it for reasoning, OpenAI Five played Dota 2 with it, and robotic manipulation systems like Learning Dexterity depend on it. Its simplicity and robustness made reinforcement learning accessible to practitioners who are not RL specialists.

a policy is like tuning a recipe while the restaurant is open. You taste the dish, adjust the seasoning, and serve again. REINFORCE throws away the dish after one taste — wasteful. TRPO hires a food scientist to calculate the exact safe adjustment — accurate but expensive. PPO is the pragmatic chef: taste, adjust, but never add more than a pinch at a time. If the seasoning change is too large, just stop — don't keep pouring. That "pinch limit" is the mechanism, and it makes PPO both stable and dead simple.

The problem: policy updates that overshoot

In reinforcement learning, a policy πθ(a∣s)\pi_\theta(a|s) maps states to action probabilities. The collects trajectories, computes returns, and updates θ\theta to increase the probability of good actions. This is the policy idea behind REINFORCE.

The core difficulty: how far should each update go?

  • Too small and learning is painfully slow — every trajectory is used once then discarded.
  • Too large and the policy can collapse: it overshoots into a bad region, collects bad data from there, and spirals downward. One bad update can undo thousands of good ones.

This is not a theoretical fear. In practice, vanilla policy gradient training is famously unstable, and tuning the is a dark art.

Open in Lab
Compare the complexity of TRPO's constrained optimization with PPO's simple clipping.
The demo wakes as you arrive…

TRPO's answer: the trust region

Policy Optimization solved the overshoot problem by adding a hard constraint: the between the old and new policies must stay below a threshold δ\delta. This guarantees monotonic improvement in theory — every update either improves the policy or does nothing.

But TRPO pays a price for this guarantee:

  • It requires computing the Fisher information matrix (second-order derivatives), making each step expensive.
  • The constrained optimization uses conjugate gradient with a line search — complex to implement and debug.
  • It is incompatible with architectures that share parameters between policy and value networks, and with regularization techniques like dropout.

PPO asks: can we get TRPO's stability with nothing more than first-order gradients and a clever ?

The core idea: clip the probability ratio

PPO starts from the same surrogate objective that TRPO uses, but replaces the hard constraint with a remarkably simple mechanism. First, define the probability ratio:

rt(θ)=πθ(at∣st)πθold(at∣st)r_t(\theta) = \frac{\pi_\theta(a_t | s_t)}{\pi_{\theta_\text{old}}(a_t | s_t)}

This ratio measures how much the new policy differs from the old one for a specific action. When r=1r = 1, the policies agree perfectly. When r>1r > 1, the new policy assigns more probability to this action; when r<1r < 1, less.

The unconstrained surrogate objective would be:

LCPI(θ)=E^t[rt(θ) A^t]L^{CPI}(\theta) = \hat{\mathbb{E}}_t \left[ r_t(\theta) \, \hat{A}_t \right]

where A^t\hat{A}_t is the estimated — how much better this action was compared to the average. Maximizing this objective pushes the policy to take good actions more often. But without a constraint, nothing stops rt(θ)r_t(\theta) from growing unboundedly, leading to destructively large updates.

LCLIP(θ)=E^t[min⁡(rt(θ) A^t,  clip(rt(θ),1−ε,1+ε) A^t)]L^{CLIP}(\theta) = \hat{\mathbb{E}}_t \left[ \min \left( r_t(\theta) \, \hat{A}_t, \; \text{clip}(r_t(\theta), 1-\varepsilon, 1+\varepsilon) \, \hat{A}_t \right) \right]
The PPO-Clip objective — the key equation — Take the minimum of the unclipped and clipped objectives. The clipped term flattens the objective when r(θ) strays beyond [1−ε, 1+ε], removing the incentive for the optimizer to push further. ε = 0.2 is the typical default.

Why the minimum? Think of it as a pessimistic bound. For actions with positive advantage (A^t>0\hat{A}_t > 0 — actions better than average), the clipped version caps the benefit when rr exceeds 1+ε1+\varepsilon: "yes, this action is good, but you've already increased its probability enough." For actions with negative advantage (A^t<0\hat{A}_t < 0 — actions worse than average), the clipped version prevents rr from dropping below 1−ε1-\varepsilon: "yes, this action is bad, but don't decrease its probability too fast."

The clip creates a flat region in the objective landscape — a plateau where the gradient is zero. Once you've changed the policy "enough," the optimizer sees no slope and naturally stops. No constraint solver needed, no second-order math, no Fisher matrix. Just a min and a clip.

Open in Lab
Drag the probability ratio r(θ) and watch how the clipped objective creates a flat region that prevents excessively large updates.
The demo wakes as you arrive…

The ε hyperparameter: how much change is "enough"?

The clipping threshold ε controls the trade-off between stability and speed of learning. A large ε (say 0.3) allows bigger policy changes per update, learning faster but risking instability. A small ε (say 0.1) constrains updates tightly, making training safer but slower. The paper uses ε = 0.2 as the default, which works remarkably well across diverse tasks — from Atari games to simulated robotic locomotion.

Think of ε as the speed limit on a highway. Too low and traffic crawls; too high and accidents happen. The paper found that 0.2 is the sweet spot — fast enough to learn efficiently, slow enough to avoid crashes.

Open in Lab
Adjust ε and watch how the clipping region expands or contracts around r = 1.
The demo wakes as you arrive…

The alternative: adaptive KL penalty

PPO also proposes a second variant — PPO-Penalty — that uses a KL divergence penalty instead of clipping. The idea: add a penalty term β⋅KL[πθold,πθ]\beta \cdot KL[\pi_{\theta_\text{old}}, \pi_\theta] to the objective, then adapt β\beta automatically:

  • If the actual KL divergence after an update is too large (above 1.5×dtarg1.5 \times d_\text{targ}), increase β\beta — tighten the leash.
  • If the actual KL is too small (below dtarg/1.5d_\text{targ} / 1.5), decrease β\beta — loosen the leash.

This way the penalty strength self-tunes. In practice, PPO-Clip tends to be preferred because it is simpler and slightly more robust, but PPO-Penalty can be useful when you want more direct control over policy divergence.

Open in Lab
Watch β adapt over training steps — it increases when KL overshoots and decreases when KL is too small.
The demo wakes as you arrive…

The PPO training loop: simple and GPU-friendly

The full PPO algorithm is beautifully simple. In each iteration:

  1. Collect — run the current policy in the for TT timesteps across NN parallel actors, producing a batch of (state, action, ) tuples.
  2. Estimate advantages — use to compute A^t\hat{A}_t for each timestep. GAE balances bias and variance through its λ\lambda .
  3. Optimize — for KK epochs, shuffle the batch into minibatches and run SGD on the combined objective:

Lt(θ)=LtCLIP(θ)−c1LtVF(θ)+c2S[πθ](st)L_t(\theta) = L_t^{CLIP}(\theta) - c_1 L_t^{VF}(\theta) + c_2 S[\pi_\theta](s_t)

where LVFL^{VF} is the (training the critic) and SS is an bonus that encourages . 4. Repeat — set θold←θ\theta_\text{old} \leftarrow \theta and go to step 1.

The key insight: because the clipped objective prevents destructive updates, you can safely reuse the same batch for multiple epochs (K=3–10K = 3\text{–}10) instead of discarding it after one gradient step. This dramatically improves sample efficiency.

Open in Lab
Step through each phase of the PPO training loop and see how data flows.
The demo wakes as you arrive…

The clipped objective in code

PPO-Clip objective — the core in ~15 linespython

Simplified to show the idea — not the real implementation.

import torch

def ppo_clip_objective(
    log_probs_new,    # log π_θ(a|s) from current policy
    log_probs_old,    # log π_θ_old(a|s) from data-collection policy
    advantages,       # Â_t from GAE
    epsilon=0.2       # clipping threshold
):
    # Probability ratio: how much has the policy changed?
    ratio = torch.exp(log_probs_new - log_probs_old)  # r(θ)

    # Unclipped objective: reward big changes (dangerous!)
    unclipped = ratio * advantages

    # Clipped objective: cap the ratio to [1-ε, 1+ε]
    clipped = torch.clamp(ratio, 1 - epsilon, 1 + epsilon) * advantages

    # Take the pessimistic (lower) bound
    loss = -torch.mean(torch.min(unclipped, clipped))

    return loss  # negate because optimizers minimize

Results: beating the competition with simplicity

The paper tested PPO on two domains:

  • Continuous control (MuJoCo): simulated robots learning to walk, run, and hop. PPO matched or exceeded TRPO on almost every task, while being much simpler.
  • Atari games: 49 classic Arcade games. PPO matched or exceeded the performance of A2C (a synchronous version of A3C) and ACER (an off-policy method) on the majority of games, with a single set of hyperparameters.

The critical result: PPO achieves a favorable balance between sample complexity (how much data it needs), simplicity (how easy it is to implement), and wall-time (how fast training runs). No other algorithm at the time matched all three.

Open in Lab
Compare PPO against baselines across different environments.
The demo wakes as you arrive…

Why PPO became the default

The timeline below shows how PPO became the backbone of modern AI alignment and beyond. What started as a robotics optimization trick turned into the engine behind the most consequential AI systems of the 2020s.

  1. 2017

    PPO published

    Schulman et al. introduce PPO at OpenAI. Immediately adopted for robotics and game playing. Replaces TRPO as the go-to algorithm.

  2. 2018

    OpenAI Five — Dota 2

    PPO trained a team of five neural networks to defeat human world champions at Dota 2, a complex multi-agent strategy game requiring long-horizon planning.

  3. 2019

    Learning Dexterity — robotic hand

    A robotic hand learned to manipulate a Rubik's cube using PPO in simulation, then transferred to a physical robot — a landmark in sim-to-real transfer.

  4. 2020

    Learning to Summarize with Human Feedback

    OpenAI used PPO to fine-tune language models from human preferences — an early blueprint for RLHF that directly led to InstructGPT and ChatGPT.

  5. 2022

    InstructGPT & ChatGPT

    PPO became the RL engine inside RLHF: a reward model scores outputs, and PPO optimizes the language model to maximize human-preferred responses. ChatGPT brought this pipeline to 100 million users.

  6. 2025

    DeepSeek-R1

    DeepSeek-R1 used pure RL (GRPO, a PPO variant without a critic) to train reasoning abilities, showing PPO's influence extends to next-generation reasoning models.

Key design decisions unpacked

CitationSchulman, Wolski, Dhariwal, Radford, Klimov. Proximal Policy Optimization Algorithms. arXiv, 2017.

Terms in this paper