Reinforcement Learning2017intermediate10 min read
Proximal Policy Optimization Algorithms
خوارزمية تحسين السياسة القريبة (PPO)
Schulman, J. · Wolski, F. · Dhariwal, P. · Radford, A. · Klimov, O. — arXiv
The problem
methods in are powerful but fragile. Standard REINFORCE uses each data sample once then discards it — terribly sample-inefficient. TRPO guarantees monotonic improvement through a hard KL-divergence constraint, but its second-order optimization is complex, hard to implement with shared architectures, and incompatible with noise or . Practitioners needed an algorithm with TRPO's stability but the simplicity of first-order methods like .
The contribution
PPO introduces a clipped that prevents destructively large updates without any second-order optimization. The r(θ) = π_θ(a|s) / π_old(a|s) is clipped to [1−ε, 1+ε], creating a flat ceiling that stops the optimizer once the policy has changed "enough." Combined with multiple epochs of minibatch updates on the same data, PPO achieves TRPO-level stability with the simplicity of a few lines of code. An adaptive KL penalty variant is also proposed.
The impact
PPO became the default RL algorithm of the deep learning era. OpenAI used it to train ChatGPT via RLHF, DeepSeek-R1 used it for reasoning, OpenAI Five played Dota 2 with it, and robotic manipulation systems like Learning Dexterity depend on it. Its simplicity and robustness made reinforcement learning accessible to practitioners who are not RL specialists.
a policy is like tuning a recipe while the restaurant is open. You taste the dish, adjust the seasoning, and serve again. REINFORCE throws away the dish after one taste — wasteful. TRPO hires a food scientist to calculate the exact safe adjustment — accurate but expensive. PPO is the pragmatic chef: taste, adjust, but never add more than a pinch at a time. If the seasoning change is too large, just stop — don't keep pouring. That "pinch limit" is the mechanism, and it makes PPO both stable and dead simple.
The problem: policy updates that overshoot
In reinforcement learning, a policy maps states to action probabilities. The collects trajectories, computes returns, and updates to increase the probability of good actions. This is the policy idea behind REINFORCE.
The core difficulty: how far should each update go?
- Too small and learning is painfully slow — every trajectory is used once then discarded.
- Too large and the policy can collapse: it overshoots into a bad region, collects bad data from there, and spirals downward. One bad update can undo thousands of good ones.
This is not a theoretical fear. In practice, vanilla policy gradient training is famously unstable, and tuning the is a dark art.
TRPO's answer: the trust region
Policy Optimization solved the overshoot problem by adding a hard constraint: the between the old and new policies must stay below a threshold . This guarantees monotonic improvement in theory — every update either improves the policy or does nothing.
But TRPO pays a price for this guarantee:
- It requires computing the Fisher information matrix (second-order derivatives), making each step expensive.
- The constrained optimization uses conjugate gradient with a line search — complex to implement and debug.
- It is incompatible with architectures that share parameters between policy and value networks, and with regularization techniques like dropout.
PPO asks: can we get TRPO's stability with nothing more than first-order gradients and a clever ?
The core idea: clip the probability ratio
PPO starts from the same surrogate objective that TRPO uses, but replaces the hard constraint with a remarkably simple mechanism. First, define the probability ratio:
This ratio measures how much the new policy differs from the old one for a specific action. When , the policies agree perfectly. When , the new policy assigns more probability to this action; when , less.
The unconstrained surrogate objective would be:
where is the estimated — how much better this action was compared to the average. Maximizing this objective pushes the policy to take good actions more often. But without a constraint, nothing stops from growing unboundedly, leading to destructively large updates.
Why the minimum? Think of it as a pessimistic bound. For actions with positive advantage ( — actions better than average), the clipped version caps the benefit when exceeds : "yes, this action is good, but you've already increased its probability enough." For actions with negative advantage ( — actions worse than average), the clipped version prevents from dropping below : "yes, this action is bad, but don't decrease its probability too fast."
The clip creates a flat region in the objective landscape — a plateau where the gradient is zero. Once you've changed the policy "enough," the optimizer sees no slope and naturally stops. No constraint solver needed, no second-order math, no Fisher matrix. Just a min and a clip.
The ε hyperparameter: how much change is "enough"?
The clipping threshold ε controls the trade-off between stability and speed of learning. A large ε (say 0.3) allows bigger policy changes per update, learning faster but risking instability. A small ε (say 0.1) constrains updates tightly, making training safer but slower. The paper uses ε = 0.2 as the default, which works remarkably well across diverse tasks — from Atari games to simulated robotic locomotion.
Think of ε as the speed limit on a highway. Too low and traffic crawls; too high and accidents happen. The paper found that 0.2 is the sweet spot — fast enough to learn efficiently, slow enough to avoid crashes.
The alternative: adaptive KL penalty
PPO also proposes a second variant — PPO-Penalty — that uses a KL divergence penalty instead of clipping. The idea: add a penalty term to the objective, then adapt automatically:
- If the actual KL divergence after an update is too large (above ), increase — tighten the leash.
- If the actual KL is too small (below ), decrease — loosen the leash.
This way the penalty strength self-tunes. In practice, PPO-Clip tends to be preferred because it is simpler and slightly more robust, but PPO-Penalty can be useful when you want more direct control over policy divergence.
The PPO training loop: simple and GPU-friendly
The full PPO algorithm is beautifully simple. In each iteration:
- Collect — run the current policy in the for timesteps across parallel actors, producing a batch of (state, action, ) tuples.
- Estimate advantages — use to compute for each timestep. GAE balances bias and variance through its .
- Optimize — for epochs, shuffle the batch into minibatches and run SGD on the combined objective:
where is the (training the critic) and is an bonus that encourages . 4. Repeat — set and go to step 1.
The key insight: because the clipped objective prevents destructive updates, you can safely reuse the same batch for multiple epochs () instead of discarding it after one gradient step. This dramatically improves sample efficiency.
The clipped objective in code
Simplified to show the idea — not the real implementation.
import torch
def ppo_clip_objective(
log_probs_new, # log π_θ(a|s) from current policy
log_probs_old, # log π_θ_old(a|s) from data-collection policy
advantages, # Â_t from GAE
epsilon=0.2 # clipping threshold
):
# Probability ratio: how much has the policy changed?
ratio = torch.exp(log_probs_new - log_probs_old) # r(θ)
# Unclipped objective: reward big changes (dangerous!)
unclipped = ratio * advantages
# Clipped objective: cap the ratio to [1-ε, 1+ε]
clipped = torch.clamp(ratio, 1 - epsilon, 1 + epsilon) * advantages
# Take the pessimistic (lower) bound
loss = -torch.mean(torch.min(unclipped, clipped))
return loss # negate because optimizers minimizeResults: beating the competition with simplicity
The paper tested PPO on two domains:
- Continuous control (MuJoCo): simulated robots learning to walk, run, and hop. PPO matched or exceeded TRPO on almost every task, while being much simpler.
- Atari games: 49 classic Arcade games. PPO matched or exceeded the performance of A2C (a synchronous version of A3C) and ACER (an off-policy method) on the majority of games, with a single set of hyperparameters.
The critical result: PPO achieves a favorable balance between sample complexity (how much data it needs), simplicity (how easy it is to implement), and wall-time (how fast training runs). No other algorithm at the time matched all three.
Why PPO became the default
The timeline below shows how PPO became the backbone of modern AI alignment and beyond. What started as a robotics optimization trick turned into the engine behind the most consequential AI systems of the 2020s.
2017
PPO published
Schulman et al. introduce PPO at OpenAI. Immediately adopted for robotics and game playing. Replaces TRPO as the go-to algorithm.
2018
OpenAI Five — Dota 2
PPO trained a team of five neural networks to defeat human world champions at Dota 2, a complex multi-agent strategy game requiring long-horizon planning.
2019
Learning Dexterity — robotic hand
A robotic hand learned to manipulate a Rubik's cube using PPO in simulation, then transferred to a physical robot — a landmark in sim-to-real transfer.
2020
Learning to Summarize with Human Feedback
OpenAI used PPO to fine-tune language models from human preferences — an early blueprint for RLHF that directly led to InstructGPT and ChatGPT.
2022
InstructGPT & ChatGPT
PPO became the RL engine inside RLHF: a reward model scores outputs, and PPO optimizes the language model to maximize human-preferred responses. ChatGPT brought this pipeline to 100 million users.
2025
DeepSeek-R1
DeepSeek-R1 used pure RL (GRPO, a PPO variant without a critic) to train reasoning abilities, showing PPO's influence extends to next-generation reasoning models.
Key design decisions unpacked
CitationSchulman, Wolski, Dhariwal, Radford, Klimov. Proximal Policy Optimization Algorithms. arXiv, 2017.
Terms in this paper
- Proximal Policy Optimizationالتحسين التقريبي للسياسة
- Clippingالقصّ
- Surrogate Objectiveدالة الهدف البديلة
- Trust Regionمنطقة الثقة
- Generalized Advantage Estimation (GAE)التقدير المعمّم للميزة
- Probability Ratioنسبة الاحتمال