Reinforcement Learning2017intermediate10 min read
Deep Reinforcement Learning from Human Preferences
التعلم المعزز العميق من التفضيلات البشرية
Christiano, P. · Leike, J. · Brown, T. · Martic, M. · Legg, S. · Amodei, D. — NeurIPS
The problem
agents need a to learn, but many real-world goals — "clean a table," "do a backflip," "drive safely" — are nearly impossible to express as a mathematical formula. Hand-crafting rewards often causes : the finds a loophole that maximizes the score without achieving the intended behavior. And using a human to provide direct reward signals at every timestep would require thousands of hours of continuous supervision.
The contribution
A three-part asynchronous system that learns complex behaviors from human preferences alone. (1) A interacts with the ; (2) a human compares short pairs of segments (1–2 second clips); (3) a is trained via the on those comparisons. The policy then optimizes this learned reward using standard RL. With just 700 comparisons (~1 hour of human time), the system matches hand-crafted reward RL on MuJoCo robotics, and even learns novel behaviors like backflips that have no predefined reward function at all.
The impact
This paper is the conceptual ancestor of RLHF — the technique that transformed raw language models (GPT-3) into conversational assistants (ChatGPT, Claude). It proved that human preferences, not hand-written rules, could scale as a training signal for deep learning. The reward-from-comparisons framework was later adopted by Learning to Summarize, InstructGPT, Constitutional AI, and DPO, each extending the idea from simulated robotics to language, safety, and beyond.
Traditional RL is like training a dog with a clicker and treat counter — you need a precise formula for every micro-behavior. But what if you want the dog to do something you can recognize but can't quantify, like "look cute"?
This paper replaces the clicker with a talent show judge: you watch two short performances, point to the one you liked more, and the system figures out what you're rewarding. After a few hundred comparisons — roughly an hour of pointing — the dog is doing backflips.
The problem: reward functions are hard to write
Reinforcement learning works beautifully when you have a clean reward signal — points in Pong, distance walked in a simulator. But most real-world goals resist mathematical specification. Consider "clean a table" or "drive safely": writing a formula that captures every nuance of the goal is at best tedious and at worst impossible.
When we try to approximate the goal with a simple formula, the agent often finds loopholes. This is reward hacking — the agent maximizes the number but not the intent. The Concrete Problems in AI Safety paper catalogued these failure modes and called for better approaches. This paper provides one.
The idea: ask a human to compare, not to score
The key insight is that humans are much better at comparing two behaviors than scoring a single one. Which clip looks better? That's easy. What's the exact reward for this frame? That's nearly impossible.
The system shows a human two short video clips (1–2 seconds) of the agent's behavior and asks: "Which do you prefer?" The human can also say "about equal" or "I can't tell." These comparison labels are far cheaper than per-timestep rewards, and they carry rich information about the human's intent.
Critically, this works with non-expert humans. The contractors labeling preferences had no knowledge of RL or the algorithm — they just watched clips and pointed.
The system: three asynchronous loops
The algorithm runs three processes simultaneously, like three musicians in a jazz trio who listen to each other and adjust in real time:
- Process 1 — The Policy interacts with the environment, producing trajectories. It uses whatever RL algorithm fits the domain: PPO/TRPO for robotics, A2C for Atari. But instead of maximizing the real reward (which it never sees), it maximizes the predicted reward from the reward model.
- Process 2 — The Human watches pairs of short clips selected from those trajectories and labels which one is better. An -based strategy selects the most informative pairs — the ones where the reward model's own predictors disagree the most.
- Process 3 — The Reward Model is a neural network trained via on the human's comparisons. It learns to predict which of two clips a human would prefer, and its internal reward prediction becomes the signal the policy optimizes.
How comparisons become rewards: the Bradley-Terry model
The mathematical bridge from "I prefer clip A" to a numerical reward function is the Bradley-Terry model — a classic framework from the 1950s originally used to rank chess players. The idea is elegant: if clip A's total reward is higher, a human is more likely to prefer it. Specifically, the probability of preferring segment 1 over segment 2 is modeled as:
Read this as a competition: each segment's total reward is like a contestant's score. The softmax converts those scores into a probability — the segment with higher total reward wins more often, but not always, mirroring the noise in real human judgments.
The reward model is then trained by minimizing the between these predicted probabilities and the actual human labels. This is identical to training a classifier, except the "classes" are preference orderings.
Three important practical refinements make this work at scale: (1) an ensemble of reward predictors is trained, with their predictions averaged for stability; (2) 10% label noise is assumed — modeling the fact that humans sometimes make mistakes; (3) prevents the reward model from to the limited comparisons.
Smart querying: ask the most informative questions
Human time is precious. Instead of showing random pairs, the system uses to find the most informative comparisons. Think of it as a committee of reward predictors: when they all agree on which clip is better, there's nothing to learn from asking. When they disagree, that's where human input has the most value.
The system samples many candidate pairs from recent trajectories, has each ensemble member predict which segment is preferred, and selects the pairs with highest prediction variance. Ablation experiments showed this active query selection improves sample efficiency, though in some tasks random selection works nearly as well.
Results: matching hand-crafted rewards with an hour of pointing
The paper tested on two domains: MuJoCo simulated robotics (Hopper, Walker, Cheetah, Ant, and others) and Atari games (Pong, Breakout, BeamRider, Enduro, and others).
On MuJoCo, 700 human comparisons — roughly one hour — were enough to nearly match the performance of RL trained with the true reward function. Surprisingly, on the Ant task, actually outperformed the true reward, because humans gave better-shaped feedback by rewarding "standing upright" in a way the hand-crafted reward didn't capture as well.
On Atari, the results were more mixed: the method matched or approached true-reward RL on BeamRider and Pong, but struggled on games like Qbert where short clips are confusing to evaluate. On Enduro, human labelers outperformed synthetic feedback by rewarding partial progress (approaching other cars), essentially providing better than the sparse game score.
Beyond benchmarks: teaching novel behaviors
The real power of the approach shines when there is no predefined reward function at all. The authors demonstrated three novel behaviors that no RL benchmark captures:
- Backflips on Hopper: 900 queries, less than an hour. The robot learns to consistently perform a backflip, land upright, and repeat — a behavior impossible to specify with a simple reward formula.
- One-legged running on Half-Cheetah: 800 queries, under an hour. The robot runs forward while balancing on a single leg.
- Driving alongside traffic on Enduro: ~1,300 queries. The agent learns to stay even with other cars rather than pass them — a behavior that contradicts the game's scoring system.
These experiments show that human preferences can define goals that are beyond the reach of hand-crafted reward functions while still using standard RL infrastructure under the hood.
What matters: ablation studies
The authors systematically removed components to measure their contribution:
- Online queries are essential. Training the reward model only on early data leads to bizarre behaviors — on Pong, the agent learned to avoid losing points without ever scoring them, producing infinitely long volleys. The reward model must see current behavior to stay calibrated.
- Comparisons beat absolute scores on continuous control tasks, because the scale of rewards varies across tasks and complicates . For Atari, where rewards are clipped, comparisons and scores performed comparably.
- Trajectory segments beat single frames. Longer clips (1–2 seconds) provide context that makes comparisons easier and more informative per unit of human time.
- Ensembles help but are not critical everywhere. On tasks like Hopper they improve stability significantly; on simpler tasks the effect is modest.
The reward model in code
Simplified to show the idea — not the real implementation.
import numpy as np
def preference_loss(reward_model, segment1, segment2, human_label):
"""
Compute cross-entropy loss for a single preference comparison.
human_label: 0 = prefer seg1, 1 = prefer seg2, 0.5 = tie
"""
# Sum predicted rewards over each trajectory segment
r1 = sum(reward_model(obs, act) for obs, act in segment1)
r2 = sum(reward_model(obs, act) for obs, act in segment2)
# Bradley-Terry: P(seg1 preferred) = exp(r1) / (exp(r1) + exp(r2))
log_p1 = r1 - np.logaddexp(r1, r2) # log P(prefer seg1)
log_p2 = r2 - np.logaddexp(r1, r2) # log P(prefer seg2)
# Cross-entropy against human label
# 10% uniform noise to model human error
mu1 = 0.9 * (1 - human_label) + 0.05
mu2 = 0.9 * human_label + 0.05
loss = -(mu1 * log_p1 + mu2 * log_p2)
return lossWhy it mattered
2017
Deep RL from Human Preferences (this paper)
Proved that ~700 comparisons from non-expert humans are enough to train RL agents on complex MuJoCo and Atari tasks, and can learn novel behaviors like backflips.
2019
Fine-Tuning Language Models from Human Preferences
Extended the framework from robotics to text generation, using human preferences to fine-tune GPT-2 for summarization and continuation.
2020
Learning to Summarize with Human Feedback
Applied preference-based reward learning to summarization at scale, proving the framework generalizes to language tasks and outperforms supervised fine-tuning alone.
2022
InstructGPT
Scaled the preference learning pipeline to GPT-3, adding supervised fine-tuning before RLHF. The direct ancestor of ChatGPT.
2022
Constitutional AI
Replaced human labelers with an AI that critiques and revises its own outputs using written principles — reducing human cost even further while maintaining alignment.
2023
Direct Preference Optimization (DPO)
Eliminated the separate reward model entirely, showing that a language model is implicitly a reward model and can be optimized directly from preference data.
The path from this paper to Claude and ChatGPT is direct. Christiano et al. showed that human preferences are a scalable training signal; Learning to Summarize proved it works for language; InstructGPT scaled it to production. Every time you talk to an AI assistant that understands your intent, the ancestry traces back to this 2017 paper and its simple insight: comparing is easier than scoring.
CitationChristiano, Leike, Brown, Martic, Legg, Amodei. Deep Reinforcement Learning from Human Preferences. NeurIPS, 2017.
Terms in this paper
- Reward Modelنموذج المكافأة
- Human Feedbackالتغذية الراجعة البشرية
- Preference Pairزوج التفضيل
- Reward Hackingاختراق المكافأة
- Bradley-Terry Modelنموذج برادلي-تيري
- Reward Shapingصياغة وهندسة دالة المكافأة
- Ensemble Disagreementاختلاف المجموعة
- Active Learningالتعلم النشط الموجه