Reinforcement Learning2023intermediate10 min read

RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

RLAIF: توسيع التعلّم المعزَّز بتغذية راجعة آلية بدلاً من البشر

Lee, H. · Phatale, S. · Mansoor, H. · Mesnard, T. · Ferret, J. · Lu, K. · Bishop, C. · Hall, E. · Carbune, V. · Rastogi, A. · Prakash, S. — ICML

The problem

is the dominant method for aligning LLMs with human preferences — it powers ChatGPT, Bard, and most modern conversational AI. But it depends on thousands of human annotators who read pairs of AI responses and label which is better. This human labeling process is expensive, slow, noisy, and hard to scale. As models improve and the number of tasks grows, the bottleneck tightens: you need more labels, faster, across more domains.

The contribution

A head-to-head comparison showing that — where an off-the-shelf LLM labels preferences instead of humans — matches RLHF performance. On , humans prefer both RLAIF and RLHF over SFT ~70% of the time, and prefer RLAIF vs RLHF at equal rates. The paper also introduces direct-RLAIF (d-RLAIF), which skips entirely by using the LLM as a live reward signal during RL, achieving even better results. Additionally, it demonstrates : RLAIF works even when the labeler is the same size or same as the .

The impact

RLAIF showed the AI community that human labeling is not the only path to . It validated AI-as-judge at scale and inspired a wave of follow-up work — from refinements to DPO and self-play methods. By making alignment 10x cheaper, it democratized access to RLHF-quality training for teams without massive annotation budgets. The d-RLAIF variant foreshadowed the move toward online, reward-model-free alignment that defines the current frontier.

Imagine a cooking school where every student dish must be judged by a panel of professional food critics before the student can improve. The critics are expensive, slow, and often disagree with each other. One day, the school discovers that an experienced chef — who has tasted thousands of dishes — can judge student dishes almost as well as the critics.

Better yet, the chef can taste dishes live during the cooking class, giving feedback in real time instead of waiting for the critics' scorecard. RLAIF replaces the panel of human critics with an AI chef — and the dishes turn out just as good.

The bottleneck: human labels do not scale

The standard RLHF pipeline has three stages. First, a language model is supervised fine-tuned (SFT) on high-quality demonstrations. Second, human annotators compare pairs of model responses and a reward model (RM) is trained on their preferences. Third, the SFT policy is optimized with using the RM as a proxy for human judgment, with a penalty to prevent the policy from drifting too far from the original model.

The second stage is the bottleneck. Collecting human preferences is expensive — roughly $0.67 per example for the summarization task. Annotators must be trained, they disagree with each other (inter-annotator agreement is only 73–77%), and the process takes weeks. For every new task or domain, you start over. As models get better, they need more nuanced feedback, which requires more skilled — and more expensive — annotators.

Open in Lab
Compare the RLHF and RLAIF pipelines side by side. The only difference is who labels preferences.
The demo wakes as you arrive…

How RLAIF works: replacing humans with an LLM judge

The core idea is simple: instead of asking a human "Which response is better?", ask a . Given a context (e.g. a Reddit post) and two candidate responses (e.g. two summaries), the LLM is prompted to evaluate and express a preference. The prompt has four parts: a preamble describing the task, optional few-shot exemplars, the sample to annotate, and an ending that elicits the preference.

Rather than extracting a binary label, the system computes a soft preference distribution by taking the of the log-probabilities of generating the tokens "1" and "2". For example, a preference of [0.6, 0.4] means the LLM slightly prefers the first response. This soft label conveys richer information than a hard one-hot choice.

Open in Lab
See how the LLM produces a soft preference distribution instead of a hard binary label.
The demo wakes as you arrive…

Two critical techniques improve the quality of AI-generated labels:

mitigation. LLMs can be biased toward whichever response appears first (or second). To counter this, each pair is scored twice with the order reversed, and the results are averaged. Smaller models show dramatically more position bias — PaLM 2 XS prefers the same position 56% of the time even after swapping, compared to only 18% for PaLM 2 L.

Chain-of-thought reasoning. Before emitting a preference, the LLM is asked to explain its reasoning. This two-step approach — first generate a rationale, then score — consistently improves alignment with human preferences by up to +1.9%.

Open in Lab
Explore how position bias varies by model size and how order-swapping mitigates it.
The demo wakes as you arrive…

From preferences to policy: reward modeling and RL

Once the AI-generated preferences are collected, the standard RLHF pipeline resumes. A reward model is trained on the soft AI labels using cross-entropy loss — the same loss used in RLHF, except the targets are soft distributions instead of hard labels.

The RL phase uses with a baseline value function, initialized from the SFT checkpoint. The optimization objective balances reward maximization against KL divergence from the SFT policy:

J(θ)=Ey∼πθ(⋅∣x)[(1−β) rϕ(y∣x)−β DKL(πθRL(y∣x)∥πSFT(y∣x))]J(\theta) = \mathbb{E}_{y \sim \pi_\theta(\cdot|x)} \Big[ (1-\beta)\, r_\phi(y|x) - \beta\, D_{\text{KL}}\big(\pi_\theta^{RL}(y|x) \| \pi^{SFT}(y|x)\big) \Big]
RLAIF / RLHF Objective — The policy π_θ generates response y given input x. The reward r_ϕ comes from the RM (trained on either human or AI preferences). β controls how much the policy is penalized for diverging from the SFT baseline — preventing reward hacking where the model exploits RM quirks.

Direct-RLAIF: skipping the reward model entirely

Canonical RLAIF still trains a separate reward model — a process that takes time and introduces a "staleness" problem. As the RL policy improves, its outputs drift away from the data the RM was trained on, causing the RM's scores to become unreliable. Iteratively retraining the RM is one fix, but it is expensive.

Direct-RLAIF (d-RLAIF) offers an elegant alternative: skip the reward model entirely and use the off-the-shelf LLM as a live scorer during RL. For each generated response, the LLM is prompted to rate its quality on a scale of 1–10. The probability of each score token is computed, a weighted average gives the raw score, and finally the score is normalized to [−1, 1] and used directly as the reward signal.

This eliminates both the staleness issue and the time cost of RM training. In experiments, d-RLAIF matched or outperformed canonical RLAIF on every task. On summarization, d-RLAIF achieved a 74% over SFT, compared to 68% for same-size canonical RLAIF.

Open in Lab
Compare canonical RLAIF (with RM training) vs. d-RLAIF (direct scoring). d-RLAIF removes the RM bottleneck and avoids reward staleness.
The demo wakes as you arrive…
s(y∣x)=∑i=110i⋅P(i∣y,x)s(y|x) = \sum_{i=1}^{10} i \cdot P(i \mid y, x)
d-RLAIF Scoring — The LLM is prompted to rate a response from 1 to 10. The probability P(i|y,x) of each score token i is computed, and the weighted sum gives the raw score. This is then normalized to [−1, 1] and used as the reward during RL.

Results: RLAIF matches RLHF across all tasks

The paper evaluates on three tasks — summarization (Reddit TL;DR), helpful dialogue generation, and harmless dialogue generation — using human evaluation as the gold standard.

For summarization, RLAIF achieves a 71% win rate over SFT while RLHF achieves 73% — a difference that is not statistically significant. In a direct head-to-head comparison, RLAIF and RLHF are preferred at equal rates by human evaluators (50% win rate).

For helpful dialogue, the pattern holds: RLAIF at 63% and RLHF at 64% over SFT, again not statistically different.

For harmless dialogue, RLAIF actually outperforms RLHF, achieving a harmless rate of 88% compared to RLHF's 76% and SFT's 64%.

Open in Lab
Human evaluation results across all three tasks — RLAIF consistently matches or beats RLHF.
The demo wakes as you arrive…

Scaling the AI labeler and self-improvement

A key question is how the quality of the AI labeler affects downstream performance. The paper finds a clear scaling relationship: larger LLM labelers produce better-aligned preferences. PaLM 2 L achieves 78% alignment with human labels, PaLM 2 S drops to 73.8%, and PaLM 2 XS falls to 62.7%.

But here is the surprising finding: even when the AI labeler is the same size as the policy being trained (PaLM 2 XS labeling for PaLM 2 XS policy), RLAIF still achieves a 68% win rate over SFT — only slightly below the 71% achieved with a much larger labeler. This is a step toward self-improvement: a model improving itself using its own judgments.

The d-RLAIF setup pushes this even further. For helpful dialogue, the LLM providing rewards and the starting policy are the exact same model checkpoint — and it still improves. This constitutes a strict example of LLM self-improvement.

Open in Lab
As the AI labeler size increases, alignment with human preferences improves predictably.
The demo wakes as you arrive…

Prompting techniques: what works and what does not

The paper systematically studies three prompt variations for AI labeling:

Chain-of-thought (CoT) consistently improves alignment. By asking the LLM to first explain its reasoning before expressing a preference, alignment improves by +1.3% to +1.9% depending on the task. This mirrors findings from other reasoning tasks — thinking before answering helps.

Detailed preambles help for complex tasks (summarization gains +1.3%) but show mixed results for simpler tasks like judging helpfulness or harmlessness, where the task is already intuitive for the LLM.

Few-shot examples surprisingly hurt performance on summarization and helpfulness, with alignment monotonically decreasing as exemplar count increases. Only the harmlessness task benefits from (+2.7% with 2-shot). The hypothesis is that the LLM already understands these tasks well enough that exemplars become distracting rather than helpful.

Open in Lab
AI labeler alignment across different prompting strategies and tasks.
The demo wakes as you arrive…

Cost: 10x cheaper than human annotation

Beyond matching quality, RLAIF dramatically reduces cost. The paper estimates that AI labeling costs approximately $0.06 per example (using GPT-4 pricing at the time), compared to $0.67 per example for human annotation through Google Cloud's labeling service. That is over 10x cheaper — and the AI labeler works around the clock without fatigue, training time, or inter-annotator calibration.

Since the AI labeler in canonical RLAIF is only used once to generate the preference dataset (not called during RL), using an even larger, more capable AI labeler is not prohibitively expensive. The cost is amortized across all the preferences it labels.

Qualitative observations

Manual inspection of summarization outputs revealed interesting differences between RLAIF and RLHF policies. In some cases, RLHF hallucinated plausible-sounding details not present in the source text — for example, stating a user was "20 years old" when no age was mentioned. RLAIF summaries were less prone to this.

On the other hand, RLAIF sometimes produced less fluent summaries — run-on sentences or repeated phrases like "How do I get over this?" that did not reflect the original text. However, when human annotators blindly ranked the policies on accuracy, coverage, and coherence, the differences were not statistically significant.

What RLAIF enabled

  1. 2017

    RLHF for dialogue

    Christiano et al. introduce deep RL from human preferences, training reward models from pairwise comparisons. The foundation for all modern alignment work.

  2. 2022

    InstructGPT — RLHF at scale

    Ouyang et al. apply the SFT → RM → RL pipeline to GPT-3, creating InstructGPT. Demonstrates that RLHF can make models dramatically more helpful and honest. Powers ChatGPT.

  3. 2022

    Constitutional AI introduces RLAIF concept

    Bai et al. at Anthropic combine AI feedback with a "constitution" of principles. First use of LLM-generated preferences for RM training, but mixed with human labels.

  4. 2023

    RLAIF — this paper

    Lee et al. provide the first head-to-head comparison showing RLAIF matches RLHF. Introduce d-RLAIF and demonstrate self-improvement. AI labeling shown to be 10x cheaper.

  5. 2023

    DPO — Direct Preference Optimization

    Rafailov et al. show the RL objective can be reformulated as a classification loss, eliminating the RM entirely. Combines naturally with AI-generated preferences from RLAIF.

  6. 2024

    Self-play and AI judges become standard

    Building on RLAIF's validation, techniques like self-play, LLM-as-judge, and reward-model-free alignment become widespread in both open-source and commercial models.

RLAIF's deepest contribution is not a specific method — it is a proof of concept. It showed that the boundary between human judgment and AI judgment, for the purpose of training reward models, is far thinner than assumed. This opened the door to a new generation of alignment methods that reduce or eliminate human annotation entirely, from DPO to self-play to Constitutional AI's iterative refinement. The practical implication is clear: alignment at scale is no longer gated by the speed of human annotators.

CitationLee, Phatale, Mansoor, Mesnard, Ferret, Lu, Bishop, Hall, Carbune, Rastogi, Prakash. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. ICML, 2024.

Terms in this paper