AI Safety2022intermediate11 min read

Constitutional AI: Harmlessness from AI Feedback

الذكاء الاصطناعي الدستوري: نحو نماذج آمنة بإشراف ذاتي مبني على مبادئ

Bai, Y. · Kadavath, S. · Kundu, S. · Askell, A. · Kernion, J. · Jones, A. · Chen, A. · Goldie, A. · Amodei, D. · Kaplan, J. — arXiv

The problem

Training AI assistants to be harmless using required tens of thousands of human-labeled examples of harmful outputs — expensive, slow, opaque, and often producing evasive models that simply refuse to engage with sensitive topics rather than addressing them thoughtfully. There was no transparent, scalable way to encode safety rules and have the AI learn to follow them.

The contribution

(CAI): a two-phase method that trains a harmless but non-evasive AI assistant using only a short list of natural-language principles (a "constitution") instead of human labels. Phase 1 (Supervised): the model critiques and revises its own harmful responses according to constitutional principles, then is fine-tuned on the revisions. Phase 2 (RL): the model judges which of two responses is more harmless, creating AI preference labels that train a for — a technique called (RL from AI Feedback). Chain-of-thought reasoning improves both phases.

The impact

CAI became the foundation of Claude's safety training and influenced the entire field of AI alignment. It demonstrated that AI can supervise AI — a crucial step toward . It resolved the -harmlessness tradeoff, proving models could refuse harmful requests while still engaging thoughtfully rather than being evasive. The constitutional approach made safety goals transparent and editable, and RLAIF eliminated the need for massive human annotation campaigns for harmlessness.

Training a safe AI with RLHF is like hiring a team of inspectors to check every response: effective, but expensive, slow, and whatever implicit standard they apply stays hidden in the data.

Constitutional AI replaces the inspectors with a rulebook. Hand the model a short set of written principles — "don't help with illegal activity", "be respectful", "explain your reasoning" — and ask it to grade its own work. The model becomes its own editor: it reads its draft, critiques what violates a rule, rewrites the draft, and eventually learns to judge which of two responses is closer to the spirit of the rulebook.

The result is a model that doesn't just refuse harmful requests — it explains why, like a thoughtful colleague rather than a locked door.

The problem: RLHF doesn't scale for safety

By 2022, Reinforcement Learning from Human Feedback (RLHF) was the standard recipe for making language models helpful and harmless. But RLHF for harmlessness had three painful weaknesses:

  • Cost and speed. Training required tens of thousands of human comparison labels just for harmlessness. Each label needed a crowdworker to read a conversation and judge which response was less harmful — expensive, slow, and hard to iterate on.

  • Opacity. The collective impact of thousands of implicit human judgments was impossible to summarize or audit. You couldn't point to a clear set of safety goals.

  • . Models trained on human harmlessness feedback learned a shortcut: refuse everything remotely sensitive. "I can't answer that" is technically harmless, but completely useless — and it makes the model worse at its primary job of being helpful.

Open in Lab
Compare how RLHF and Constitutional AI handle a sensitive question. Toggle between the two approaches.
The demo wakes as you arrive…

The idea: give the model a constitution

The insight behind Constitutional AI is simple: instead of hiring thousands of annotators to demonstrate safe behavior implicitly, write down the rules explicitly and let the AI learn to follow them on its own.

The "constitution" is a small set of natural-language principles — roughly 16 — covering different aspects of harmlessness. For example: "Identify specific ways in which the response is harmful, unethical, racist, sexist, toxic, dangerous, or illegal", or "Choose the response that a wise, ethical, polite and friendly person would more likely say."

These principles serve two purposes. In the supervised phase they guide critique and revision. In the RL phase they guide preference labeling. The same small set of human-readable rules replaces tens of thousands of opaque labels.

Phase 1: critique, revise, learn

The first phase turns a helpful-but-unsafe model into a helpful-and-harmless one through a loop. Think of it as the model editing its own rough drafts until they meet the constitution's standards.

The process works in four steps:

Step 1 — Generate harmful responses. Start with a helpful-only RLHF model and feed it red-team prompts — adversarial questions designed to bait harmful answers. The model, having been trained only for helpfulness, obliges.

Step 2 — Critique. Append a constitutional principle to the conversation and ask the model: "What's wrong with your response?" The model writes a critique identifying specific harms.

Step 3 — Revise. Ask the model to rewrite its response to fix the problems the critique identified. A different principle can be sampled at each revision step, creating diversity.

Step 4 — Fine-tune. Collect the final revised (prompt, response) pairs and fine-tune a pretrained model on them. This produces the SL-CAI model.

Open in Lab
Walk through the critique-revision pipeline step by step. Watch a harmful response get transformed into a thoughtful, non-evasive one.
The demo wakes as you arrive…

Crucially, the revisions can be applied multiple times in sequence. Each round uses a different randomly-sampled principle, progressively removing more harm. The paper found that the first revision removes most harmful content, while later revisions make smaller improvements.

A key observation: the revised responses were rarely evasive. Instead of shutting down sensitive conversations, the model learned to engage thoughtfully — explaining why a request is problematic while still being helpful.

Open in Lab
Drag the slider to see how harmlessness and helpfulness scores change with each revision round.
The demo wakes as you arrive…
SL-CAI=FineTune(θpretrained,  {(xi,rirevised)}i=1N)\text{SL-CAI} = \text{FineTune}\bigl(\theta_{\text{pretrained}},\; \{(x_i, r_i^{\text{revised}})\}_{i=1}^{N}\bigr)
SL-CAI — the supervised constitutional model — Fine-tune a pretrained model θ on revised (prompt, response) pairs. The revised responses come from the critique-revision pipeline, not from human annotations.

Phase 2: RL from AI Feedback (RLAIF)

The supervised phase gets the model into the right neighborhood of behavior. The RL phase refines it further — the same way RLHF works, but with AI judges instead of human ones.

The key idea is to use the model itself as a harmlessness evaluator, guided by constitutional principles. The process mirrors standard RLHF but replaces human comparison labels for harmlessness with AI-generated ones:

Step 1 — Generate response pairs. For each red-team prompt, sample two responses from the SL-CAI model.

Step 2 — AI evaluation. Present each pair to a feedback model with a constitutional principle in a multiple-choice format: "Which response is less harmful?" The model's log probabilities become soft preference labels. A different principle is randomly sampled for each comparison, creating an ensemble effect.

Step 3 — Train a . Combine the AI harmlessness labels with human helpfulness labels to train a preference model (PM).

Step 4 — RL training. Use the preference model as a to train the final RL-CAI model via reinforcement learning.

Open in Lab
Click each stage to see how data flows through the full Constitutional AI pipeline — from red-team prompt to final RL-CAI model.
The demo wakes as you arrive…
P(A≻B∣principlej)=elog⁡P(A)elog⁡P(A)+elog⁡P(B)P(A \succ B \mid \text{principle}_j) = \frac{e^{\log P(A)}}{e^{\log P(A)} + e^{\log P(B)}}
AI preference label — soft probability from constitutional evaluation — The feedback model assigns a soft probability that response A is preferred over B, given constitutional principle j. These calibrated probabilities serve as training targets for the preference model — no human labels needed for harmlessness.

Chain-of-thought: making AI reasoning transparent

Both phases can be enhanced with chain-of-thought (CoT) reasoning — asking the model to "think step by step" before making a judgment. In the RL phase, instead of directly choosing which response is less harmful, the model first writes out an argument explaining its reasoning, then commits to a choice.

CoT has two benefits. First, it improves accuracy: larger models with CoT approach the performance of preference models trained on hundreds of thousands of human labels. Second, it makes the AI's decision-making transparent and auditable — you can read why the model preferred one response over another.

One subtlety: CoT reasoning produces very confident labels (near 0 or 1) because the step-by-step argument commits the model to a choice. The authors found that clamping probabilities to the 40–60% range produced better, less extreme model behavior during RL training.

Open in Lab
See how chain-of-thought reasoning helps the model choose between two responses. Click "Show Reasoning" to reveal the step-by-step argument.
The demo wakes as you arrive…

Results: harmless without being evasive

The paper evaluated models using Elo scores from crowdworker preference tests — the same framework used in chess ratings. The key findings:

  • RL-CAI beat human-feedback RLHF on harmlessness while maintaining comparable helpfulness. This was the core result: AI supervision matched or exceeded human supervision for safety.

  • The helpfulness-harmlessness tradeoff was reduced. Standard RLHF models showed a painful tradeoff — more harmless meant less helpful. Constitutional AI models achieved a : more harmless and more helpful at the same time.

  • Non-evasiveness. When asked "Why are women bad at science and math?", the RLHF model responded "I'm sorry. I won't respond." The RL-CAI model explained why the premise is a harmful stereotype, provided factual context, and condemned bigotry — a far more useful and educational response.

  • Chain-of-thought helped. RL-CAI with CoT was slightly more harmless than without, though slightly less helpful.

Open in Lab
Explore the helpfulness vs. harmlessness Pareto frontier. Each point is a model snapshot during RL training.
The demo wakes as you arrive…

The paper explored two important design questions:

Are critiques necessary? The authors compared the full critique→revision pipeline against direct revision (skipping the critique). For smaller models, critiques significantly improved harmlessness scores. For larger models (52B parameters), the difference was small — the large model could revise well without explicit critique. However, critiques were kept because they provide transparency into the model's reasoning process.

How many principles? Increasing the number of constitutional principles from 1 to 16 did not significantly change harmlessness scores, but it increased response diversity — which is valuable for during RL training.

AI evaluation scales with model size. Larger feedback models produced better preference labels. With chain-of-thought, models approached the accuracy of preference models trained on hundreds of thousands of human labels.

The same idea in code

Constitutional critique-revision loop (simplified)python

Simplified to show the idea — not the real implementation.

import random

CONSTITUTION = [
    "Identify specific ways in which this response is harmful or unethical.",
    "Rewrite to remove any dangerous, illegal, or toxic content.",
    "Choose the response a wise, ethical, polite person would more likely say.",
    # ... ~16 principles total
]

def critique_and_revise(model, prompt, response, n_revisions=4):
    """Apply constitutional critique-revision loop."""
    current = response
    for i in range(n_revisions):
        principle = random.choice(CONSTITUTION)

        # Step 1: Critique — ask model what's wrong
        critique = model.generate(
            f"{prompt}\nAssistant: {current}\n"
            f"Critique Request: {principle}\nCritique:"
        )

        # Step 2: Revise — ask model to fix it
        revision = model.generate(
            f"{prompt}\nAssistant: {current}\n"
            f"Critique: {critique}\n"
            f"Revision Request: Please rewrite to remove harmful content.\n"
            f"Revision:"
        )
        current = revision  # each round builds on the last

    return current  # final revised response for SL fine-tuning

# After collecting revised responses for all red-team prompts,
# fine-tune a pretrained model on (prompt, revised_response) pairs.
# This produces the SL-CAI model — the starting point for RL.

Why it mattered

  1. 2017

    RLHF for language models

    Christiano et al. introduced deep RL from human preferences — training a reward model from human comparisons and using it to guide policy optimization.

  2. 2020

    Learning to summarize from human feedback

    Stiennon et al. scaled RLHF to summarization, establishing the preference model → RL pipeline used by later alignment work.

  3. 2022

    Training a Helpful and Harmless Assistant (HH-RLHF)

    Bai et al. identified the helpfulness-harmlessness tradeoff and the evasiveness problem in RLHF-trained models — the direct predecessors to Constitutional AI.

  4. 2022

    Constitutional AI (this paper)

    Replaced human harmlessness labels with constitutional principles and AI self-evaluation. Introduced RLAIF and demonstrated that AI feedback can match or exceed human feedback.

  5. 2023

    Claude — Constitutional AI in production

    Anthropic deployed Claude, built on Constitutional AI principles. The model demonstrated that CAI could scale to production quality while remaining helpful, harmless, and honest.

  6. 2024

    RLAIF becomes industry standard

    Multiple labs adopted AI feedback techniques for alignment training. Constitutional-style approaches became a standard tool in the alignment toolkit alongside RLHF.

Constitutional AI sits at the intersection of capability and safety. It showed that the frontier of both can advance together — not as a zero-sum game. Its descendants include Transformer-based systems that use constitutional principles at scale, and the RL from AI Feedback paradigm it pioneered continues to evolve.

CitationBai, Kadavath, Kundu, Askell, Kernion, Jones, Chen, Goldie, Mirhoseini, McKinnon, et al.. Constitutional AI: Harmlessness from AI Feedback. arXiv, 2022.

Terms in this paper