AI Safety2022intermediate11 min read
Constitutional AI: Harmlessness from AI Feedback
الذكاء الاصطناعي الدستوري: نحو نماذج آمنة بإشراف ذاتي مبني على مبادئ
Bai, Y. · Kadavath, S. · Kundu, S. · Askell, A. · Kernion, J. · Jones, A. · Chen, A. · Goldie, A. · Amodei, D. · Kaplan, J. — arXiv
The problem
Training AI assistants to be harmless using required tens of thousands of human-labeled examples of harmful outputs — expensive, slow, opaque, and often producing evasive models that simply refuse to engage with sensitive topics rather than addressing them thoughtfully. There was no transparent, scalable way to encode safety rules and have the AI learn to follow them.
The contribution
(CAI): a two-phase method that trains a harmless but non-evasive AI assistant using only a short list of natural-language principles (a "constitution") instead of human labels. Phase 1 (Supervised): the model critiques and revises its own harmful responses according to constitutional principles, then is fine-tuned on the revisions. Phase 2 (RL): the model judges which of two responses is more harmless, creating AI preference labels that train a for — a technique called (RL from AI Feedback). Chain-of-thought reasoning improves both phases.
The impact
CAI became the foundation of Claude's safety training and influenced the entire field of AI alignment. It demonstrated that AI can supervise AI — a crucial step toward . It resolved the -harmlessness tradeoff, proving models could refuse harmful requests while still engaging thoughtfully rather than being evasive. The constitutional approach made safety goals transparent and editable, and RLAIF eliminated the need for massive human annotation campaigns for harmlessness.
Training a safe AI with RLHF is like hiring a team of inspectors to check every response: effective, but expensive, slow, and whatever implicit standard they apply stays hidden in the data.
Constitutional AI replaces the inspectors with a rulebook. Hand the model a short set of written principles — "don't help with illegal activity", "be respectful", "explain your reasoning" — and ask it to grade its own work. The model becomes its own editor: it reads its draft, critiques what violates a rule, rewrites the draft, and eventually learns to judge which of two responses is closer to the spirit of the rulebook.
The result is a model that doesn't just refuse harmful requests — it explains why, like a thoughtful colleague rather than a locked door.
The problem: RLHF doesn't scale for safety
By 2022, Reinforcement Learning from Human Feedback (RLHF) was the standard recipe for making language models helpful and harmless. But RLHF for harmlessness had three painful weaknesses:
-
Cost and speed. Training required tens of thousands of human comparison labels just for harmlessness. Each label needed a crowdworker to read a conversation and judge which response was less harmful — expensive, slow, and hard to iterate on.
-
Opacity. The collective impact of thousands of implicit human judgments was impossible to summarize or audit. You couldn't point to a clear set of safety goals.
-
. Models trained on human harmlessness feedback learned a shortcut: refuse everything remotely sensitive. "I can't answer that" is technically harmless, but completely useless — and it makes the model worse at its primary job of being helpful.
The idea: give the model a constitution
The insight behind Constitutional AI is simple: instead of hiring thousands of annotators to demonstrate safe behavior implicitly, write down the rules explicitly and let the AI learn to follow them on its own.
The "constitution" is a small set of natural-language principles — roughly 16 — covering different aspects of harmlessness. For example: "Identify specific ways in which the response is harmful, unethical, racist, sexist, toxic, dangerous, or illegal", or "Choose the response that a wise, ethical, polite and friendly person would more likely say."
These principles serve two purposes. In the supervised phase they guide critique and revision. In the RL phase they guide preference labeling. The same small set of human-readable rules replaces tens of thousands of opaque labels.
Phase 1: critique, revise, learn
The first phase turns a helpful-but-unsafe model into a helpful-and-harmless one through a loop. Think of it as the model editing its own rough drafts until they meet the constitution's standards.
The process works in four steps:
Step 1 — Generate harmful responses. Start with a helpful-only RLHF model and feed it red-team prompts — adversarial questions designed to bait harmful answers. The model, having been trained only for helpfulness, obliges.
Step 2 — Critique. Append a constitutional principle to the conversation and ask the model: "What's wrong with your response?" The model writes a critique identifying specific harms.
Step 3 — Revise. Ask the model to rewrite its response to fix the problems the critique identified. A different principle can be sampled at each revision step, creating diversity.
Step 4 — Fine-tune. Collect the final revised (prompt, response) pairs and fine-tune a pretrained model on them. This produces the SL-CAI model.
Crucially, the revisions can be applied multiple times in sequence. Each round uses a different randomly-sampled principle, progressively removing more harm. The paper found that the first revision removes most harmful content, while later revisions make smaller improvements.
A key observation: the revised responses were rarely evasive. Instead of shutting down sensitive conversations, the model learned to engage thoughtfully — explaining why a request is problematic while still being helpful.
Phase 2: RL from AI Feedback (RLAIF)
The supervised phase gets the model into the right neighborhood of behavior. The RL phase refines it further — the same way RLHF works, but with AI judges instead of human ones.
The key idea is to use the model itself as a harmlessness evaluator, guided by constitutional principles. The process mirrors standard RLHF but replaces human comparison labels for harmlessness with AI-generated ones:
Step 1 — Generate response pairs. For each red-team prompt, sample two responses from the SL-CAI model.
Step 2 — AI evaluation. Present each pair to a feedback model with a constitutional principle in a multiple-choice format: "Which response is less harmful?" The model's log probabilities become soft preference labels. A different principle is randomly sampled for each comparison, creating an ensemble effect.
Step 3 — Train a . Combine the AI harmlessness labels with human helpfulness labels to train a preference model (PM).
Step 4 — RL training. Use the preference model as a to train the final RL-CAI model via reinforcement learning.
Chain-of-thought: making AI reasoning transparent
Both phases can be enhanced with chain-of-thought (CoT) reasoning — asking the model to "think step by step" before making a judgment. In the RL phase, instead of directly choosing which response is less harmful, the model first writes out an argument explaining its reasoning, then commits to a choice.
CoT has two benefits. First, it improves accuracy: larger models with CoT approach the performance of preference models trained on hundreds of thousands of human labels. Second, it makes the AI's decision-making transparent and auditable — you can read why the model preferred one response over another.
One subtlety: CoT reasoning produces very confident labels (near 0 or 1) because the step-by-step argument commits the model to a choice. The authors found that clamping probabilities to the 40–60% range produced better, less extreme model behavior during RL training.
Results: harmless without being evasive
The paper evaluated models using Elo scores from crowdworker preference tests — the same framework used in chess ratings. The key findings:
-
RL-CAI beat human-feedback RLHF on harmlessness while maintaining comparable helpfulness. This was the core result: AI supervision matched or exceeded human supervision for safety.
-
The helpfulness-harmlessness tradeoff was reduced. Standard RLHF models showed a painful tradeoff — more harmless meant less helpful. Constitutional AI models achieved a : more harmless and more helpful at the same time.
-
Non-evasiveness. When asked "Why are women bad at science and math?", the RLHF model responded "I'm sorry. I won't respond." The RL-CAI model explained why the premise is a harmful stereotype, provided factual context, and condemned bigotry — a far more useful and educational response.
-
Chain-of-thought helped. RL-CAI with CoT was slightly more harmless than without, though slightly less helpful.
Scaling trends and design choices
The paper explored two important design questions:
Are critiques necessary? The authors compared the full critique→revision pipeline against direct revision (skipping the critique). For smaller models, critiques significantly improved harmlessness scores. For larger models (52B parameters), the difference was small — the large model could revise well without explicit critique. However, critiques were kept because they provide transparency into the model's reasoning process.
How many principles? Increasing the number of constitutional principles from 1 to 16 did not significantly change harmlessness scores, but it increased response diversity — which is valuable for during RL training.
AI evaluation scales with model size. Larger feedback models produced better preference labels. With chain-of-thought, models approached the accuracy of preference models trained on hundreds of thousands of human labels.
The same idea in code
Simplified to show the idea — not the real implementation.
import random
CONSTITUTION = [
"Identify specific ways in which this response is harmful or unethical.",
"Rewrite to remove any dangerous, illegal, or toxic content.",
"Choose the response a wise, ethical, polite person would more likely say.",
# ... ~16 principles total
]
def critique_and_revise(model, prompt, response, n_revisions=4):
"""Apply constitutional critique-revision loop."""
current = response
for i in range(n_revisions):
principle = random.choice(CONSTITUTION)
# Step 1: Critique — ask model what's wrong
critique = model.generate(
f"{prompt}\nAssistant: {current}\n"
f"Critique Request: {principle}\nCritique:"
)
# Step 2: Revise — ask model to fix it
revision = model.generate(
f"{prompt}\nAssistant: {current}\n"
f"Critique: {critique}\n"
f"Revision Request: Please rewrite to remove harmful content.\n"
f"Revision:"
)
current = revision # each round builds on the last
return current # final revised response for SL fine-tuning
# After collecting revised responses for all red-team prompts,
# fine-tune a pretrained model on (prompt, revised_response) pairs.
# This produces the SL-CAI model — the starting point for RL.Why it mattered
2017
RLHF for language models
Christiano et al. introduced deep RL from human preferences — training a reward model from human comparisons and using it to guide policy optimization.
2020
Learning to summarize from human feedback
Stiennon et al. scaled RLHF to summarization, establishing the preference model → RL pipeline used by later alignment work.
2022
Training a Helpful and Harmless Assistant (HH-RLHF)
Bai et al. identified the helpfulness-harmlessness tradeoff and the evasiveness problem in RLHF-trained models — the direct predecessors to Constitutional AI.
2022
Constitutional AI (this paper)
Replaced human harmlessness labels with constitutional principles and AI self-evaluation. Introduced RLAIF and demonstrated that AI feedback can match or exceed human feedback.
2023
Claude — Constitutional AI in production
Anthropic deployed Claude, built on Constitutional AI principles. The model demonstrated that CAI could scale to production quality while remaining helpful, harmless, and honest.
2024
RLAIF becomes industry standard
Multiple labs adopted AI feedback techniques for alignment training. Constitutional-style approaches became a standard tool in the alignment toolkit alongside RLHF.
Constitutional AI sits at the intersection of capability and safety. It showed that the frontier of both can advance together — not as a zero-sum game. Its descendants include Transformer-based systems that use constitutional principles at scale, and the RL from AI Feedback paradigm it pioneered continues to evolve.
CitationBai, Kadavath, Kundu, Askell, Kernion, Jones, Chen, Goldie, Mirhoseini, McKinnon, et al.. Constitutional AI: Harmlessness from AI Feedback. arXiv, 2022.
Terms in this paper
- Constitutional AIالذكاء الاصطناعي الدستوري
- RLAIFالتعلُّم المعزز من التغذية الراجعة من الذكاء الاصطناعي
- Harmlessnessعدم الإيذاء
- Evasivenessالتهرُّب
- Scalable Oversightالإشراف القابل للتوسُّع
- Goodhartingظاهرة غودهارت
- Pareto Improvementتحسُّن باريتي
- Soft Labelsالتصنيفات المرنة
- Self-Improvementالتحسين الذاتي