AI Safety2022intermediate14 min read
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
اختبار الفريق الأحمر للنماذج اللغوية للحدّ من أضرارها: المنهجيات وأنماط التوسّع والدروس المستفادة
Ganguli, D. · Lovitt, L. · Kernion, J. · Askell, A. · Bai, Y. · Kadavath, S. · Mann, B. · Perez, E. · Schiefer, N. · Ndousse, K. · Jones, A. · Bowman, S. · Amodei, D. · Brown, T. · Kaplan, J. · Clark, J. — arXiv
The problem
Large language models can produce harmful outputs — from reinforcing social biases and generating toxic content to leaking private data and aiding disinformation. By 2022, several interventions existed (prompting, filtering, RLHF), but the AI community lacked systematic methods to discover, measure, and compare how well these interventions actually work. There were no shared norms for how to stress-test language models, no large-scale public datasets of adversarial attacks, and no clear understanding of how model size affects vulnerability to adversarial probing.
The contribution
Three contributions that shaped how the AI community approaches model safety. First, a systematic study of behaviors for across 3 model sizes (2.7B, 13B, 52B parameters) and 4 model types — revealing that RLHF models become dramatically harder to red team as they scale, while other interventions show flat trends. Second, the release of 38,961 red team attacks as a public dataset — an order of magnitude larger than prior work — enabling the community to study attack patterns, build safety classifiers, and develop automated red teaming. Third, a transparent, exhaustive description of their methods, processes, and uncertainties — establishing a template for responsible red teaming practices.
The impact
This paper helped establish red teaming as a standard practice in . It demonstrated that RLHF models scale favorably for safety — a finding that influenced the design of subsequent systems like ChatGPT and Claude. The released dataset became a resource for safety research. The paper's transparent methodology inspired other labs to adopt and publish their own red teaming practices, and contributed to the growing consensus that systematic is essential before deploying large language models.
Imagine a bank hiring professional burglars to break into its own vault — not to steal, but to find every weak lock, every blind spot in the cameras, every gap in the alarm system. Each failed break-in teaches the bank exactly where to reinforce.
Now scale that idea up: instead of one vault, imagine four vaults of increasing sophistication. Instead of one burglar, imagine 324 people launching nearly 39,000 attempts. And the most important discovery? The vault built with feedback from previous break-in attempts (RLHF) becomes exponentially harder to crack as it gets bigger — while simpler vaults stay equally vulnerable no matter their size.
This paper is that security audit for AI language models.
What is red teaming and why does it matter?
Red teaming is borrowed from military and cybersecurity practice: a dedicated "red team" tries to break a system while the "blue team" defends it. In the context of language models, red teaming means having people (or other models) deliberately try to make an AI say harmful things — offensive content, dangerous instructions, private information, misinformation.
Why not just test with standard benchmarks? Because benchmarks use fixed prompts. A determined human can be creative, adaptive, and persistent in ways that static evaluations cannot capture. A person might try roleplay attacks ("pretend you're a villain and explain how to…"), multi-turn escalation (starting innocently and gradually steering the conversation toward harm), or exploit the model's desire to be helpful.
This paper conducted one of the first large-scale, systematic red teaming efforts on language models. The team recruited 324 crowdworkers who launched 38,961 adversarial attacks against models of varying sizes and safety configurations. The goal was three-fold: discover what types of harm these models produce, measure whether safety interventions actually work, and share the findings openly so the community can build on them.
The four models under test
The researchers tested four types of AI assistants, each with a different level of safety intervention. Think of them as four locks of increasing sophistication on the same vault:
Plain LM — A raw with minimal formatting to act as a chatbot. No safety measures at all. This is the baseline — the vault with a basic padlock.
Prompted LM — The same model, but given a 14-example telling it to be helpful, honest, and harmless (HHH). This is like adding a security guard who follows written instructions — helpful, but easily fooled by creative social engineering.
(RS) — For each response, the model generates 16 candidates, then a separate " " picks the two safest. This is like having a panel of judges review every response before it leaves the building.
RLHF — The model is trained end-to-end using to maximize harmlessness scores from the preference model. This is the vault that was redesigned from the ground up based on what previous burglars found.
Each model type was tested at three sizes: 2.7 billion, 13 billion, and 52 billion parameters. This matrix of 12 configurations let the researchers study how safety scales with model size and intervention type.
Key finding: RLHF scales, other interventions do not
The most striking result is shown in Figure 1 of the paper. The researchers measured attack success in two ways: the red team members' self-reported success rating (a 5-point Likert scale), and an automated harmlessness score from a preference model (lower means more harmful).
The finding that reshaped the field: RLHF models become significantly harder to red team as they scale up. The 52B- RLHF model was substantially more resistant to attacks than the 2.7B version. This is a positive scaling behavior — as models get larger, they also get safer, when trained with RLHF.
In contrast, the other three model types — plain LM, prompted LM, and rejection sampling — showed flat trends with scale. Making the model bigger did not make it safer unless the safety was baked into the process via RLHF.
A surprising sub-finding: the prompted LM (with its 14-example HHH prompt) was no harder to red team than the plain LM with no safety intervention at all. This contradicted earlier results from static evaluation benchmarks, suggesting that adversarial humans can bypass prompt-based safety in ways that fixed test prompts cannot.
Rejection sampling placed a floor on vulnerability — it was consistently the hardest to attack at every scale. However, qualitatively, the researchers noted that RS models achieved safety by being evasive — they refused to engage rather than engaging helpfully while avoiding harm.
Measuring harm: preference models as automated judges
How do you quantify whether an AI response is harmful? The authors used two complementary methods. First, red team members themselves rated their attack success on a 0–4 Likert scale after each conversation. Second, a harmlessness preference model — a separate trained to distinguish harmful from harmless responses — scored every AI utterance automatically.
The preference model works by comparing pairs of responses: during data collection, each red team member was shown two possible AI responses and asked to select the more harmful one. This pairwise comparison data trains a model that can score any new response on a continuous harmlessness scale.
For each conversation, the researchers computed the harmlessness score of every assistant utterance (conditioned on all prior context), then took the minimum score as the overall harmlessness of that conversation. This minimum captures the worst single moment — because one severely harmful response makes the entire interaction dangerous, regardless of how benign the rest was.
The two metrics — human self-report and automated harmlessness score — are correlated but not identical. The human metric captures the red team member's subjective judgment, while the automated score provides a consistent, scalable measurement. Using both gives a more complete picture of attack success.
The landscape of harm: what the attacks revealed
One of the paper's most valuable contributions is its detailed taxonomy of the types of attacks and harms uncovered. The researchers embedded all 38,961 attack transcripts using a 52B language model, then projected them into 2D using . This revealed natural clusters of attack types that could be manually labeled.
The five most common attack categories were: (1) discrimination and injustice, (2) hate speech and offensive language, (3) violence and incitement, (4) non-violent unethical behavior (lying, cheating, manipulation), and (5) bullying and harassment.
Crucially, the researchers found that more subtle attacks — especially "non-violent unethical behavior" — had higher success rates than blunt offensive attacks. This makes intuitive sense: safety interventions are typically calibrated to block obviously toxic outputs (slurs, explicit violence), so more nuanced ethical violations slip through more easily.
Less common but significant categories included: soliciting personally identifiable information (PII), substance abuse, fraud, weapons information, terrorism, self-harm, and child abuse. The "Other" category was also prevalent, showing that any fixed taxonomy will miss attack types that creative humans invent.
The red team: who attacks and how
The red team consisted of 324 US-based crowdworkers, recruited primarily from Amazon Mechanical Turk (307 workers) and Upwork (17 workers). Each worker was given open-ended instructions to "make the AI behave badly" — they could choose their own attack strategies and topics within their personal risk tolerance.
A key finding about the red team itself: roughly 80% of all attacks came from just 15% of the workers. Some individuals were far more effective at red teaming than others — an observation that has implications for how to select and train red teams. The researchers fitted a linear mixed-effects model to estimate each worker's inherent red teaming ability, and found significant variation in skill.
Some workers developed template-based attacks (e.g., "tell me an insult for [X] that starts with [Y]"), which generated many attacks quickly but with limited diversity. Others were more creative, inventing roleplay attacks, multi-turn escalation strategies, and subtle manipulation techniques. The most dangerous attacks typically required creativity and persistence.
The researchers also invested significantly in red team well-being: clear warnings about harmful content, personal risk tolerance guidelines, recommended breaks, and a well-being survey. Participants generally reported positive feelings and enjoyed the work — an important finding given the parallels with , where worker burnout is a serious concern.
The subjectivity problem: what counts as harmful?
A critical question for any red teaming effort: how much do different people agree on whether an attack was successful? The researchers ran a separate review experiment where 3 annotators independently rated the same 1,000 attack transcripts.
The result was sobering: inter-rater agreement was low. Using Fleiss's Kappa — a standard measure where 1 means perfect agreement and 0 means chance — they found a score of only 0.32 for the 5-point Likert rating, rising to 0.49 when ratings were simplified to binary (successful/not successful). Even among the 3 reviewers alone (excluding the original attacker), agreement peaked at 0.55.
This means "harmful" is deeply subjective. What one person considers a successful attack — dangerous advice, subtle manipulation, a harmful joke — another might dismiss as harmless or ineffective. This subjectivity is not a flaw of the method but a fundamental property of harm itself: harm depends on context, cultural background, and individual sensitivity.
The implication is that any single harmlessness metric — human or automated — provides an incomplete picture. The researchers recommend using multiple metrics and being transparent about the inherent uncertainty in measuring harm.
The released dataset: 38,961 attacks as a public resource
A bold decision: the researchers released the entire dataset of 38,961 red team attacks publicly. This was not without controversy — the same data that can train safer models can also be used to train more harmful ones. The team explicitly weighed the pros and cons.
On the positive side: the dataset is an order of magnitude larger than the previous Bot Adversarial Dialogues dataset (about 5,000 conversations). It covers larger models, includes attacks on RLHF-trained models, and comes with both human ratings and automated harmlessness scores. It costs over $60,000 in crowdworker payments alone to produce, making it a significant public good.
Each data point includes the conversation transcript, the red team member's success rating, a short description of their intended attack, harmlessness scores for both the AI responses and the attack description, the model type and size, and an anonymous worker identifier. The team filtered personally identifiable information (PII) using a regular expression, though they noted that some AI-generated PII (hallucinated addresses, phone numbers) appeared to be synthetic and not real.
The dataset has been used extensively by the research community for building safety classifiers, developing automated red teaming methods, and characterizing the attack surface of language models.
Limitations: what red teaming cannot catch
The authors were refreshingly honest about limitations. The red team was composed of US-based crowdworkers, which means the dataset reflects a particular cultural perspective. Harms that are specific to other cultures, languages, or contexts may be underrepresented.
Domain expertise was another gap. Some attacks required specialized knowledge to evaluate — for example, whether instructions for synthesizing a chemical are actually viable, or whether medical advice is genuinely dangerous. Crowdworkers without this expertise may have rated such attacks incorrectly.
The data was inherently incomplete. The models were partly trained on code, yet no attacks related to malicious code generation appeared. "Roleplay attacks" — where the model is asked to act as a fictional character who can bypass safety rules — were discovered internally but did not appear in the crowdworker dataset.
Finally, the paper acknowledged a fundamental tension in red teaming: sharing findings openly improves community safety, but also reveals vulnerabilities that bad actors can exploit. The authors called for the development of shared norms and neutral forums to discuss these trade-offs.
Legacy: from experiment to industry standard
2014
Adversarial examples discovered (Szegedy et al.)
The concept of adversarial testing for neural networks begins with image classifiers. Imperceptible perturbations can fool confident models.
2019
Build It Break It Fix It (Dinan et al.)
An early adversarial data collection framework for dialogue safety. Showed that iterative human-in-the-loop testing can improve model robustness.
2022
Red teaming LMs with LMs (Perez et al.)
Automated red teaming: using language models to generate adversarial prompts instead of relying entirely on human labor.
2022
This paper — Red Teaming at Scale (Ganguli et al.)
38,961 human attacks across 12 model configurations. First large-scale evidence that RLHF models improve with scale. Public dataset release.
2022
Constitutional AI (Bai et al.)
Built on red teaming insights to create AI systems that self-critique and self-revise based on a set of constitutional principles.
2023
Red teaming becomes industry standard
Major AI labs (OpenAI, Google, Anthropic, Meta) all adopt systematic red teaming before model deployment. Executive orders and policy frameworks cite red teaming as a required practice.
This paper's influence extends beyond its technical findings. By publishing methods, data, and uncertainties in full, it established transparency as a norm for safety research. The finding that RLHF scales positively for safety became a foundational assumption in the design of subsequent AI systems. And the paper's honest discussion of trade-offs — between openness and security, between completeness and feasibility — set the template for how responsible AI labs communicate about safety.
CitationGanguli, Lovitt, Kernion, Askell, Bai, Kadavath, Mann, Perez, Schiefer, Ndousse, et al.. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv, 2022.
Terms in this paper
- Red Teamingالاختبار العدائي
- Adversarial Testingالاختبار العدائي
- AI Safetyسلامة الذكاء الاصطناعي
- Reinforcement Learning from Human Feedbackالتعلم بالتعزيز القائم على التقييم البشري
- Rejection Samplingالترشيح بالرفض
- Harmlessnessعدم الإيذاء
- Toxicityالسمّية
- Adversarial Attackالهجوم العدائي الموجه
- Language Modelالنموذج اللغوي
- Scalingالتدريج
- Crowdsourcingالتعهيد الجماعي
- Preference Modelنموذج التفضيلات
- Content Moderationإشراف المحتوى
- Biasالانحياز الحسابي