Red Teaming
الاختبار العدائي
ممارسة يحاول فيها مقيّمون بشريون أو آليون كسر نموذج ذكاء اصطناعي عمداً بمطالبات مصمَّمة لاستدراج مخرجات ضارّة أو متحيّزة أو غير أمينة، بهدف كشف الثغرات قبل النشر واستخدام النتائج في تحسين تدريب النموذج ومحاذاته.
A practice in which human or automated evaluators deliberately try to break an AI model with prompts crafted to elicit harmful, biased, or dishonest outputs, aiming to surface vulnerabilities before deployment and to feed the findings back into model training and alignment.
Also translated asRed Teaming، اختبار الفريق الأحمر، اختبار الاختراق، تقييم أمن النماذج
First appears in this corpus in: Concrete Problems in AI Safety (2016)
Appears in these papers
- Concrete Problems in AI Safety2016in the sky ✦
- Constitutional AI: Harmlessness from AI Feedback2022in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2022in the sky ✦
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2022in the sky ✦
- Sparks of Artificial General Intelligence: Early Experiments with GPT-42023in the sky ✦