Glossary

Red Teaming

الاختبار العدائي

ممارسة يحاول فيها مقيّمون بشريون أو آليون كسر نموذج ذكاء اصطناعي عمداً بمطالبات مصمَّمة لاستدراج مخرجات ضارّة أو متحيّزة أو غير أمينة، بهدف كشف الثغرات قبل النشر واستخدام النتائج في تحسين تدريب النموذج ومحاذاته.

A practice in which human or automated evaluators deliberately try to break an AI model with prompts crafted to elicit harmful, biased, or dishonest outputs, aiming to surface vulnerabilities before deployment and to feed the findings back into model training and alignment.

Also translated asRed Teaming، اختبار الفريق الأحمر، اختبار الاختراق، تقييم أمن النماذج

First appears in this corpus in: Concrete Problems in AI Safety (2016)

Appears in these papers