Glossary

Evasiveness

التهرُّب

ميل النماذج المُدرَّبة على السلامة إلى رفض الخوض في المواضيع الحسّاسة كلياً بعبارات جاهزة مثل «لا أستطيع الإجابة»، بدلاً من الشرح المُتعقِّل لأسباب الرفض.

The tendency of safety-trained models to refuse to engage with sensitive topics entirely using canned phrases like 'I can't answer that', instead of thoughtfully explaining why.

Also translated asالتملُّص، الاستجابة المُراوِغة

First appears in this corpus in: Constitutional AI: Harmlessness from AI Feedback (2022)

Appears in these papers