Glossary

Alignment Tax

ضريبة المحاذاة

التراجع الطفيف في أداء المقاييس المعيارية الناتج عن تدريب RLHF، إذ يُقايض النموذج بعض القدرة العامة مقابل سلوك أفضل.

The small regression in standard NLP benchmark performance caused by RLHF training, as the model trades some general capability for better behavior.

Also translated asتكلفة المحاذاة، كُلفة التوافق

First appears in this corpus in: Training Language Models to Follow Instructions with Human Feedback (2022)

Appears in these papers