MMLU
MMLU
معيار تقييم يضم أسئلة متعددة الاختيارات في 57 تخصصاً أكاديمياً ومهنياً، يختبر اتساع معرفة النموذج عبر مجالات متنوعة من القانون إلى الطب إلى الرياضيات.
An evaluation benchmark of multiple-choice questions across 57 academic and professional subjects, testing the breadth of a model's knowledge from law to medicine to mathematics.
Also translated asمعيار الفهم اللغوي متعدد المهام
First appears in this corpus in: Emergent Abilities of Large Language Models (2022)
Appears in these papers