BIG-Bench
BIG-Bench
مجموعة معايير واسعة النطاق تضم أكثر من 200 مهمة جُمعت بشكل تعاوني لتقييم النماذج اللغوية الكبيرة، تشمل مهام حسابية ولغوية ومنطقية متنوعة.
A large-scale crowd-sourced benchmark suite of over 200 tasks for evaluating large language models, including arithmetic, linguistic, and logical reasoning tasks.
Also translated asمعيار BIG-Bench، مجموعة BIG-Bench
First appears in this corpus in: Emergent Abilities of Large Language Models (2022)
Appears in these papers