HumanEval
HumanEval
معيار يقيس قدرة النموذج على تخليق دوال Python بدرجات تعقيد متفاوتة، وهو من أبرز مقاييس القدرة البرمجية للنماذج اللغوية.
A benchmark measuring the ability to synthesize Python functions of varying complexity, one of the key measures of language model coding capability.
Also translated asمعيار التقييم البشري للبرمجة
First appears in this corpus in: Textbooks Are All You Need (2023)
Appears in these papers