المعجم

معيار SWE-bench

SWE-bench

معيار من برينستون قدّمه خيمينيث وزملاؤه، يقيس قدرة النماذج اللغوية على حل مسائل هندسة برمجيات واقعية عبر 2,294 قضية GitHub من 12 مستودع بايثون شهير، إذ يجب على النموذج تعديل قاعدة الكود لتمرير اختبارات المستودع.

A Princeton benchmark introduced by Jimenez and colleagues that measures language models' ability to resolve real-world software engineering problems via 2,294 GitHub issues drawn from 12 popular Python repositories, requiring the model to edit the codebase so the repository's own tests pass.

تُرجم أيضاًSWE-bench، معيار سوي بينش، معيار برينستون لهندسة البرمجيات