Benchmark
المعيار المرجعي
مجموعة مهام أو أسئلة موحَّدة تُستخدم لقياس أداء النماذج ومقارنتها بشكل عادل ومتسق عبر الزمن.
A standardized set of tasks or questions used to measure and compare model performance fairly and consistently over time.
Also translated asمعيار الأداء، مقياس التقييم
First appears in this corpus in: ImageNet: A Large-Scale Hierarchical Image Database (2009)
Appears in these papers
- Parameter-Efficient Transfer Learning for NLP2019in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- Microsoft COCO: Common Objects in Context2014in the sky ✦
- Evaluating Large Language Models Trained on Code2021in the sky ✦
- Evaluating Large Language Models Trained on Code2021in the sky ✦
- Contriever: Unsupervised Dense Information Retrieval with Contrastive Learning2022in the sky ✦
- Conservative Q-Learning for Offline Reinforcement Learning2020in the sky ✦
- DeBERTa: Decoding-Enhanced BERT with Disentangled Attention2020in the sky ✦
- DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter2019in the sky ✦
- Are Transformers Effective for Time Series Forecasting?2023in the sky ✦
- Emergent Abilities of Large Language Models2022in the sky ✦
- Emergent Abilities of Large Language Models2022in the sky ✦
- Are Emergent Abilities of Large Language Models a Mirage?2023in the sky ✦
- Are Emergent Abilities of Large Language Models a Mirage?2023in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- Flamingo: a Visual Language Model for Few-Shot Learning2022in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding2018in the sky ✦
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding2018in the sky ✦
- First Return, Then Explore2021in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- HuggingGPT: Solving AI Tasks with ChatGPT and Its Friends in Hugging Face2023in the sky ✦
- ImageNet: A Large-Scale Hierarchical Image Database2009in the sky ✦
- LIMA: Less Is More for Alignment2023in the sky ✦
- Llama 2: Open Foundation and Fine-Tuned Chat Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale2022in the sky ✦
- Mistral 7B2023in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2024in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- Textbooks Are All You Need2023in the sky ✦
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2022in the sky ✦
- Reflexion: Language Agents with Verbal Reinforcement Learning2023in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Self-Instruct: Aligning Language Models with Self-Generated Instructions2022in the sky ✦
- Sparks of Artificial General Intelligence: Early Experiments with GPT-42023in the sky ✦
- Sparks of Artificial General Intelligence: Early Experiments with GPT-42023in the sky ✦
- SQuAD: 100,000+ Questions for Machine Comprehension of Text2016in the sky ✦
- SQuAD: 100,000+ Questions for Machine Comprehension of Text2016in the sky ✦
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2024in the sky ✦
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2024in the sky ✦
- Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows2021in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- Toolformer: Language Models Can Teach Themselves to Use Tools2023in the sky ✦
- Toolformer: Language Models Can Teach Themselves to Use Tools2023in the sky ✦
- Universal Language Model Fine-Tuning for Text Classification2018in the sky ✦
- VQA: Visual Question Answering2015in the sky ✦
- WebArena: A Realistic Web Environment for Building Autonomous Agents2023in the sky ✦
- WebArena: A Realistic Web Environment for Building Autonomous Agents2023in the sky ✦
- XLNet: Generalized Autoregressive Pretraining for Language Understanding2019in the sky ✦