Scaling Law
قانون التحجيم
معادلة تجريبية تتنبأ بكيفية تغيُّر أداء النموذج كدالة لحجمه وبيانات التدريب والميزانية الحوسبية، وتُمكِّن من تحديد التوزيع الأمثل للموارد.
An empirical equation that predicts how model performance changes as a function of model size, training data, and compute budget, enabling optimal resource allocation.
Also translated asقانون القياس، قانون التدريج
First appears in this corpus in: Improving Language Understanding by Generative Pre-Training (2018)
Appears in these papers
- Training Compute-Optimal Large Language Models2022in the sky ✦
- Training Compute-Optimal Large Language Models2022in the sky ✦
- Evaluating Large Language Models Trained on Code2021in the sky ✦
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks2019in the sky ✦
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks2019in the sky ✦
- Are Emergent Abilities of Large Language Models a Mirage?2023in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- Improving Language Understanding by Generative Pre-Training2018in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- Language Models Are Unsupervised Multitask Learners2019in the sky ✦
- Language Models Are Few-Shot Learners2020in the sky ✦
- Language Models Are Few-Shot Learners2020in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- GPT-4 Technical Report2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer2022in the sky ✦
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks2018in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- RETRO: Improving Language Models by Retrieving from Trillions of Tokens2022in the sky ✦
- Scaling Laws for Neural Language Models2020in the sky ✦
- Video Generation Models as World Simulators2024in the sky ✦
- Sparks of Artificial General Intelligence: Early Experiments with GPT-42023in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer2019in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- A Decoder-Only Foundation Model for Time Series Forecasting2024in the sky ✦
- An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale2020in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦