Warmup
الإحماء
مرحلة في بداية التدريب يرتفع فيها معدل التعلم تدريجياً من قيمة قريبة من الصفر إلى قيمته المستهدفة. تمنع الانحراف المبكر الناتج عن خطوات كبيرة حين لا تزال أوزان النموذج عشوائية.
A phase at the start of training where the learning rate gradually increases from near-zero to its target value. Prevents early divergence caused by large steps when model weights are still near random initialization.
Also translated asمرحلة الإحماء، التسخين التدريجي
First appears in this corpus in: Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour (2017)
Appears in these papers
- The Falcon Series of Open Language Models2023in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes2019in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks2019in the sky ✦