Mixed Precision Training
التدريب بالدقة المختلطة
تقنية تدريب تُجرى فيها معظم العمليات بأنواع فاصلة عائمة قصيرة (FP16 أو BF16) مع الاحتفاظ بنسخة رئيسية من الأوزان وتدرّج الخسارة بدقة FP32، وتُستخدم قياس الخسارة لتفادي الفيض، فتُقلِّص الذاكرة والوقت مع الحفاظ على الدقة النهائية.
A training technique in which most operations run in short floating-point types (FP16 or BF16) while a master copy of the weights and the loss gradient are kept in FP32, with loss scaling to avoid underflow, cutting memory and time while preserving final accuracy.
Also translated asMixed Precision، Mixed Precision Training، التدريب بدقّات مختلطة، تدريب FP16/FP32
First appears in this corpus in: Mixed Precision Training (2018)
Appears in these papers
- Alpaca: A Strong, Replicable Instruction-Following Model2023in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLM.int8(): 8-Bit Matrix Multiplication for Transformers at Scale2022in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Mixed Precision Training2018in the sky ✦
- RoBERTa: A Robustly Optimized BERT Pretraining Approach2019in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦