Data Parallelism
توازي البيانات
استراتيجية تدريب تُوزّع دُفعات البيانات على عدة وحدات معالجة رسومية، كل منها تحتفظ بنسخة من النموذج وتُجمع التدرجات بعد كل خطوة. تُستخدم مع توازي النموذج في تدريب LLaMA.
A training strategy that distributes data batches across multiple GPUs, each holding a copy of the model and aggregating gradients after each step. Used alongside model parallelism in LLaMA training.
Also translated asالتوازي في البيانات، توزيع مصفوفات البيانات عبر المعالجات، معالجة البيانات بالتوازي، معالجة حزم المدخلات المختلفة بالتزامن
First appears in this corpus in: Deep Speech 2: End-to-End Speech Recognition in English and Mandarin (2015)
Appears in these papers
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model2022in the sky ✦
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin2015in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- The Falcon Series of Open Language Models2023in the sky ✦
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning2023in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures2018in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour2017in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- OPT: Open Pre-Trained Transformer Language Models2022in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦