Model Parallelism
توازي النموذج
استراتيجية تدريب تُقسِّم طبقات النموذج أو معاملاته على عدة وحدات معالجة رسومية لأن النموذج أكبر من أن يسع ذاكرة وحدة واحدة. النموذج ذو الـ65 مليار في LLaMA تدرّب على 2048 وحدة A100.
A training strategy that splits model layers or parameters across multiple GPUs because the model is too large for a single device's memory. LLaMA-65B trained on 2048 A100 GPUs.
Also translated asالتوازي في النموذج، التوزيع الهيكلي للأوزان حوسبياً، تقسيم النموذج، تقسيم طبقات الشبكة عبر معالجات متعددة
First appears in this corpus in: Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation (2016)
Appears in these papers
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Gemma: Open Models Based on Gemini Research and Technology2024in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation2016in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher2022in the sky ✦
- LLaMA: Open and Efficient Foundation Language Models2023in the sky ✦
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism2019in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- PaLM: Scaling Language Modeling with Pathways2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models2019in the sky ✦