Glossary

Model Parallelism

توازي النموذج

استراتيجية تدريب تُقسِّم طبقات النموذج أو معاملاته على عدة وحدات معالجة رسومية لأن النموذج أكبر من أن يسع ذاكرة وحدة واحدة. النموذج ذو الـ65 مليار في LLaMA تدرّب على 2048 وحدة A100.

A training strategy that splits model layers or parameters across multiple GPUs because the model is too large for a single device's memory. LLaMA-65B trained on 2048 A100 GPUs.

Also translated asالتوازي في النموذج، التوزيع الهيكلي للأوزان حوسبياً، تقسيم النموذج، تقسيم طبقات الشبكة عبر معالجات متعددة

First appears in this corpus in: Google's Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation (2016)

Appears in these papers