Expert Capacity
سعة الخبير
الحجم الأقصى لدفعة الرموز التي يستطيع كل خبير معالجتها في طبقة خليط الخبراء. تُحسب بقسمة إجمالي الرموز على عدد الخبراء مع ضربها في معامل السعة. الرموز التي تتجاوز هذا الحدّ تُسقَط وتمرّ عبر الاتصال المتبقي.
The maximum batch size of tokens each expert can process in an MoE layer. Calculated as (total tokens / num experts) × capacity factor. Tokens exceeding this limit are dropped and pass through the residual connection.
Also translated asالطاقة الاستيعابية للخبير
First appears in this corpus in: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (2022)
Appears in these papers