Glossary

Mixture of Experts

مزيج الخبراء

بنية نموذج تحتوي على عدد كبير من المعاملات لكن تُفعّل جزءاً صغيراً منها فقط لكل رمز عبر آلية توجيه، مما يُتيح طاقة استيعابية ضخمة بتكلفة حوسبة معقولة.

A model architecture containing a large number of parameters but activating only a small fraction per token via a routing mechanism, enabling massive capacity at reasonable compute cost.

Also translated asMoE، الطبقات متعددة التخصصات الموجهة، شبكات تفعيل الخبراء الحوسبية، نموذج مزيج الخبراء

First appears in this corpus in: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (2022)

Appears in these papers