Mixture of Experts
مزيج الخبراء
بنية نموذج تحتوي على عدد كبير من المعاملات لكن تُفعّل جزءاً صغيراً منها فقط لكل رمز عبر آلية توجيه، مما يُتيح طاقة استيعابية ضخمة بتكلفة حوسبة معقولة.
A model architecture containing a large number of parameters but activating only a small fraction per token via a routing mechanism, enabling massive capacity at reasonable compute cost.
Also translated asMoE، الطبقات متعددة التخصصات الموجهة، شبكات تفعيل الخبراء الحوسبية، نموذج مزيج الخبراء
First appears in this corpus in: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (2022)
Appears in these papers
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning2025in the sky ✦
- DeepSeek-V3 Technical Report2024in the sky ✦
- Gemini: A Family of Highly Capable Multimodal Models2023in the sky ✦
- The Llama 3 Herd of Models2024in the sky ✦
- Mistral 7B2023in the sky ✦
- Mixtral of Experts2024in the sky ✦
- Mixtral of Experts2024in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- π₀: A Vision-Language-Action Flow Model for General Robot Control2024in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity2022in the sky ✦