مزيج الخبراء المتناثر
Sparse Mixture of Experts
بنية يُستبدل فيها كل كتلة تغذية أمامية بمجموعة من الشبكات الخبيرة، وموجِّه يختار عدداً صغيراً منها لكل رمز. هذا يفصل السعة الإجمالية للنموذج عن تكلفة الحوسبة لكل رمز.
An architecture where each feedforward block is replaced by a set of expert networks and a router selects a small subset per token. This decouples total model capacity from per-token compute cost.
تُرجم أيضاًمزيج الخبراء المبعثَر، SMoE