Glossary

Sparse Mixture of Experts

مزيج الخبراء المتناثر

بنية يُستبدل فيها كل كتلة تغذية أمامية بمجموعة من الشبكات الخبيرة، وموجِّه يختار عدداً صغيراً منها لكل رمز. هذا يفصل السعة الإجمالية للنموذج عن تكلفة الحوسبة لكل رمز.

An architecture where each feedforward block is replaced by a set of expert networks and a router selects a small subset per token. This decouples total model capacity from per-token compute cost.

Also translated asمزيج الخبراء المبعثَر، SMoE