Language Models2024intermediate11 min read

Mixtral of Experts

Mixtral: مزيج الخبراء

Jiang, A. Q. · Sablayrolles, A. · Roux, A. · Mensch, A. · Savary, B. · Bamford, C. · Chaplot, D. S. · de las Casas, D. · Bou Hanna, E. · Bressand, F. · Lengyel, G. · Bour, G. · Lample, G. · Renard Lavaud, L. · Saulnier, L. · Lachaux, M.-A. · Stock, P. · Subramanian, S. · Yang, S. · Antoniak, S. · Le Scao, T. · Gervet, T. · Lavril, T. · Wang, T. · Lacroix, T. · El Sayed, W. — arXiv

The problem

By late 2023, scaling language models meant scaling all parameters for every — a brute-force approach where cost grows linearly with model size. Llama 2 70B delivers strong performance, but every single token activates all 70 billion parameters, making expensive and slow. The fundamental question was: can we build a model that accesses a vast pool of knowledge while only paying the compute cost of a much smaller model for each token?

The contribution

Mixtral 8x7B: a Sparse that replaces each feedforward block in a Mistral 7B-style with 8 expert networks and a learned that selects 2 experts per token. The result is a model with 47B total parameters but only 13B active per token. With a 32k and open weights under Apache 2.0, Mixtral matches or outperforms Llama 2 70B and GPT-3.5 on most benchmarks, and vastly outperforms them on mathematics, code generation, and multilingual tasks. The instruction-tuned variant, Mixtral Instruct, surpasses GPT-3.5 Turbo, Claude-2.1, and Gemini Pro on human evaluation benchmarks.

The impact

Mixtral proved that sparse Mixture of Experts is a practical, production-ready architecture for open-weight LLMs. It demonstrated that routing tokens to a subset of experts can match dense models five times larger in active parameters, reshaping the cost-performance frontier. It directly influenced DeepSeek-V3 and the broader adoption of MoE in frontier models, and established Mistral AI as a major open-source competitor. The Apache 2.0 license made MoE accessible to the entire community for the first time at this scale.

Imagine a law firm with 8 specialist lawyers — one expert in tax, another in patents, a third in contracts, and so on. When a new case arrives, a senior partner glances at it and routes it to the 2 most relevant specialists. The firm has the collective expertise of 8 lawyers, but the billable hours for each case are only those of 2.

Now imagine every case is different, so the partner might pick a completely different pair each time. The tax expert might handle one case, sit idle for the next three, then be called again — the assignment is dynamic, not fixed.

Mixtral works exactly like this firm. Each layer has 8 feedforward "expert" networks, and a lightweight router picks 2 for each token. The model accesses 47 billion parameters of knowledge, but any single token only activates 13 billion — the cost of a small model, with the capacity of a large one.

The architecture: replacing dense layers with a panel of experts

Mixtral starts from the same Transformer used in Mistral 7B — with , , and positional encodings — but makes one decisive change: every feedforward () sub-block is replaced by a Mixture-of-Experts (MoE) layer.

In a standard Transformer, the FFN block is a single large network that every token must pass through. Think of it as a single highway lane: no matter the vehicle, everyone merges into the same lane. In Mixtral, that single lane is replaced by 8 parallel lanes (experts), and a traffic controller (the router) directs each vehicle (token) to the 2 lanes that best match its needs.

Each expert is a SwiGLU feedforward network — the same architecture used inside Mistral 7B — with its own independent set of weights. The key insight is that although the model has 8 experts per layer (totaling ~47B parameters), any individual token only activates 2 experts (~13B active parameters). This is why the approach is called sparse: most of the model is dormant for any given token.

Open in Lab
Click on different tokens to see which 2 of the 8 experts the router selects, and how their outputs are combined with learned weights.
The demo wakes as you arrive…

The router: deciding which experts to consult

The router (also called the gating network) is the brain of the MoE layer. For each token, it must answer one question: "Which 2 of these 8 experts should handle this token?"

Mechanically, the router is a simple linear layer followed by a Top-K selection and . The input token's xx is multiplied by a weight matrix WgW_g to produce 8 scores (one per expert). The Top-2 scores are kept; all others are set to −∞-\infty. A softmax over the surviving scores produces the gating weights — these weights determine both which experts are activated and how much each expert's output contributes to the final result.

The beauty of this design is its simplicity. The router has no memory, no recurrence, no special tricks — just a single matrix multiply and a selection step. Yet through , it learns to route tokens to appropriate experts based on their content.

G(x)=Softmax ⁣(Top2(x⋅Wg))G(x) = \text{Softmax}\!\bigl(\text{Top2}(x \cdot W_g)\bigr)
Router gating function — selecting the top-2 experts — The hidden state xx is projected by WgW_g into n=8n = 8 logits. Top2 keeps the two largest and sets the rest to −∞-\infty. Softmax normalizes the two surviving logits into weights that sum to 1. These weights control how much each selected expert contributes to the output.
y=∑i=0n−1G(x)i⋅SwiGLUi(x)y = \sum_{i=0}^{n-1} G(x)_i \cdot \text{SwiGLU}_i(x)
MoE layer output — weighted sum of selected experts — The final output yy for a token is the weighted sum of the expert outputs. Because G(x)G(x) is sparse (only 2 non-zero entries), we only compute the SwiGLU forward pass for 2 of the 8 experts — the other 6 contribute nothing and cost nothing.
Open in Lab
Adjust the input to see how the router assigns gating weights across the 8 experts. Notice that only the top-2 are activated.
The demo wakes as you arrive…

Sparse vs. dense: doing more with less

The central claim of Mixtral is that a sparse model with 47B total parameters and only 13B active parameters per token can match or beat a dense model with 70B parameters that are all active for every token.

This distinction between sparse count (total) and (per token) is fundamental. Think of it as the difference between a library's total collection and the books you actually take off the shelf for a given research question. The library is vast, but your reading cost is small.

In practice, Mixtral's inference compute is proportional to its active parameters (13B), not its total parameters (47B). This means it processes tokens at roughly the speed of a 13B dense model, while drawing on the knowledge capacity of a much larger one. The trade-off is memory: the full 47B parameters must be stored, though this is still smaller than Llama 2 70B.

Open in Lab
Compare the active parameter count, total parameters, and relative speed of Mixtral against dense models like Llama 2 13B, 70B, and Mistral 7B.
The demo wakes as you arrive…

Results: punching above its weight class

Mixtral was evaluated across a comprehensive set of benchmarks covering commonsense reasoning, world knowledge, reading comprehension, mathematics, code generation, and multilingual understanding.

The headline result is striking: with only 13B active parameters, Mixtral matches or outperforms Llama 2 70B — a model that uses 5× more active parameters — on almost every . On mathematics (GSM-8K: 74.4% vs 69.6%) and code generation (HumanEval: 40.2% vs 29.3%; MBPP: 60.7% vs 49.8%), the gap is enormous.

Compared to GPT-3.5, Mixtral performs similarly on MMLU (70.6% vs 70.0%) and outperforms it on code (MBPP: 60.7% vs 52.2%). The instruction-tuned variant, Mixtral Instruct, achieves an MT-Bench score of 8.30 — surpassing GPT-3.5 Turbo (8.32 is comparable), Claude-2.1, and Gemini Pro on human evaluation via the LMSys Chatbot Arena.

Open in Lab
Radar chart comparing Mixtral 8x7B against Llama 2 70B and GPT-3.5 across key benchmarks. Hover over each axis for exact scores.
The demo wakes as you arrive…

Multilingual strength: beyond English

Mixtral was pretrained with a significantly increased proportion of multilingual data compared to Mistral 7B. The extra capacity provided by the 8 experts allows the model to absorb multilingual knowledge without sacrificing English performance.

The results on French, German, Spanish, and Italian benchmarks are particularly impressive. On MMLU in French, Mixtral scores 70.9% compared to Llama 2 70B's 64.3%. On German MMLU, it achieves 71.5% versus 64.2%. The pattern repeats across all four languages and all three benchmark types (ARC-Challenge, HellaSwag, MMLU), with Mixtral consistently outperforming a model that uses 5× its active compute.

This suggests that the additional expert capacity provides "room" for the model to store language-specific knowledge in specialized experts, while shared linguistic structure is captured by experts activated across languages.

Routing analysis: what do experts specialize in?

A natural question is whether experts specialize by domain — does one expert handle mathematics while another handles biology? The authors investigated this by measuring expert selection distributions across different subsets of The Pile validation dataset: ArXiv, DM Mathematics, GitHub, Gutenberg, PhilPapers, PubMed Abstracts, StackExchange, and Wikipedia.

The surprising answer is no — there is no clear domain specialization. The distribution of expert assignments looks remarkably similar across topics like ArXiv (LaTeX papers), PubMed (biology), and PhilPapers (philosophy). Only DM Mathematics shows a slightly different pattern, likely because it is a synthetic dataset with limited natural language diversity.

Instead of domain specialization, the router exhibits syntactic behavior. When visualizing token-level routing on code, mathematics, and English text, the patterns follow syntax rather than semantics: Python's self keyword consistently routes to the same expert, indentation tokens cluster together, and consecutive tokens within the same syntactic unit often share experts. This "positional locality" is especially pronounced in the middle and final layers of the model.

Open in Lab
See how the router assigns experts across different domains. Toggle between ArXiv, GitHub, and English text to observe syntactic — not semantic — patterns.
The demo wakes as you arrive…

Mixtral Instruct: from base model to assistant

The base Mixtral model is a next-token predictor. To turn it into a helpful assistant, the authors applied a two-stage fine-tuning recipe: (SFT) on an instruction dataset, followed by (DPO) on a paired feedback dataset.

DPO is a simpler alternative to : instead of training a separate and running reinforcement learning, DPO directly optimizes the language model on pairs of (preferred response, rejected response). The result is a model aligned with human preferences without the complexity of PPO.

Mixtral Instruct achieved an MT-Bench score of 8.30 and reached an Elo rating of 1121 on the LMSys Chatbot Arena (as of December 2023) — placing it above Claude-2.1, every version of GPT-3.5 Turbo, Gemini Pro, and Llama 2 70B Chat. It was the highest-ranked open-weight model at the time.

The authors also measured biases using BBQ and BOLD benchmarks. Mixtral shows higher accuracy on BBQ (56.0% vs 51.5% for Llama 2 70B), indicating fewer social biases, and displays more positive sentiment with similar variance on BOLD.

Long context: 32k tokens with full retrieval

Mixtral was trained with a context window of 32,768 tokens and supports fully dense across the entire window. To test whether the model can actually use all of this context, the authors ran the passkey retrieval task: a random passkey is inserted at a random position in a long prompt, and the model must retrieve it.

Mixtral achieves 100% retrieval accuracy regardless of where the passkey is placed or how long the prompt is. This confirms that the model does not suffer from the "lost in the middle" problem that affects some long-context models.

Additionally, on the proof-pile dataset decreases monotonically as context length increases — evidence that the model genuinely leverages longer context rather than ignoring distant tokens.

Architecture at a glance

Mixtral 8x7B — Architecture hyperparameterspython

Simplified to show the idea — not the real implementation.

# Mixtral 8x7B architecture configuration
config = {
    "dim":           4096,     # Hidden dimension
    "n_layers":      32,       # Number of Transformer layers
    "head_dim":      128,      # Per-head dimension
    "hidden_dim":    14336,    # FFN intermediate size (per expert)
    "n_heads":       32,       # Number of attention heads
    "n_kv_heads":    8,        # KV heads (Grouped Query Attention)
    "context_len":   32768,    # 32k context window
    "vocab_size":    32000,    # Vocabulary size
    "num_experts":   8,        # Experts per MoE layer
    "top_k_experts": 2,        # Experts activated per token
}
# Total params:  ~47B (sparse)
# Active params: ~13B (per token)

The MoE lineage: from idea to Mixtral

  1. 1991

    Mixture of Experts (Jacobs et al.)

    The original MoE concept: divide-and-conquer with a gating network that learns to partition the input space among specialized expert modules.

  2. 2017

    Sparsely-Gated MoE (Shazeer et al.)

    Scaled MoE to billions of parameters with a sparsely-gated layer that routes each token to a small number of experts. Proved that sparse routing can achieve enormous models at manageable compute.

  3. 2022

    Switch Transformer (Fedus et al.)

    Simplified MoE routing to a single expert per token (K=1), scaling to trillions of parameters. Demonstrated that simpler routing can work at extreme scale.

  4. 2023

    Mistral 7B

    Dense 7B model with grouped query attention and sliding window attention. Outperformed Llama 2 13B. Became the architectural backbone for Mixtral.

  5. 2024

    Mixtral 8x7B

    Replaced every FFN in Mistral 7B with 8 SwiGLU experts and a top-2 router. 47B total, 13B active. Matched Llama 2 70B with 5× fewer active parameters. Open weights under Apache 2.0.

Mixtral's contribution is not any single technique — MoE existed, SwiGLU existed, top-K routing existed. The contribution is the engineering proof that these ideas, combined carefully in an open-weight model, produce a system that reshapes the cost-performance frontier. It showed the community that sparse models are not just a research curiosity but a practical path to deploying powerful language models at lower cost.

The line from Mixtral to DeepSeek-V3 — which scaled MoE further — demonstrates that this was not a dead end but the beginning of a new architectural paradigm for frontier LLMs.

CitationJiang, Sablayrolles, Roux, Mensch, Savary, Bamford, Chaplot, de las Casas, Bou Hanna, Bressand, Lengyel, Bour, Lample, Renard Lavaud, Saulnier, Lachaux, Stock, Subramanian, Yang, Antoniak, Le Scao, Gervet, Lavril, Wang, Lacroix, El Sayed. Mixtral of Experts. arXiv, 2024.

Terms in this paper