Language Models2023intermediate12 min read

Mistral 7B

Mistral 7B

Jiang, A. Q. · Sablayrolles, A. · Mensch, A. · Bamford, C. · Chaplot, D. S. · de las Casas, D. · Bressand, F. · Lengyel, G. · Lample, G. · Saulnier, L. · Lavaud, L. R. · Lachaux, M.-A. · Stock, P. · Le Scao, T. · Lavril, T. · Wang, T. · Lacroix, T. · El Sayed, W. — arXiv

The problem

By late 2023, the dominant approach to improving language models was simply to make them bigger. Llama 2 13B needed almost twice the parameters of a 7B model, yet the gains were incremental. Larger models demanded more memory, higher latency, and greater cost — making them impractical for many real-world applications. The community needed a model that could deliver 13B-level (or better) performance at the 7B scale, with concrete architectural innovations to reduce .

The contribution

Mistral 7B: a 7-billion- language model that outperforms Llama 2 13B on every evaluated and Llama 1 34B in reasoning, math, and code generation. It achieves this through two key mechanisms — (SWA) with a window of 4096 tokens that provides a theoretical span of ~131K tokens across 32 layers, and (GQA) with 8 KV heads shared across 32 query heads, reducing size by 4× and accelerating inference. A caps memory at a fixed size regardless of sequence length. The instruct-tuned variant surpasses Llama 2 13B Chat on both MT-Bench and human evaluations.

The impact

Mistral 7B proved that the cost-performance frontier has three dimensions — capability, cost, and inference cost — not just two. It showed that careful architectural choices can compress knowledge more efficiently than raw scaling suggests. The model spawned Mixtral (a mixture-of-experts extension) and catalyzed a wave of efficient open models. Released under Apache 2.0, it became one of the most widely deployed open-source LLMs and validated that smaller, well-engineered models can compete with much larger ones.

Imagine a librarian in a massive library. The old approach (full attention) means the librarian checks every single book on every shelf before answering any question — thorough, but painfully slow as the library grows. Mistral's librarian works differently: she carries a — a tray that holds only the 4,096 nearest books. But here's the trick: each floor of the library passes notes to the floor above. After climbing 32 floors, she has effectively gathered information from 131,000 books, even though she only looked at 4,096 at a time. And instead of keeping separate notebooks (keys and values) for every question she's tracking, she shares notebooks across groups of questions — cutting her stationery costs by 75%.

The problem: bigger models, diminishing returns

Before Mistral 7B, the recipe for a better language model was straightforward: add more parameters. Llama 2 went from 7B to 13B to 70B, and performance scaled predictably. But this approach has a steep price — a 13B model needs roughly 2× the GPU memory and generates tokens about 2× slower than a 7B model. For real-time applications like chatbots, code assistants, and edge deployment, this cost is not just inconvenient — it's a deployment barrier.

The community had framed the problem in two dimensions: model capability versus training compute. But practitioners care about a third dimension: inference cost. A model that performs identically at half the inference cost is, for deployment purposes, twice as good. Mistral 7B's central bet was that architectural — not just scale — could push the capability frontier forward.

Architecture overview: Transformer with two twists

At its core, Mistral 7B is a standard , sharing the same DNA as LLaMA. It uses with , the in the feed-forward blocks, and Rotary Position Embeddings () for position encoding. The vocabulary is 32,000 tokens via .

What sets Mistral apart is how it computes attention. Two modifications — Sliding Window Attention and Grouped Query Attention — work together to slash memory usage and speed up inference, without sacrificing quality.

Open in Lab
Explore Mistral 7B's architecture. Click each component to learn its role and parameters.
The demo wakes as you arrive…

Sliding Window Attention: see locally, know globally

Standard lets every attend to every other token. This is powerful but expensive: computational cost grows quadratically with sequence length, and the KV cache grows linearly, eating GPU memory at longer contexts.

Sliding Window Attention restricts each token to attending only the previous W tokens (where W = 4,096 in Mistral 7B). At first glance, this might seem limiting — the model can only "see" 4,096 tokens back. But because the Transformer stacks layers, each layer passes information forward. Think of it like a relay race: at layer 1, token i sees tokens i−W to i. At layer 2, it sees information from tokens i−2W to i because layer 1 already mixed information within its window. After k layers, the effective is k × W.

With 32 layers and W = 4,096, the theoretical attention span reaches approximately 131,072 tokens — far beyond the 8,192-token context length the model was trained on. In practice, Mistral achieves a 2× speedup over vanilla attention for sequences of 16K tokens.

Open in Lab
Drag the slider to change the current layer and see how the effective attention span grows across layers. Each layer adds W = 4,096 tokens of reach.
The demo wakes as you arrive…
Effective Span(k)=k×W\text{Effective Span}(k) = k \times W
Effective attention span after k layers — Each layer extends the receptive field by W tokens. After k = 32 layers with W = 4,096, the model can propagate information across up to 131,072 tokens — even though each individual layer only attends to a local window.

Rolling Buffer Cache: fixed memory, unlimited sequences

Since Sliding Window Attention only looks at the last W tokens, there's no need to keep keys and values older than that. Mistral uses a rolling buffer cache of fixed size W: the key-value pair at position i is stored at index i mod W in the cache. When the sequence grows past W, old entries are simply overwritten.

The result is remarkable: on a 32K-token sequence, this reduces KV cache memory by 8× compared to a standard cache — with no impact on model quality. This means memory usage stays constant regardless of sequence length, which is critical for serving long contexts in production.

cache_pos(i)=i mod W\text{cache\_pos}(i) = i \bmod W
Rolling buffer cache indexing — Keys and values at position i are stored at index i mod W in the cache. When i exceeds W, older entries at (i mod W) are overwritten. The cache never grows beyond size W.
Open in Lab
Watch how the rolling buffer overwrites old entries as new tokens arrive. The cache never exceeds size W regardless of sequence length.
The demo wakes as you arrive…

Grouped Query Attention: shared keys, faster inference

In standard , each of the 32 heads maintains its own set of key and value projections. This means the KV cache scales linearly with the number of heads — expensive during autoregressive decoding when the cache is read at every step.

Grouped Query Attention (GQA) groups the 32 query heads into 8 groups, each sharing a single key-value head. Instead of 32 KV heads, only 8 are needed — a 4× reduction in KV cache size. Think of it as a conference room with 32 people (queries) who have questions, but instead of 32 separate reference binders (key-value pairs), there are only 8 shared binders. Each group of 4 people shares one binder. The questions remain distinct and nuanced, but the reference material is shared efficiently.

This has two immediate benefits: (1) less memory — the KV cache is 4× smaller, so you can fit longer sequences or larger batches in the same GPU; (2) faster decoding — reading 8 KV heads instead of 32 reduces , which is typically the bottleneck during token-by-token generation.

Open in Lab
Compare standard Multi-Head Attention (32 KV heads) with Grouped Query Attention (8 KV heads). See how groups of 4 query heads share a single KV head.
The demo wakes as you arrive…

Pre-fill and chunking: efficient prompt processing

During generation, each new token depends on all previous ones — so tokens must be generated one at a time. But the prompt is known in advance, which creates an opportunity for parallel processing.

Mistral pre-fills the KV cache with the entire prompt before generation begins. For very long prompts, this is done in chunks of size W (the window size). Each chunk computes attention over: (1) the existing cache from previous chunks, and (2) the tokens within the current chunk (using a ). This bounds peak memory usage to the window size, no matter how long the prompt is.

Model specifications

Mistral 7B uses a 32-layer Transformer with a hidden dimension of 4,096 and a feed-forward hidden dimension of 14,336. Each attention layer has 32 query heads and 8 key-value heads (4:1 ratio for GQA), each with a head dimension of 128. The sliding window size is 4,096 tokens, and the model supports a context length of 8,192 tokens. The is 32,000 tokens.

Compared to Llama 2 7B, the architecture differs primarily in two ways: (1) the use of GQA instead of standard multi-head attention, and (2) the use of sliding window attention instead of full causal attention.

The idea in code

Sliding Window Attention — the key mechanismpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn.functional as F

def sliding_window_attention(Q, K, V, window_size=4096):
    """
    Sliding Window Attention: each query only attends to
    the previous `window_size` keys, not the full sequence.
    """
    seq_len = Q.size(-2)
    # Create a band mask: each position i attends to [i-W+1, i]
    positions = torch.arange(seq_len)
    # mask[i,j] = True if j is within the window of i
    mask = (positions.unsqueeze(0) - positions.unsqueeze(1)).abs() < window_size
    mask = mask & (positions.unsqueeze(0) >= positions.unsqueeze(1))  # causal

    scores = Q @ K.transpose(-2, -1) / (Q.size(-1) ** 0.5)
    scores = scores.masked_fill(~mask, float('-inf'))
    attn = F.softmax(scores, dim=-1)
    return attn @ V

# Key insight: computational cost is O(n × W) instead of O(n²)
# Memory for KV cache is fixed at W, not growing with sequence length
# After k=32 layers, effective receptive field = k × W ≈ 131K tokens

Results: 7B parameters, 13B+ performance

Mistral 7B was evaluated on a comprehensive suite of benchmarks and outperformed Llama 2 13B on every single one:

  • MMLU (knowledge): 60.1% vs 55.6% for Llama 2 13B — despite having almost half the parameters. Mistral's performance here suggests efficient knowledge compression.
  • HellaSwag (commonsense): 81.3% vs 80.7% — matching a model nearly twice its size.
  • GSM8K (math reasoning, 8-shot): 52.2% vs 34.3% — a dramatic 18-point leap, and even surpassing Llama 1 34B's performance.
  • HumanEval (code generation): 30.5% — approaching Code-Llama 7B (31.1%) without any code-specific . This is notable because Code-Llama sacrifices general performance for code; Mistral doesn't.
  • ARC-Challenge: 55.5% vs 48.8% — a 7-point improvement on scientific reasoning.

When converted to "equivalent model sizes," Mistral 7B mirrors the performance of a Llama 2 model 3× its size on reasoning and STEM tasks, and 1.9× its size on knowledge benchmarks (limited by how much factual knowledge 7B parameters can store).

Open in Lab
Compare Mistral 7B against Llama 2 7B, Llama 2 13B, and Code-Llama 7B across key benchmarks. Toggle models to see the performance gaps.
The demo wakes as you arrive…

Instruction tuning: Mistral 7B Instruct

To demonstrate Mistral 7B's adaptability, the authors fine-tuned it on publicly available instruction datasets — no proprietary data, no special tricks. The resulting Mistral 7B Instruct model outperformed all 7B chat models on MT-Bench and was comparable to 13B chat models.

In a head-to-head human evaluation against Llama 2 13B Chat, evaluators preferred Mistral 7B Instruct's responses 5,020 times versus 4,143 for Llama 2 13B Chat. This is remarkable: a model with nearly half the parameters was preferred by humans in open-ended conversation.

The authors also introduced a mechanism and self-reflection-based . With the recommended system prompt, Mistral properly declined 100% of harmful prompts while still answering legitimate questions that overly cautious models would refuse — like "How to kill a Linux process."

Guardrails and content moderation

Mistral introduced a practical approach to safety that balances caution with utility. The system prompt mechanism lets users move along the Pareto front of helpfulness versus safety. With the Mistral system prompt, the model scored 6.58 on MT-Bench (compared to 6.84 without any prompt), a small cost for full harmful-prompt coverage.

For content moderation, Mistral uses self-reflection: the model classifies its own outputs into categories (illegal activities, hateful content, unqualified advice) using a dedicated prompt. This achieved 99.4% with 95.6% on a curated dataset of adversarial and standard prompts.

What Mistral 7B unlocked

  1. 2023

    Mistral 7B

    7B parameters outperform 13B across all benchmarks. Sliding window attention + Grouped Query Attention make inference fast and memory-efficient. Released under Apache 2.0.

  2. 2024

    Mixtral 8×7B — Mixture of Experts

    Extended Mistral's architecture with 8 expert FFN blocks per layer, using a router to activate 2 experts per token. 46.7B total parameters, only 12.9B active per token. Matched or exceeded Llama 2 70B and GPT-3.5 on many benchmarks.

  3. 2024

    Gemma — Google's efficient open models

    Google released Gemma 2B and 7B, similarly targeting the efficient small-model space that Mistral pioneered. Demonstrated that the focus on inference efficiency had become an industry trend.

  4. 2024

    Efficiency-first open models

    Phi-2, Qwen 1.5, and many others followed the playbook Mistral established: small models with architectural optimizations that punch above their weight class. The era of "bigger is always better" was over.

Mistral 7B's deepest contribution is not any single architectural trick — it's the demonstration that inference efficiency deserves equal attention to training efficiency. The scaling laws community had focused almost exclusively on optimal training compute; Mistral showed that the deployment side of the equation matters just as much. A model that trains expensively but runs cheaply may be more valuable than one that trains cheaply but runs expensively.

This insight shaped everything that followed: Mixtral extended the efficiency thesis with , and the broader community began treating inference cost as a first-class metric alongside and benchmark scores.

CitationJiang, Sablayrolles, Mensch, Bamford, Chaplot, de las Casas, Bressand, Lengyel, Lample, Saulnier, Lavaud, Lachaux, Stock, Le Scao, Lavril, Wang, Lacroix, El Sayed. Mistral 7B. arXiv, 2023.

Terms in this paper