RNNs & Sequence Models2023advanced11 min read
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
مامبا: نمذجة التسلسلات في زمن خطّي باستخدام فضاءات حالة انتقائية
Gu, A. · Dao, T. — COLM
The problem
Transformers scale quadratically with sequence length — processing 10× more tokens costs 100× more compute. Their grows linearly, making slow. Prior models (linear attention, global convolutions, structured state space models) scale better but fail on language because they cannot selectively focus on relevant tokens — they treat every input identically regardless of content.
The contribution
Selective State Space Models (S6): making the SSM parameters (Δ, B, C) functions of the input so the model can choose what to remember and what to forget at each step. A hardware-aware algorithm computes this selective efficiently on GPUs without materializing the full state in slow memory. The Mamba architecture wraps this selective SSM in a simplified block that merges the SSM with gated MLPs — no attention, no separate block. Mamba-3B matches -quality on language while being 5× faster at inference.
The impact
Mamba proved that attention is not the only path to Transformer-quality language modeling. It sparked a wave of SSM-based architectures (Mamba-2, Vision Mamba, Jamba) and hybrid attention-SSM models. Its selection principle — making dynamics input-dependent — became a design axiom for all subsequent state space models. The hardware-aware scan technique influenced efficient implementations beyond SSMs.
A Transformer reading a book is a conference room: every word sits at one table and can address any other word directly — powerful, but the table must grow with the number of guests, and a million-word book needs a million seats.
An is a single courier carrying a suitcase through the book: fast to move, but the suitcase has fixed capacity and the courier cannot choose what to pack — everything gets the same compression.
Mamba is a smart courier with a suitcase that has a selective lock: at each word, the courier decides — based on the word itself — whether to snap it into the suitcase or let it pass. Important facts get locked in; filler words slide off. The suitcase never grows, so the courier always moves at the same speed regardless of book length.
The problem: attention is powerful but expensive
By 2023, virtually every — GPT, Claude, Gemini, BERT — was built on the Transformer. Its mechanism connects every to every other token in a single step, which is why it models language so well.
But this power comes at a steep price:
-
Quadratic compute. Self-attention compares every pair of tokens: doubling the sequence length quadruples the work. A 100K-token document costs 10,000× more than a 1K-token paragraph.
-
Linear KV cache. During inference, the model stores a Key and Value vector for every past token. As context grows, so does this cache, consuming memory and slowing down generation.
Many alternatives tried to fix this — linear attention, global convolutions, structured state space models like S4 — but all shared the same blind spot: they treated every input identically, regardless of whether it was a crucial fact or meaningless noise.
Background: state space models in 60 seconds
Before understanding what Mamba adds, we need the foundation it builds on: structured state space models (S4).
Think of a as a pipeline with hidden memory. An input signal enters, updates a , and produces an output . The hidden state is like the water level in a tank — it accumulates input over time and the output reads from it.
Four matrices control this pipeline: governs how the state evolves on its own (the leak rate of the tank), controls how new input flows in, controls what we read out, and sets the step size (). Crucially, in prior SSMs all four are fixed — the same for every input. This is called (LTI).
The core insight: selection as compression
Gu and Dao identified a fundamental tension: sequence modeling is about compressing context into a finite state. Attention avoids compression entirely by storing every token (the KV cache) — effective but expensive. Recurrent models compress into a fixed-size state — efficient but lossy.
The key question is: what gets compressed? In LTI models, the compression is content-blind: the same matrix governs every step, so the model cannot choose to remember "Harry" more than "um". This is why LTI models fail on tasks requiring , like the Selective Copying task: copying specific tokens (not all tokens) from scattered positions.
The solution is selection: make the SSM parameters , , and depend on the current input. Now the model can:
- Remember a relevant token by opening the gate ( large → state resets and absorbs the new input)
- Ignore noise by closing the gate ( small → state persists, input slides off)
- Control what enters the state ( input-dependent) and what exits ( input-dependent)
The selective SSM: from S4 to S6
The change from S4 to Mamba's S6 is remarkably simple. Instead of parameters being fixed constants, three of them become functions of the input:
- — the input is now input-dependent
- — the output projection is now input-dependent
- — the step size is now input-dependent
The matrix stays fixed (it only matters through its interaction with via discretization, so selectivity in already provides selectivity in ).
This is the entire innovation at the mathematical level. But it has a profound consequence: the model is no longer time-invariant, which means the shortcut no longer works.
Δ is a learned gate: connection to LSTMs
The is not as alien as it might seem — it has deep roots in classical gating. The paper proves (Theorem 1) that when the state dimension is 1, the selective SSM reduces exactly to a gated RNN:
, then
This is the familiar LSTM-style gate: near 1 means "forget the old state, absorb the new input"; near 0 means "keep the old state, ignore the current input."
Mamba generalizes this in two ways: the state dimension is much larger (typically 16), giving the model a richer hidden representation; and the gating comes from principled discretization of a continuous system rather than heuristic design.
Making it fast: the hardware-aware scan
Making SSM parameters input-dependent breaks the convolution shortcut. The naive recurrence is sequential and requires materializing a state of shape — far larger than the input. On a GPU, this means constant trips between slow global memory () and fast local memory (), which is the real bottleneck for non-matmul operations.
Mamba solves this with three classical techniques fused together:
-
. Load the small parameters from HBM to SRAM once, perform discretization and the full scan and the output multiplication all in SRAM, then write only the final output back to HBM. This reduces memory reads by a factor of (the state dimension).
-
Parallel scan. The recurrence is associative, so it can be parallelized with a work-efficient prefix sum (Blelloch scan), avoiding the sequential bottleneck.
-
. During the backward pass, instead of storing the full intermediate states, recompute them on the fly from the small inputs already in SRAM. This keeps memory usage equal to FlashAttention.
The Mamba block: simplicity by design
Most SSM architectures (H3, Hyena) alternate between an SSM block and a standard MLP block, similar to how Transformers alternate attention and MLP. Mamba simplifies this by merging both into a single homogeneous block, inspired by the Gated Attention Unit.
Each Mamba block has two parallel branches:
- Branch 1 (main): Linear projection → 1D convolution → SiLU activation → Selective SSM
- Branch 2 (gate): Linear projection → SiLU activation
The outputs of both branches are multiplied element-wise, then projected back to the model dimension. The 1D convolution is a short local convolution (kernel size 4) that provides local context before the SSM processes global context.
Two Mamba blocks have the same count as one Transformer attention+MLP pair (about parameters), making them directly comparable.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def softplus(x):
return np.log1p(np.exp(x))
def selective_ssm(x, A, W_B, W_C, W_delta):
"""Selective SSM (S6): parameters depend on input.
x: (L, D) input sequence
A: (D, N) state matrix (fixed)
W_B, W_C: (D, N) projection weights
W_delta: (D, 1) step size projection
"""
L, D = x.shape
N = A.shape[1]
h = np.zeros((D, N)) # hidden state
outputs = []
for t in range(L):
# ── Selection: parameters depend on x[t] ──
B_t = x[t] @ W_B # (N,) — input-dependent
C_t = x[t] @ W_C # (N,) — input-dependent
delta_t = softplus(x[t] @ W_delta) # (D,) — input-dependent
# ── Discretize ──
A_bar = np.exp(delta_t[:, None] * A) # (D, N)
B_bar = delta_t[:, None] * B_t[None, :] # (D, N)
# ── Recurrence ──
h = A_bar * h + B_bar * x[t, :, None] # (D, N)
y_t = (h * C_t[None, :]).sum(axis=-1) # (D,)
outputs.append(y_t)
return np.stack(outputs) # (L, D)
# The only difference from S4: B, C, and delta are computed FROM x[t],
# not fixed. This lets the model choose what to remember and forget.
# That's it. That's Mamba's core innovation.Results: matching Transformers at half the size
Mamba was evaluated across three modalities — language, DNA, and audio — and set new records for sub-quadratic models on all of them:
-
Language: Mamba-3B outperformed Transformers of the same size (Pythia-3B) by 4 points on common-sense and even exceeded Pythia-7B — a model more than twice its size. Against strong baselines using the LLaMA recipe (Transformer++), Mamba matched at all scales from 125M to 1.3B parameters.
-
DNA: Mamba was the only model whose perplexity improved continuously with longer context, all the way up to million-length sequences. On a task distinguishing five great ape species (which share 99% of their DNA), Mamba significantly outperformed HyenaDNA.
-
Audio: On speech generation (SC09), a small 6M-parameter Mamba outperformed much larger GAN and diffusion models. A parameter-matched 24M model cut the FID score to 0.67 — the best autoregressive result to date.
-
Speed: Mamba achieves 5× higher inference than Transformers of similar size, because it has no KV cache — just a fixed-size hidden state. This means a Mamba-6.9B at inference is faster than a Transformer-1.3B.
Why it mattered
2020
S4 — structured state spaces conquer long range
Gu et al. introduced S4, which solved the Long Range Arena benchmark by combining continuous-time dynamics with efficient convolutional computation. But S4 struggled on language — its LTI nature could not do content-based reasoning.
2023
H3 and Hyena — SSMs get closer to Transformers
H3 sandwiched S4 between gated connections; Hyena replaced S4 with an MLP-parameterized convolution. Both narrowed the gap with Transformers but still fell short on language.
2023
Mamba — selection breaks the quality barrier
By making SSM parameters input-dependent and solving the resulting computational challenge with a hardware-aware scan, Mamba became the first linear-time model to match Transformer quality on language while being 5× faster.
2024
Mamba-2 — SSM-attention duality
Dao and Gu proved that selective SSMs and attention are mathematically related through structured matrix decompositions, enabling even faster algorithms and hybrid architectures.
2024
Jamba — first production hybrid SSM-Transformer
AI21 Labs released Jamba, interleaving Mamba and attention layers with mixture-of-experts to handle a 256K context window efficiently — proving the SSM-Transformer hybrid is practical at scale.
Mamba showed that making recurrence content-aware — a conceptually simple change — is enough to close the gap between linear-time models and Transformers. The principle of selection extends beyond SSMs: any recurrent architecture benefits from input-dependent dynamics. This is now an active frontier of research, with Mamba-3, state-space-attention hybrids, and vision/audio Mamba variants building on this foundation.
CitationGu, Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. COLM, 2024.
Terms in this paper
- Selective State Space Modelنموذج فضاء الحالة الانتقائي
- Linear Time Invarianceثبات الزمن الخطّي
- Discretizationالتمييز
- Parallel Scanالمسح المتوازي
- Kernel Fusionدمج النواة
- High Bandwidth Memory (HBM)الذاكرة عالية النطاق
- SRAMالذاكرة الساكنة
- Selection Mechanismآلية الانتقائية
- Zero-Order Holdالاستمساك بالمرتبة صفر
- Content-Based Reasoningالاستدلال القائم على المحتوى
- Global Convolutionالالتفاف الشامل
- Recomputationإعادة الحساب
- Subquadraticدون التربيعي
- Induction Headsرؤوس الاستقراء