RNNs & Sequence Models2023advanced11 min read

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

مامبا: نمذجة التسلسلات في زمن خطّي باستخدام فضاءات حالة انتقائية

Gu, A. · Dao, T. — COLM

The problem

Transformers scale quadratically with sequence length — processing 10× more tokens costs 100× more compute. Their grows linearly, making slow. Prior models (linear attention, global convolutions, structured state space models) scale better but fail on language because they cannot selectively focus on relevant tokens — they treat every input identically regardless of content.

The contribution

Selective State Space Models (S6): making the SSM parameters (Δ, B, C) functions of the input so the model can choose what to remember and what to forget at each step. A hardware-aware algorithm computes this selective efficiently on GPUs without materializing the full state in slow memory. The Mamba architecture wraps this selective SSM in a simplified block that merges the SSM with gated MLPs — no attention, no separate block. Mamba-3B matches -quality on language while being 5× faster at inference.

The impact

Mamba proved that attention is not the only path to Transformer-quality language modeling. It sparked a wave of SSM-based architectures (Mamba-2, Vision Mamba, Jamba) and hybrid attention-SSM models. Its selection principle — making dynamics input-dependent — became a design axiom for all subsequent state space models. The hardware-aware scan technique influenced efficient implementations beyond SSMs.

A Transformer reading a book is a conference room: every word sits at one table and can address any other word directly — powerful, but the table must grow with the number of guests, and a million-word book needs a million seats.

An is a single courier carrying a suitcase through the book: fast to move, but the suitcase has fixed capacity and the courier cannot choose what to pack — everything gets the same compression.

Mamba is a smart courier with a suitcase that has a selective lock: at each word, the courier decides — based on the word itself — whether to snap it into the suitcase or let it pass. Important facts get locked in; filler words slide off. The suitcase never grows, so the courier always moves at the same speed regardless of book length.

The problem: attention is powerful but expensive

By 2023, virtually every — GPT, Claude, Gemini, BERT — was built on the Transformer. Its mechanism connects every to every other token in a single step, which is why it models language so well.

But this power comes at a steep price:

  • Quadratic compute. Self-attention compares every pair of tokens: doubling the sequence length quadruples the work. A 100K-token document costs 10,000× more than a 1K-token paragraph.

  • Linear KV cache. During inference, the model stores a Key and Value vector for every past token. As context grows, so does this cache, consuming memory and slowing down generation.

Many alternatives tried to fix this — linear attention, global convolutions, structured state space models like S4 — but all shared the same blind spot: they treated every input identically, regardless of whether it was a crucial fact or meaningless noise.

Open in Lab
Drag the sequence length slider and watch attention's cost explode while Mamba stays flat.
The demo wakes as you arrive…

Background: state space models in 60 seconds

Before understanding what Mamba adds, we need the foundation it builds on: structured state space models (S4).

Think of a as a pipeline with hidden memory. An input signal x(t)x(t) enters, updates a h(t)h(t), and produces an output y(t)y(t). The hidden state is like the water level in a tank — it accumulates input over time and the output reads from it.

Four matrices control this pipeline: A\boldsymbol{A} governs how the state evolves on its own (the leak rate of the tank), B\boldsymbol{B} controls how new input flows in, C\boldsymbol{C} controls what we read out, and Δ\Delta sets the step size (). Crucially, in prior SSMs all four are fixed — the same for every input. This is called (LTI).

ht=A‾ ht−1+B‾ xtyt=C hth_t = \overline{\boldsymbol{A}}\, h_{t-1} + \overline{\boldsymbol{B}}\, x_t \qquad y_t = \boldsymbol{C}\, h_t
The discrete SSM recurrence — the engine of S4 — At each step, the hidden state mixes its old value (scaled by Ā) with the new input (scaled by B̄), and the output reads from the state through C. In LTI models, Ā, B̄, C are the same at every step.
Open in Lab
Toggle between convolution mode (parallel, for training) and recurrence mode (sequential, for inference) to see how the same SSM is computed differently.
The demo wakes as you arrive…

The core insight: selection as compression

Gu and Dao identified a fundamental tension: sequence modeling is about compressing context into a finite state. Attention avoids compression entirely by storing every token (the KV cache) — effective but expensive. Recurrent models compress into a fixed-size state — efficient but lossy.

The key question is: what gets compressed? In LTI models, the compression is content-blind: the same matrix A\boldsymbol{A} governs every step, so the model cannot choose to remember "Harry" more than "um". This is why LTI models fail on tasks requiring , like the Selective Copying task: copying specific tokens (not all tokens) from scattered positions.

The solution is selection: make the SSM parameters Δ\Delta, B\boldsymbol{B}, and C\boldsymbol{C} depend on the current input. Now the model can:

  • Remember a relevant token by opening the gate (Δ\Delta large → state resets and absorbs the new input)
  • Ignore noise by closing the gate (Δ\Delta small → state persists, input slides off)
  • Control what enters the state (B\boldsymbol{B} input-dependent) and what exits (C\boldsymbol{C} input-dependent)
Open in Lab
Watch how LTI models fail on selective copying (random spacing), while Mamba's selection mechanism solves it perfectly.
The demo wakes as you arrive…

The selective SSM: from S4 to S6

The change from S4 to Mamba's S6 is remarkably simple. Instead of parameters being fixed constants, three of them become functions of the input:

  • Bt=LinearN(xt)\boldsymbol{B}_t = \text{Linear}_N(x_t) — the input is now input-dependent
  • Ct=LinearN(xt)\boldsymbol{C}_t = \text{Linear}_N(x_t) — the output projection is now input-dependent
  • Δt=softplus(Linear(xt))\Delta_t = \text{softplus}(\text{Linear}(x_t)) — the step size is now input-dependent

The matrix A\boldsymbol{A} stays fixed (it only matters through its interaction with Δ\Delta via discretization, so selectivity in Δ\Delta already provides selectivity in A\boldsymbol{A}).

This is the entire innovation at the mathematical level. But it has a profound consequence: the model is no longer time-invariant, which means the shortcut no longer works.

Bt=sB(xt),Ct=sC(xt),Δt=τΔ ⁣(Parameter+sΔ(xt))\boldsymbol{B}_t = s_B(x_t), \quad \boldsymbol{C}_t = s_C(x_t), \quad \Delta_t = \tau_\Delta\!\bigl(\text{Parameter} + s_\Delta(x_t)\bigr)
The selection mechanism — parameters become input-dependent — B and C are projected from the input via simple linear layers. Δ is projected from the input and passed through softplus to ensure positivity. This is the only change from S4 — but it transforms the model from content-blind to content-aware.
Open in Lab
See how Δ acts as a gate: large Δ resets the state and absorbs the new token, small Δ preserves the state and ignores the token.
The demo wakes as you arrive…

Δ is a learned gate: connection to LSTMs

The is not as alien as it might seem — it has deep roots in classical gating. The paper proves (Theorem 1) that when the state dimension is 1, the selective SSM reduces exactly to a gated RNN:

gt=σ(Linear(xt))g_t = \sigma(\text{Linear}(x_t)), then ht=(1−gt) ht−1+gt xth_t = (1 - g_t)\, h_{t-1} + g_t\, x_t

This is the familiar LSTM-style gate: gtg_t near 1 means "forget the old state, absorb the new input"; gtg_t near 0 means "keep the old state, ignore the current input."

Mamba generalizes this in two ways: the state dimension NN is much larger (typically 16), giving the model a richer hidden representation; and the gating comes from principled discretization of a continuous system rather than heuristic design.

gt=σ(Linear(xt)),ht=(1−gt) ht−1+gt xtg_t = \sigma(\text{Linear}(x_t)), \qquad h_t = (1 - g_t)\, h_{t-1} + g_t\, x_t
Theorem 1 — the selective SSM is a generalized gated RNN — When N=1, A=−1, B=1, with softplus activation on Δ, the selective SSM recurrence becomes exactly this gated update rule. The gate g decides the balance between old state and new input.
Open in Lab
Adjust the gate value and watch how the hidden state balance shifts between preserving old memory and absorbing new input.
The demo wakes as you arrive…

Making it fast: the hardware-aware scan

Making SSM parameters input-dependent breaks the convolution shortcut. The naive recurrence is sequential and requires materializing a state of shape (B,L,D,N)(B, L, D, N) — far larger than the input. On a GPU, this means constant trips between slow global memory () and fast local memory (), which is the real bottleneck for non-matmul operations.

Mamba solves this with three classical techniques fused together:

  • . Load the small parameters (Δ,A,B,C)(Δ, A, B, C) from HBM to SRAM once, perform discretization and the full scan and the output multiplication all in SRAM, then write only the final output back to HBM. This reduces memory reads by a factor of NN (the state dimension).

  • Parallel scan. The recurrence ht=Aˉtht−1+Bˉtxth_t = \bar A_t h_{t-1} + \bar B_t x_t is associative, so it can be parallelized with a work-efficient prefix sum (Blelloch scan), avoiding the sequential bottleneck.

  • . During the backward pass, instead of storing the full intermediate states, recompute them on the fly from the small inputs already in SRAM. This keeps memory usage equal to FlashAttention.

Open in Lab
See how the fused kernel avoids round-trips to slow GPU memory. Parameters load once into SRAM, and only the final output goes back to HBM.
The demo wakes as you arrive…

The Mamba block: simplicity by design

Most SSM architectures (H3, Hyena) alternate between an SSM block and a standard MLP block, similar to how Transformers alternate attention and MLP. Mamba simplifies this by merging both into a single homogeneous block, inspired by the Gated Attention Unit.

Each Mamba block has two parallel branches:

  • Branch 1 (main): Linear projection → 1D convolution → SiLU activation → Selective SSM
  • Branch 2 (gate): Linear projection → SiLU activation

The outputs of both branches are multiplied element-wise, then projected back to the model dimension. The 1D convolution is a short local convolution (kernel size 4) that provides local context before the SSM processes global context.

Two Mamba blocks have the same count as one Transformer attention+MLP pair (about 12D212D^2 parameters), making them directly comparable.

Open in Lab
Click any layer in the Mamba block to learn what it does.
The demo wakes as you arrive…

The idea in code

Selective SSM: S4 → S6 in ~25 linespython

Simplified to show the idea — not the real implementation.

import numpy as np

def softplus(x):
    return np.log1p(np.exp(x))

def selective_ssm(x, A, W_B, W_C, W_delta):
    """Selective SSM (S6): parameters depend on input.
    x: (L, D) input sequence
    A: (D, N) state matrix (fixed)
    W_B, W_C: (D, N) projection weights
    W_delta: (D, 1) step size projection
    """
    L, D = x.shape
    N = A.shape[1]
    h = np.zeros((D, N))               # hidden state
    outputs = []

    for t in range(L):
        # ── Selection: parameters depend on x[t] ──
        B_t = x[t] @ W_B               # (N,) — input-dependent
        C_t = x[t] @ W_C               # (N,) — input-dependent
        delta_t = softplus(x[t] @ W_delta)  # (D,) — input-dependent

        # ── Discretize ──
        A_bar = np.exp(delta_t[:, None] * A)     # (D, N)
        B_bar = delta_t[:, None] * B_t[None, :]  # (D, N)

        # ── Recurrence ──
        h = A_bar * h + B_bar * x[t, :, None]    # (D, N)
        y_t = (h * C_t[None, :]).sum(axis=-1)    # (D,)
        outputs.append(y_t)

    return np.stack(outputs)  # (L, D)

# The only difference from S4: B, C, and delta are computed FROM x[t],
# not fixed. This lets the model choose what to remember and forget.
# That's it. That's Mamba's core innovation.

Results: matching Transformers at half the size

Mamba was evaluated across three modalities — language, DNA, and audio — and set new records for sub-quadratic models on all of them:

  • Language: Mamba-3B outperformed Transformers of the same size (Pythia-3B) by 4 points on common-sense and even exceeded Pythia-7B — a model more than twice its size. Against strong baselines using the LLaMA recipe (Transformer++), Mamba matched at all scales from 125M to 1.3B parameters.

  • DNA: Mamba was the only model whose perplexity improved continuously with longer context, all the way up to million-length sequences. On a task distinguishing five great ape species (which share 99% of their DNA), Mamba significantly outperformed HyenaDNA.

  • Audio: On speech generation (SC09), a small 6M-parameter Mamba outperformed much larger GAN and diffusion models. A parameter-matched 24M model cut the FID score to 0.67 — the best autoregressive result to date.

  • Speed: Mamba achieves 5× higher inference than Transformers of similar size, because it has no KV cache — just a fixed-size hidden state. This means a Mamba-6.9B at inference is faster than a Transformer-1.3B.

Why it mattered

  1. 2020

    S4 — structured state spaces conquer long range

    Gu et al. introduced S4, which solved the Long Range Arena benchmark by combining continuous-time dynamics with efficient convolutional computation. But S4 struggled on language — its LTI nature could not do content-based reasoning.

  2. 2023

    H3 and Hyena — SSMs get closer to Transformers

    H3 sandwiched S4 between gated connections; Hyena replaced S4 with an MLP-parameterized convolution. Both narrowed the gap with Transformers but still fell short on language.

  3. 2023

    Mamba — selection breaks the quality barrier

    By making SSM parameters input-dependent and solving the resulting computational challenge with a hardware-aware scan, Mamba became the first linear-time model to match Transformer quality on language while being 5× faster.

  4. 2024

    Mamba-2 — SSM-attention duality

    Dao and Gu proved that selective SSMs and attention are mathematically related through structured matrix decompositions, enabling even faster algorithms and hybrid architectures.

  5. 2024

    Jamba — first production hybrid SSM-Transformer

    AI21 Labs released Jamba, interleaving Mamba and attention layers with mixture-of-experts to handle a 256K context window efficiently — proving the SSM-Transformer hybrid is practical at scale.

Mamba showed that making recurrence content-aware — a conceptually simple change — is enough to close the gap between linear-time models and Transformers. The principle of selection extends beyond SSMs: any recurrent architecture benefits from input-dependent dynamics. This is now an active frontier of research, with Mamba-3, state-space-attention hybrids, and vision/audio Mamba variants building on this foundation.

CitationGu, Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. COLM, 2024.

Terms in this paper