Mamba & Mamba-2 Architecture
Mamba makes state space models selectively content-aware by generating transition matrices directly from input tokens, while Mamba-2 unifies this recurrence with attention via matrix duality.
Why Does This Exist?
Prior Structured State Space models such as s4 established that linear time-invariant (LTI) systems can train in parallel via Fast Fourier Transform (FFT) convolutions while serving autoregressively in constant time. Despite their computational elegance, LTI models hit a fundamental capability ceiling on language and reasoning tasks. Because their state transition matrices are fixed constants across all time steps, an LTI system cannot decide to retain or discard information based on what it reads. It compresses every incoming token into its state vector with identical weight, making it mathematically incapable of solving selective copying or associative retrieval ("needle in a haystack").
Transformers solve this effortlessly via input-dependent attention weights (), but pay an computation cost and require hundreds of gigabytes of Key-Value cache memory at scale.
Mamba (Mamba-1) broke this trade-off by making the state space parameters dynamic functions of the current input token . However, making parameters time-varying destroys the convolutional representation: the convolution kernel can no longer be precomputed, rendering standard FFT parallelization impossible. Mamba resolved this with a hardware-aware parallel associative scan executed directly in fast GPU on-chip SRAM.
Mamba-2 pushed this paradigm further by proving State Space Duality (SSD): a theoretical equivalence showing that selective SSMs and masked linear attention are dual perspectives of the same 1-semiseparable matrix transformation. By reformulating the recurrence as structured block-matrix multiplications, Mamba-2 utilizes GPU Tensor Cores natively, achieving 2 to 8 times higher training throughput than Mamba-1.
Before analyzing selective mechanics, make sure you understand the foundational continuous formulation in state-space models and quadratic attention in transformers.
Think of It Like This
A legal stenographer with selective highlighter pens
An LTI state space model like S4 is like a mechanical dictaphone that records room audio at an unvarying sampling rate. It records background traffic noise, throat-clearing, and legal arguments with the exact same fidelity. When storage fills up, all historical sounds are compressed equally, drowning the critical testimony in ambient noise.
Mamba-1 gives the stenographer cognitive discretion. When irrelevant banter or procedural noise enters the courtroom, the stenographer sets the step size , effectively pressing pause on the tape recorder. When crucial testimony begins, spikes, opening the input gate and recording the statement into the memory state with high fidelity. The state vector only expends capacity on tokens that matter.
Mamba-2 recognizes that the stenographer’s selective transcript can be read in two identical ways:
- Sentence by sentence (the recurrent view, ideal for live testimony).
- As a block-structured reference ledger, where each witness's testimony is processed in chunks using high-speed optical scanners (Tensor Core matrix multiplications).
Both workflows produce the exact same final legal brief, but the structured ledger operates at industrial printing press speed.
How It Actually Works
Selective State Spaces, SRAM Fusion, and State Space Duality
1. The Mamba-1 Selection Mechanism
In Mamba-1, input sequence passes through linear projection layers to dynamically generate time-varying discretization step sizes and projection matrices:
Here, and vary for each token, and acts as a content-based gating mechanism. The continuous diagonal matrix is discretized via zero-order hold (ZOH):
The recurrent update equation becomes:
When , and : the current input is ignored, and previous state is preserved indefinitely. When is large, and absorbs : the model resets its state and overwrites memory with fresh information.
2. Hardware-Aware Parallel Scan
Because varies with time , global convolution is mathematically invalid. A naive sequential loop over steps on a GPU would cause severe memory bandwidth bottlenecks because reading and writing the state tensor across High-Bandwidth Memory (HBM, ~2 TB/s) is catastrophically slow.
Mamba bypasses HBM by fusing the scan into a single GPU kernel:
- Load parameters from HBM to fast on-chip SRAM (~20 TB/s).
- Discretize in SRAM.
- Execute a parallel associative prefix scan (Blelloch scan) across time steps directly within GPU thread blocks.
- Compute and write only the final output back to HBM.
Intermediate states are never saved to global GPU memory during the forward pass; they are recomputed on the fly during backpropagation via kernel recomputation.
3. Mamba-2 and State Space Duality (SSD)
Mamba-2 constrains the transition matrix to be a scalar multiplier per channel (, scaled by identity ). This simplifies the transformation into a 1-semiseparable matrix :
The entire SSM operation over sequence length can be expressed as:
Comparing this to causal linear attention , the mapping is exact:
- Query
- Key
- Value
- Attention Mask is weighted by cumulative decay factors .
Because is a structured matrix multiplication, Mamba-2 partitions the sequence into blocks (e.g., block size 64) and executes the intra-block computations using GPU Tensor Cores (Matrix Multiply-Accumulate / MMA instructions). This bridges the hardware utilization gap between SSMs and Transformers.
Worked Example
Consider a single state dimension () over 2 time steps () showing how selectively retains or forgets memories:
Parameters:
Time step 1 (Important token):
Let the model predict a large step size: .
Assuming initial state :
Token is absorbed into the state vector.
Time step 2 (Irrelevant distractor token):
The model encounters filler content and selects a tiny step size: .
Updating the state:
Notice that even though distractor input had twice the magnitude of ( vs ), the selective gate suppressed 's contribution to just , while preserving 99% of prior memory ().
Code
Here is a minimal, self-contained Python implementation of the selective associative scan showing how content-dependent modulates state evolution:
import numpy as np
def selective_ssm_step( x: np.ndarray, delta: np.ndarray, a_scalar: float, b: np.ndarray, c: np.ndarray) -> np.ndarray: """Simulate a selective SSM forward pass across sequence length L.""" seq_len = x.shape[0] state_dim = b.shape[1] y = np.zeros(seq_len)
h = np.zeros(state_dim)
for t in range(seq_len): # Discretize continuous parameters based on input-dependent delta[t] a_bar = np.exp(delta[t] * a_scalar) b_bar = delta[t] * b[t]
# Selective recurrent state update h = a_bar * h + b_bar * x[t] y[t] = np.dot(c[t], h)
return y
# Sequence of 3 tokens: Token 0 (signal), Token 1 (noise), Token 2 (query)x = np.array([5.0, 100.0, 1.0])# Model chooses high delta for important token, tiny delta for noisedelta = np.array([1.5, 0.001, 1.0])
a_scalar = -1.0b = np.array([[1.0], [1.0], [1.0]]) # State dimension N = 1c = np.array([[1.0], [1.0], [1.0]])
outputs = selective_ssm_step(x, delta, a_scalar, b, c)
print(f"Token inputs: {x}")print(f"Selective delta: {delta}")print(f"Model outputs: {np.round(outputs, 4)}")# -> Token inputs: [ 5. 100. 1.]# -> Selective delta: [1.5 0.001 1. ]# -> Model outputs: [7.5 7.5925 3.7915]Watch Out For
Memory state capacity bottleneck during complex in-context retrieval
While Mamba matches Transformers on language modeling perplexity, its fixed state vector creates a fundamental information-theoretic bottleneck. Unlike Transformer attention, which stores every prior token's Key and Value vector uncompressed in memory ( capacity), an SSM must compress an arbitrary number of tokens into a fixed state vector. For tasks requiring exact retrieval across dozens of distinct, arbitrary key-value facts scattered across 100k tokens (multi-query associative recall), pure SSMs degrade. For such workloads, hybrid architectures like jamba that interleave SSM layers with sparse attention layers are required.
The Quick Version
- Mamba replaces static LTI state spaces with content-dependent parameters .
- Content-dependent gating allows Mamba to selectively filter noise and retain critical facts across long horizons.
- Because selection breaks FFT convolution, Mamba uses a hardware-aware parallel scan executed entirely in fast GPU SRAM.
- Mamba-2 introduces State Space Duality (SSD), proving the mathematical equivalence of selective SSMs and masked linear attention.
- Mamba-2 executes intra-chunk recurrence using GPU Tensor Cores, achieving dramatic training throughput gains over Mamba-1.