While Transformers dominated deep learning for years, their fundamental bottleneck remains the quadratic time and memory complexity $\mathcal{O}(N^2)$ of self-attention relative to sequence length $N$.

Mamba, created by Albert Gu and Tri Dao, introduces Selective State Space Models (SSMs). Mamba achieves sub-quadratic $\mathcal{O}(N)$ linear time scaling with sequence length and constant $\mathcal{O}(1)$ memory inference per token step, matching or exceeding Transformer quality across language, audio, and genomics.

1. Continuous-Time State Space Equations

State Space Models map an input signal $x(t) \in \mathbb{R}$ to an output $y(t) \in \mathbb{R}$ via an implicit $N$-dimensional latent state $h(t) \in \mathbb{R}^N$:

h'(t) = A h(t) + B x(t), \quad y(t) = C h(t)

Where $A \in \mathbb{R}^{N \times N}$ is the state transition matrix, $B \in \mathbb{R}^{N \times 1}$ is the input matrix, and $C \in \mathbb{R}^{1 \times N}$ is the output projection matrix.

2. Discretization via Zero-Order Hold (ZOH)

Digital computers operate on discrete sequences $(x_0, x_1, \dots)$. To process discrete tokens with timescale step size $\Delta$, the continuous matrices $(A, B)$ are discretized using Zero-Order Hold (ZOH):

# Zero-Order Hold (ZOH) Discretization Math \bar{A} = \exp(\Delta A) \bar{B} = (\Delta A)^{-1} (\exp(\Delta A) - I) \cdot (\Delta B) # Discrete Recurrence Relation: h_k = \bar{A} h_{k-1} + \bar{B} x_k y_k = C h_k

3. The Selective Mechanism: Input-Dependent Parameters

Prior SSMs (like S4) used static, time-invariant parameters $(B, C, \Delta)$ across all tokens. Mamba's core innovation is making $(B, C, \Delta)$ input-dependent dynamic functions of the current token $x_k$:

  • $B_k = \text{Linear}_B(x_k)$
  • $C_k = \text{Linear}_C(x_k)$
  • $\Delta_k = \text{Softplus}(\text{Parameter}_\Delta + \text{Linear}_\Delta(x_k))$

This selectivity enables the model to dynamically choose whether to memorize or flush information from its latent state based on token context.