While Transformers dominated deep learning for years, their fundamental bottleneck remains the quadratic time and memory complexity $\mathcal{O}(N^2)$ of self-attention relative to sequence length $N$.
Mamba, created by Albert Gu and Tri Dao, introduces Selective State Space Models (SSMs). Mamba achieves sub-quadratic $\mathcal{O}(N)$ linear time scaling with sequence length and constant $\mathcal{O}(1)$ memory inference per token step, matching or exceeding Transformer quality across language, audio, and genomics.
1. Continuous-Time State Space Equations
State Space Models map an input signal $x(t) \in \mathbb{R}$ to an output $y(t) \in \mathbb{R}$ via an implicit $N$-dimensional latent state $h(t) \in \mathbb{R}^N$:
h'(t) = A h(t) + B x(t), \quad y(t) = C h(t)
Where $A \in \mathbb{R}^{N \times N}$ is the state transition matrix, $B \in \mathbb{R}^{N \times 1}$ is the input matrix, and $C \in \mathbb{R}^{1 \times N}$ is the output projection matrix.
2. Discretization via Zero-Order Hold (ZOH)
Digital computers operate on discrete sequences $(x_0, x_1, \dots)$. To process discrete tokens with timescale step size $\Delta$, the continuous matrices $(A, B)$ are discretized using Zero-Order Hold (ZOH):
3. The Selective Mechanism: Input-Dependent Parameters
Prior SSMs (like S4) used static, time-invariant parameters $(B, C, \Delta)$ across all tokens. Mamba's core innovation is making $(B, C, \Delta)$ input-dependent dynamic functions of the current token $x_k$:
- $B_k = \text{Linear}_B(x_k)$
- $C_k = \text{Linear}_C(x_k)$
- $\Delta_k = \text{Softplus}(\text{Parameter}_\Delta + \text{Linear}_\Delta(x_k))$
This selectivity enables the model to dynamically choose whether to memorize or flush information from its latent state based on token context.
Join the Technical Discussion
Have questions about this architecture? Drop a comment below.