The release of DeepSeek-V3 represents a major watershed moment in open-weight Frontier LLM engineering. Operating with 671B total parameters and activating 37B parameters per token via sparse Mixture-of-Experts (MoE), DeepSeek-V3 achieves state-of-the-art reasoning benchmark scores while drastically reducing training compute costs to under $6 million USD.

At the core of this hardware efficiency lie two breakthroughs: Multi-Head Latent Attention (MLA)—which compresses the Key-Value (KV) cache into a low-rank latent vector—and an end-to-end FP8 Mixed-Precision Training Framework coupled with DualPipe overlapping. In this architectural breakdown, we inspect the exact mathematical formulations, CUDA memory trade-offs, and PyTorch implementations behind MLA and FP8 scaling.

1. The Memory Wall: MHA vs GQA vs MLA

In standard Multi-Head Attention (MHA), serving long context windows (e.g. 128k tokens) requires caching Query, Key, and Value vectors across every Transformer layer. For a batch size $B$, sequence length $S$, number of KV heads $n_{kv}$, and head dimension $d_{h}$, the KV cache memory footprint scales as:

Memory_{\text{KV}} = 2 \times B \times S \times n_{kv} \times d_{h} \times \text{BytesPerElement}

While Grouped-Query Attention (GQA) reduces $n_{kv}$ by sharing keys and values across query heads, it trades off model expressiveness and still requires significant GPU memory per active user stream.

2. Multi-Head Latent Attention (MLA) Mathematics

DeepSeek-V3 introduces Multi-Head Latent Attention (MLA). Instead of caching high-dimensional Key ($K$) and Value ($V$) projections directly, MLA projects hidden state $h_t \in \mathbb{R}^d$ down into a compact low-rank latent vector $\mathbf{c}_t^{KV} \in \mathbb{R}^{d_c}$ where $d_c \ll n_{h} d_h$:

# Low-Rank KV Compression Math in MLA c_t_KV = W_DKV * h_t # [b, d_c] Latent KV compression K_t_C = W_UK * c_t_KV # [b, n_h * d_h] Uncompressed Key V_t_C = W_UV * c_t_KV # [b, n_h * d_h] Uncompressed Value

A. Decoupled Rotary Position Embedding (RoPE)

Rotary Position Embeddings (RoPE) are position-sensitive. If position embeddings were applied directly to $\mathbf{c}_t^{KV}$, the uncompression matrix $W^{UK}$ could not absorb $W^{UK}$ during inference projection because RoPE rotation matrices depend on token position $t$.

To solve this, MLA decouples positional keys from content keys:

  • Content Key: $\mathbf{k}_{t,i}^C = W^{UK}_i \mathbf{c}_t^{KV}$ (Position-independent, absorbed into weight matrix during inference).
  • Position Key: $\mathbf{k}_{t}^R = \text{RoPE}(W^{KR} h_t)$ (Shared position key of dimension $d_R$).

3. Inference Memory Acceleration: Matrix Absorption

During generation, because $W^{UK}_i$ is linear and position-independent, the projection matrix $W^{UK}$ can be absorbed directly into the Query projection matrix $W^{UQ}$! Thus, GPUs only store the low-rank vector $\mathbf{c}_t^{KV}$ and position key $\mathbf{k}_t^R$ in memory:

Attention Mechanism KV Cache Bytes per Token (d_h=128, FP16) 128k Context Memory (Batch Size = 1)
Multi-Head Attention (MHA) 32,768 Bytes 4.29 GB per stream
Grouped-Query Attention (GQA - 8 heads) 4,096 Bytes 0.53 GB per stream
DeepSeek MLA (d_c=512 + d_R=64) 1,152 Bytes 0.14 GB per stream (73% memory savings over GQA!)