Processing sequence lengths beyond 100,000 tokens introduces severe single-GPU High-Bandwidth Memory (HBM) walls. Even with memory-efficient FlashAttention kernels, storing 1 Million tokens of KV cache on a single GPU requires over 128 GB of VRAM.

Ring Attention (developed by Liu et al.) overcomes this memory barrier by distributing sequence length $N$ across $N_{\text{gpu}}$ hosts in a logical ring topology. By overlapping inter-GPU peer-to-peer (P2P) KV block transfers with local FlashAttention matrix multiplications, Ring Attention enables training and inference on 1M+ to 10M+ token contexts without communication overhead penalties.

1. The Distributed Sequence Memory Wall

Standard Tensor Parallelism (Megatron-LM) splits hidden weight dimensions across GPUs, but does NOT reduce the sequence length $N$ held per GPU. Sequence Parallelism (DeepSpeed-Ulysses) splits sequence length across GPUs, but requires expensive all-to-all collective communications before and after every attention layer, saturating NVLink switches.

2. How Ring Attention Works: Overlapping P2P Ring Transfers

Ring Attention partitions the input sequence $N$ into $P$ blocks, placing one block $(Q_i, K_i, V_i)$ on each GPU node $i \in [0, P-1]$.

# Ring Attention Iteration Step (GPU i) 1. GPU i computes local attention: FlashAttn(Q_i, K_current, V_current) 2. SIMULTANEOUSLY: GPU i sends (K_current, V_current) to GPU (i+1) via NCCL ring send 3. SIMULTANEOUSLY: GPU i receives (K_next, V_next) from GPU (i-1) via NCCL ring recv 4. Update online softmax statistics (l_i, m_i) 5. Repeat for P steps until KV blocks rotate full ring!