Running 70B parameter Large Language Models (LLMs) in FP16 precision requires over 140 GB of VRAM—demanding two NVIDIA H100 or A100 GPUs. Model Quantization solves this hardware bottleneck by compressing 16-bit floating-point weights down to 4-bit integers (`INT4`) or 4-bit floating-point representations (`NF4`), dropping memory requirements down to ~38 GB while retaining 99%+ of zero-shot perplexity performance.

However, quantization algorithms differ drastically in how they select which weights to compress. In this guide, we dive deep into Activation-aware Weight Quantization (AWQ), GPTQ second-order Hessian matrix optimization, GGUF CPU/Metal quantization, and bitsandbytes 4-bit NormalFloat (NF4).

1. Naive Uniform Quantization vs Outlier Weight Preservation

In standard symmetric linear quantization, 16-bit weight matrices $W \in \mathbb{R}^{m \times n}$ are projected onto a 4-bit integer grid $[-8, 7]$ using a single scale factor $S$:

W_{\text{quant}} = \text{clip}\left( \text{round}\left( \frac{W}{S} \right), -8, 7 \right), \quad S = \frac{\max(|W|)}{7}

The Outlier Problem: LLMs contain emergent salient activations—less than 1% of channels carry 100x larger activation magnitudes than others. Naively rounding these salient channels destroys model reasoning.

2. Activation-Aware Weight Quantization (AWQ)

AWQ protects salient weights by observing activation magnitudes $X$ during a calibration forward pass. Instead of keeping 1% of weights in FP16 (which causes GPU SIMD memory fragmentation), AWQ scales up salient weight channels by a per-channel factor $s > 1$:

# AWQ Channel Scaling Formulation # Scale weight W by s, and scale input activation X by 1/s # W' = W * diag(s), X' = diag(s)^-1 * X # Result: W' * X' = W * X (Mathematically exact!)

By scaling $W$ upward, the quantization grid relative error for salient weights is minimized, preserving accuracy while enabling uniform INT4 execution in GPU CUDA kernels.

3. GPTQ: Second-Order Hessian Matrix Optimization

GPTQ is a Post-Training Quantization (PTQ) method based on Optimal Brain Surgeon (OBS) theory. GPTQ quantizes weight columns one-by-one and updates all remaining unquantized weights in the layer to compensate for the rounding error.

The compensation delta for remaining weights $\Delta W$ is calculated using the inverse Hessian matrix $H = 2 X X^T$:

\Delta W_j = - \frac{w_q - w_q^*}{[H^{-1}]_{qq}} \cdot H^{-1}_{:, q}

GPTQ achieves 4-bit quantization on a 70B model in under 15 minutes of calibration runtime.