Executing 7-billion parameter Large Language Models locally inside a web browser without cloud backend APIs was long considered impossible due to WebGL limitations and JavaScript single-threaded overhead. The standardization of WebGPU and WebAssembly (Wasm) SIMD has eliminated this barrier, enabling hardware-accelerated local AI inference directly on client GPUs.

In this web systems engineering guide, we inspect WebGPU Compute Shaders written in WGSL (WebGPU Shading Language), analyze 4-bit weight quantization (AWQ/GPTQ) decompression algorithms, examine zero-copy GPU buffer mapping, and build a working WebGPU matrix multiplication engine in JavaScript.

1. WebGL vs WebGPU: Architecture Evolution

Legacy WebGL treated compute tasks as artificial fragment shader rendering passes over 2D textures. In contrast, WebGPU provides low-overhead direct access to modern GPU hardware concepts (similar to Vulkan, Metal, and Direct3D 12):

  • Compute Shaders & Workgroups: Executes arbitrary parallel mathematical kernels across 3D grid workgroups (`@workgroup_size(X, Y, Z)`).
  • Explicit Storage Buffers (`GPUBuffer`): Provides direct read-write access to GPU VRAM array buffers without texture coordinate overhead.
  • Non-blocking Command Encoders: Batches GPU dispatches asynchronously off the main JavaScript UI loop.

2. On-The-Fly 4-Bit Decompression Math in WGSL

To fit a 3B parameter model into a client browser's 4 GB VRAM allocation, model weights are quantized from FP16 (16 bits) down to INT4 (4 bits). During matrix multiplication, a WGSL compute shader unpacks two 4-bit integers stored inside a single `u32` byte on-the-fly:

$$W_{\text{unpacked}} = (W_{\text{packed}} \gg \text{shift}) \text{ \& } 0x0F$$ $$W_{\text{dequantized}} = (W_{\text{unpacked}} - \text{zero\_point}) \times \text{scale}$$

3. Production WGSL Compute Shader for 4-Bit Quantized MatMul

@group(0) @binding(0) var inputs : array; @group(0) @binding(1) var weights_packed : array; @group(0) @binding(2) var scales : array; @group(0) @binding(3) var output : array; @compute @workgroup_size(16, 16) fn main(@builtin(global_invocation_id) global_id : vec3) { let row = global_id.y; let col = global_id.x; let K = 4096u; // Matrix inner dimension var acc : f32 = 0.0; // Iterate over inner dimension in packed u32 steps (8 4-bit weights per u32) for (var k = 0u; k < K / 8u; k++) { let packed_val = weights_packed[row * (K / 8u) + k]; let scale = scales[row]; for (var sub_k = 0u; sub_k < 8u; sub_k++) { let shift = sub_k * 4u; let int4_val = f32((packed_val >> shift) & 0x0Fu); let weight_f32 = (int4_val - 8.0) * scale; let in_val = inputs[(k * 8u + sub_k) * 4096u + col]; acc += in_val * weight_f32; } } output[row * 4096u + col] = acc; }

4. JavaScript WebGPU Pipeline Execution Engine

async function initWebGPUInferenceEngine() { if (!navigator.gpu) { throw new Error("WebGPU is not supported on this browser!"); } const adapter = await navigator.gpu.requestAdapter(); const device = await adapter.requestDevice(); // Create GPU Shader Module const shaderModule = device.createShaderModule({ code: ` @group(0) @binding(0) var A : array; @group(0) @binding(1) var B : array; @group(0) @binding(2) var C : array; @compute @workgroup_size(8, 8) fn main(@builtin(global_invocation_id) id : vec3) { let idx = id.y * 64u + id.x; C[idx] = A[idx] + B[idx]; } ` }); console.log("[SUCCESS] WebGPU Compute Pipeline compiled successfully!"); return device; } // Initialize Pipeline initWebGPUInferenceEngine().catch(console.error);