Inference & Serving
21 min
Four Bits Per Weight: How Low-Precision Quantization Stopped Hurting LLMs
A 70B model in FP16 needs 140 GB of memory it spends most of its time waiting to read. Dropping each weight to four bits cuts that to 35 GB, and for years that cut also broke the model. Here is what changed.