Inference & Serving
21 min
PagedAttention and Continuous Batching: How vLLM Stopped Wasting Your GPU
A GPU loaded with a 13B model can have most of its KV-cache memory sitting idle while requests queue for capacity. PagedAttention and continuous batching reclaim that memory, and the throughput follows.