Skip to content
∑ Praveen T N Learning Library
Overview Concepts Flashcards Writing AI Feed Graph Architecture
Portfolio ↗
Overview Concepts Flashcards Writing AI Feed Graph Search Architecture Practice Back to portfolio ↗
Library/ Writing/Tagged “quantisation”

Tagged “quantisation”

2 posts.

Clear
All Model Architecture15 Training & Alignment15 Inference & Serving15 Agents & Orchestration11 Reasoning & Evaluation9 Safety, Security & Governance3 Platforms & Practice12
Inference & Serving 29 min

Arithmetic Intensity: Why Your GPU Is Idle 99% of the Time

An H100 advertises 989 teraflops. Generating one token from an 8B model uses roughly 0.3% of that. The gap is not a bug in your code or a missing compiler flag; it is a single ratio, FLOPs per byte moved, and almost every perform…

neural-plumbing gpu performance kernels ∑ ◫
Inference & Serving 27 min

Everything Is Lossy Compression: A Rate-Distortion View of Quantisation, KV Caches, and Distillation

Weight quantisation, KV cache eviction, prompt compression and distillation are treated as four separate engineering disciplines with four separate literatures. They are one problem: choosing a point on a rate-distortion curve. S…

information-theory quantisation kv-cache distillation ∑ ◫
The library

545 concepts, 4,452 flashcards and 80 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N