Skip to content
∑ Praveen T N AI & ML
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
AI & ML/ Writing/Tagged “quantisation”

Tagged “quantisation”

2 posts.

Clear
All Model Architecture21 Training & Alignment22 Inference & Serving23 Agents & Orchestration12 Reasoning & Evaluation27 Safety, Security & Governance9 Platforms & Practice22
Inference & Serving 29 min

Arithmetic Intensity: Why Your GPU Is Idle 99% of the Time

An H100 advertises 989 teraflops. Generating one token from an 8B model uses roughly 0.3% of that. The gap is not a bug in your code or a missing compiler flag; it is a single ratio, FLOPs per byte moved, and almost every perform…

neural-plumbing gpu performance kernels ∑ ◫
Inference & Serving 27 min

Everything Is Lossy Compression: A Rate-Distortion View of Quantisation, KV Caches, and Distillation

Weight quantisation, KV cache eviction, prompt compression and distillation are treated as four separate engineering disciplines with four separate literatures. They are one problem: choosing a point on a rate-distortion curve. S…

information-theory quantisation kv-cache distillation ∑ ◫
The library

1094 concepts, 11,598 flashcards and 136 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee Contact Privacy policy Terms of use

Written and maintained by Praveen T N.

© 2026 Praveen T N