Skip to content
∑ Praveen T N AI & ML
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
AI & ML/ Writing/Tagged “roofline”

Tagged “roofline”

2 posts.

Clear
All Model Architecture22 Training & Alignment22 Inference & Serving23 Agents & Orchestration12 Reasoning & Evaluation27 Safety, Security & Governance9 Platforms & Practice22
Inference & Serving 29 min

Arithmetic Intensity: Why Your GPU Is Idle 99% of the Time

An H100 advertises 989 teraflops. Generating one token from an 8B model uses roughly 0.3% of that. The gap is not a bug in your code or a missing compiler flag; it is a single ratio, FLOPs per byte moved, and almost every perform…

neural-plumbing gpu performance kernels ∑ ◫
Inference & Serving 24 min

Running Language Models on a Phone: Memory Bandwidth, NPUs, and the Few-Billion-Parameter Ceiling

Phones ship NPUs rated in trillions of operations per second, yet the speed at which a reply appears is set by how fast LPDDR memory can hand a couple of gigabytes of weights to the processor, over and over. This post derives the…

on-device-and-edge-ai on-device-ai llm-inference memory-bandwidth ∑ ◫
The library

1099 concepts, 11,628 flashcards and 137 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee Contact Privacy policy Terms of use

Written and maintained by Praveen T N.

© 2026 Praveen T N