Term Kind Topic What it is
Context Window Budget concept LLM Application Architecture The finite token allowance per request, treated as an engineering resource to be allocated deliberately between system instructions, retrieved context, history and output.
Inference Request Path concept LLM Application Architecture The sequence of stages an LLM application request passes through, each with distinct latency, cost and failure characteristics.
KV Cache Attention Cache, Key-Value Cache, Prefix Cache concept LLM Application Architecture The per-session key and value tensors a transformer must hold in GPU memory to generate each subsequent token - the resource that limits concurrent sessions, and whose reuse across turns is the difference betw…