LLM Application Architecture
The shape of a production system with a model in the request path.
4 to work through
-
beginner Multiple choice
Your chat endpoint streams tokens to the browser. A request fails after 300 of an expected 500 tokens are already on the user's screen. The team wants to apply the same automatic retry policy the rest of their API uses. Why is this different?
3 min answer -
advanced
A workspace product adds AI features over user content. Which architectural decisions dominate, and which are commonly deferred at cost?
2 min answer -
advanced
An LLM chat product serves millions of multi-turn conversations. How do KV-cache reuse, prefix caching, continuous batching and session affinity change the architecture, and what breaks when a session lands on a different GPU?
3 min answer -
advanced
Zoom publicly describes the architecture behind its AI Companion as a federated approach - its own models used alongside third-party frontier models, with work routed by task rather than every request going to a single provider. What problem does that structure solve that a single-provider design does not, and where would copying it be a mistake?
3 min answer
3 terms in this topic
Context Window Budget
The finite token allowance per request, treated as an engineering resource to be allocated deliberately between system instructions, retrieved contex…
conceptInference Request Path
The sequence of stages an LLM application request passes through, each with distinct latency, cost and failure characteristics.
conceptKV Cache
The per-session key and value tensors a transformer must hold in GPU memory to generate each subsequent token - the resource that limits concurrent s…
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.