AI Cost Management
Token accounting, routing, caching and the context-window budget.
5 to work through
-
intermediate
A product embeds model calls in several features and spend is growing unpredictably. What controls actually bound it?
2 min answer -
advanced
A travel platform's AI feature costs vary enormously between interactions and the total is growing faster than usage. Which levers apply, in what order?
2 min answer -
advanced
An AI feature launched two months ago now costs more per month than the rest of the platform. What do you investigate?
2 min answer -
advanced
An AI feature was modelled at £0.02 per interaction. Production shows £0.11. Where did the difference come from?
2 min answer -
advanced
An inference platform receives a burst of expensive long-context requests that would starve short interactive ones. Design admission control, priority classes, token-based limits and batching so interactive latency SLOs hold.
3 min answer
3 terms in this topic
Chunked Prefill
Breaking a long prompt's prefill into pieces interleaved with other sequences' decode steps, so that one enormous request cannot stall every in-fligh…
metricInference Cost per Interaction
The fully loaded token cost of one user interaction, including retrieval, retries, guardrails and agent loops, measured rather than estimated.
practiceToken Cost Attribution
Assigning inference spend to features, tenants and users, so that cost can be managed by the people who influence it.
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.