Term Kind Topic What it is
Small Model Routing Model Cascade, Tiered Inference practice Model Selection Sending each request to the smallest model that can handle it, escalating to a larger one only when needed.
Spotify Discover Weekly: Three Models, One Playlist Discover Weekly case-study ML Platform Spotify combined collaborative filtering, natural language processing and raw audio analysis because each covers the others' blind spots.
Streaming Passthrough Gateway Token Stream Relay, SSE Passthrough Proxy pattern AI Gateways Relaying a model's token stream through a shared proxy without buffering it, which makes the gateway's capacity unit concurrent long-lived connections rather than requests per second.
Subagent Context Isolation Parallel Context Windows, Context Partitioning Across Agents concept Multi-Agent Systems The property that makes multi-agent systems worth their cost - each subagent explores with its own context window and returns only a condensed result, so the system reads far more than one context could hold.
Token Budget Enforcement practice AI Gateways Limiting token consumption per user, tenant, feature or time window at a central point, so cost cannot run away unobserved.
Token Cost Attribution practice AI Cost Management Assigning inference spend to features, tenants and users, so that cost can be managed by the people who influence it.
Tool Authorisation Boundary Model as Untrusted Proposer, Authorise Outside the Model concept Prompt Injection Defence Authorising every tool invocation against the initiating user's own permissions, outside the model - because the model cannot distinguish instructions from data and no prompt-level defence is reliable.
Tool Calling Function Calling pattern AI-Era Architecture Giving a model a set of typed function definitions it can request to invoke, with the application executing the call and returning the result.
Tool Result Budget Tool Output Capping, Result Truncation Contract pattern Tool Calling A hard cap on the tokens any single tool may return into the model's context, with pagination and summarisation behind it, so that one unlucky query cannot fill the window and end the run.
Tool Schema Design practice Tool Calling Defining the tools available to a model — names, descriptions, parameters and errors — in a way that makes correct selection likely.
Training-Serving Skew Online-Offline Skew, Feature Skew concept ML Platform A divergence between the features a model was trained on and the features computed at serving time - producing a model that performs well offline and worse in production, with nothing erroring.
Two-Stage Retrieval pattern Reranking Retrieving a broad candidate set cheaply and then reordering it with an expensive, more accurate model.
Vector Database tool AI-Era Architecture A store optimised for approximate nearest-neighbour search over high-dimensional embeddings.
Vector Index Quantisation Product Quantisation, Scalar Quantisation, Vector Compression pattern Vector Databases Storing embeddings in fewer bits so a large index fits in memory, trading a measurable amount of recall for a multiple-times reduction in RAM and a rescoring stage to win most of it back.