AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
4 to work through
-
intermediate
Several product teams call model providers directly. What does introducing an AI gateway buy, and what does it cost?
2 min answer -
intermediate
Six teams are calling model providers directly from their applications. Justify an AI gateway, and say what it should not do.
2 min answer -
advanced
A platform routes all model calls through an internal gateway. What belongs there, and what must not?
2 min answer -
advanced
An AI gateway now terminates streaming responses for six products and holds a semantic cache shared across tenants. Two properties of that design have caused real outages and real data exposure at large AI providers. What has the organisation taken on and when does the bill arrive?
3 min answer
2 terms in this topic
AI Gateway Pattern
A single control point through which all model calls pass, providing routing, cost control, caching, logging, guardrails and provider abstraction.
practiceToken Budget Enforcement
Limiting token consumption per user, tenant, feature or time window at a central point, so cost cannot run away unobserved.
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.