AI-Era Architecture
General material on architecting systems that include models.
7 to work through
-
intermediate Multiple choice
A client wants an assistant that answers questions from 50,000 internal documents which change weekly. RAG or fine-tuning? What actually determines the quality?
2 min answer -
advanced
An AI platform serves computationally expensive requests with unpredictable bursts, where one request can occupy an accelerator for seconds. What is the architectural shape, and which control is most often misplaced?
2 min answer -
advanced
An LLM feature that worked last week now gives worse answers. Nothing was deployed. How do you find out what changed, and what should have been in place?
2 min answer -
advanced
An enterprise AI platform serves many customers on shared inference infrastructure. What isolation is required, and where is the hardest boundary?
2 min answer -
advanced
An inference platform receives computationally expensive requests in unpredictable bursts. Should it use admission control, request queues, dynamic batching, autoscaling, priority classes or pre-provisioned capacity?
2 min answer -
advanced
Legal asks whether customer personal data is being sent to your model provider. What must you be able to answer?
2 min answer -
advanced
You are asked to give an internal AI agent access to the customer database, the ticketing system and outbound email so it can resolve support tickets. What is your response?
2 min answer
13 terms in this topic
AI Gateway
A shared proxy in front of model providers that centralises routing, keys, quotas, caching, logging and safety policy.
conceptContext Window
The maximum number of tokens a model can attend to in one request, holding the system prompt, history, retrieved context, tools and the answer.
conceptEmbedding
A dense numeric vector representing a piece of content, positioned so that semantically similar content sits nearby.
patternHuman in the Loop
Requiring human review or approval at a defined point in an automated flow, chosen by the reversibility and cost of the action.
practiceLLM Evaluation
A repeatable measurement of whether an AI system's outputs are good enough, on cases that reflect the actual task.
protocolModel Context Protocol
An open protocol that standardises how AI applications connect to external tools, data sources and prompts.
patternModel Router
Directing each request to a model chosen by the task's difficulty, cost and latency budget, rather than sending everything to the largest model available.
conceptPrompt Injection
An attack in which text from an untrusted source is interpreted by the model as instructions rather than as data.
toolPrompt Registry
A versioned store of production prompts with their model bindings, parameters and evaluation results, so a prompt change is a reviewable, traceable, …
patternRetrieval-Augmented Generation
Retrieving relevant documents at query time and putting them in the model's context, so answers are grounded in your data rather than in training data.
patternSemantic Cache
Caching model responses keyed by the meaning of the request rather than by its exact text, so near-duplicate questions are served without a model call.
patternTool Calling
Giving a model a set of typed function definitions it can request to invoke, with the application executing the call and returning the result.
toolVector Database
A store optimised for approximate nearest-neighbour search over high-dimensional embeddings.
Neighbouring topics
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.