AI-Era Architecture

General material on architecting systems that include models.

3Questions
13Flashcards
15Terms
Terminology

15 terms in this topic

pattern

AI Gateway

A shared proxy in front of model providers that centralises routing, keys, quotas, caching, logging and safety policy.

concept

Context Window

The maximum number of tokens a model can attend to in one request, holding the system prompt, history, retrieved context, tools and the answer.

concept

Embedding

A dense numeric vector representing a piece of content, positioned so that semantically similar content sits nearby.

pattern

Guardrail

A deterministic check applied to a model's input or output, enforcing rules that cannot be left to the model itself.

pattern

Human in the Loop

Requiring human review or approval at a defined point in an automated flow, chosen by the reversibility and cost of the action.

practice

LLM Evaluation

A repeatable measurement of whether an AI system's outputs are good enough, on cases that reflect the actual task.

protocol

Model Context Protocol

An open protocol that standardises how AI applications connect to external tools, data sources and prompts.

pattern

Model Router

Directing each request to a model chosen by the task's difficulty, cost and latency budget, rather than sending everything to the largest model available.

concept

Prompt Injection

An attack in which text from an untrusted source is interpreted by the model as instructions rather than as data.

tool

Prompt Registry

A versioned store of production prompts with their model bindings, parameters and evaluation results, so a prompt change is a reviewable, traceable, …

practice

Prompt Versioning

Treating prompts as versioned, reviewed, tested artefacts rather than as strings edited in place.

pattern

Retrieval-Augmented Generation

Retrieving relevant documents at query time and putting them in the model's context, so answers are grounded in your data rather than in training data.

pattern

Semantic Cache

Caching model responses keyed by the meaning of the request rather than by its exact text, so near-duplicate questions are served without a model call.

pattern

Tool Calling

Giving a model a set of typed function definitions it can request to invoke, with the application executing the call and returning the result.

tool

Vector Database

A store optimised for approximate nearest-neighbour search over high-dimensional embeddings.

AI-Era Architecture

Neighbouring topics

LLM Application Architecture

The shape of a production system with a model in the request path.

No content yet

RAG Architecture

Retrieval, grounding, citation and the permissions RAG can enforce.

No content yet

Vector Databases

Approximate nearest-neighbour search, filtering and re-indexing.

No content yet

Embeddings

Dense representations, model coupling and the migration they imply.

No content yet

Chunking & Retrieval

Structure-aware splitting, hybrid search and why chunking dominates quality.

No content yet

Reranking

Cross-encoders improving precision more than a bigger embedding model.

No content yet

Model Selection

Capability, latency, cost and the evaluation that decides between them.

No content yet

AI Gateways

Centralised routing, keys, quotas, caching, logging and safety policy.

No content yet

Prompt & Version Management

Prompts as reviewed, versioned, evaluated production configuration.

No content yet

Agent Architectures

Loops, planning, memory and the boundaries an agent must not cross.

No content yet

Tool Calling

Typed tool interfaces, narrow parameters and per-tool authorisation.

No content yet

Multi-Agent Systems

Coordination, hand-off and whether more agents actually help.

No content yet

LLM Evaluation

Held-out sets, rubric judging, CI gates and production sampling.

No content yet

AI Observability

Logging prompts, versions, retrieved context and cost per request.

No content yet

Guardrails

Deterministic checks on input and output that fail closed.

No content yet

Prompt Injection Defence

Breaking the private-data, untrusted-input, outbound-channel combination.

No content yet

AI Cost Management

Token accounting, routing, caching and the context-window budget.

No content yet

Human in the Loop

Gating by reversibility and blast radius, and avoiding approval fatigue.

No content yet

ML Platform

Feature stores, training pipelines, registries and deployment.

No content yet