AI Observability
Logging prompts, versions, retrieved context and cost per request.
4 to work through
-
advanced
A fraud model has been in production for eight months. Ground truth arrives weeks later. How do you know if it is still working?
2 min answer -
advanced
Anthropic reported in 2025 that a routing bug sent a share of Claude Sonnet 4 requests to servers configured for a different context length, peaking at 16% of those requests in the worst hour on 31 August, and that it took weeks to identify. Error rates never moved. Which design decision allowed the delay and what telemetry closes it?
3 min answer -
advanced
What must be observable in an AI feature that is not covered by conventional application monitoring?
2 min answer -
advanced
You are asked to log every prompt and response for an assistant handling 2 million interactions a day, so quality regressions are diagnosable. Roughly how much data does that commit you to per year, and does the number change the design?
3 min answer
4 terms in this topic
Inference Telemetry
Recording the full context of each model interaction — inputs, outputs, tokens, latency, model version and evaluation scores — so quality and cost ca…
practiceModel Drift Monitoring
Detecting that a deployed model's inputs or performance have shifted away from the conditions it was validated under.
metricQuality Proxy Metric
A continuously available production signal that moves when answer quality moves - abstention rate, regeneration rate, escalation rate - used to page …
conceptServing Path Divergence
One configuration in a fleet of otherwise identical inference paths behaving differently from its peers - detectable by comparing paths against each …
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.