ML Platform
Feature stores, training pipelines, registries and deployment.
4 to work through
-
intermediate
A machine-learning platform must ingest telemetry from many concurrent training runs. What are the workload's distinguishing characteristics?
2 min answer -
advanced
A model performs well in evaluation and poorly in production. What are the likely causes?
2 min answer -
advanced
A recommendation platform's ML systems span data pipelines, training, evaluation and online serving. Where should the boundaries be, and what causes the most costly class of bug?
2 min answer -
advanced
A training run on thousands of GPUs loses several nodes mid-run. How should checkpointing frequency, elastic training, straggler detection and scheduling minimise wasted compute, and what does each checkpoint cost?
3 min answer
4 terms in this topic
Feature Store
A system that computes, stores and serves model input features consistently for both training and inference, eliminating training-serving skew.
conceptML Platform
The infrastructure that makes machine learning repeatable — data, features, training, deployment, monitoring — where the model is the small part.
case-studySpotify Discover Weekly: Three Models, One Playlist
Spotify combined collaborative filtering, natural language processing and raw audio analysis because each covers the others' blind spots.
conceptTraining-Serving Skew
A divergence between the features a model was trained on and the features computed at serving time - producing a model that performs well offline and…
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.