Model Selection
Capability, latency, cost and the evaluation that decides between them.
5 to work through
-
intermediate Multiple choice
A client wants a model that follows their specific document formatting conventions and answers from their knowledge base. What do you recommend?
2 min answer -
intermediate
A content platform must choose between a large hosted model, a smaller hosted model and a self-hosted open model for a moderation task. What decides it?
2 min answer -
intermediate
A recommendation service's model server is down at peak. Should the application fall back to an older model, precomputed candidates, popularity baselines or cached per-user results — and how is the fallback tested?
2 min answer -
intermediate
Your music recommender uses collaborative filtering. Newly released tracks are never recommended. Why, and what do you do?
2 min answer -
advanced
A platform must choose which model serves which request class. What should drive the decision, and what changes over time?
2 min answer
3 terms in this topic
Model Cascade
A cheap fast model handling the clear majority of cases with escalation to a larger model or a human for the uncertain ones - usually a large cost re…
patternModel Routing
Directing each request to the cheapest model capable of handling it, rather than sending all traffic to the most capable one.
practiceSmall Model Routing
Sending each request to the smallest model that can handle it, escalating to a larger one only when needed.
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.