1. ML Platform intermediate

    A machine-learning platform must ingest telemetry from many concurrent training runs. What are the workload's distinguishing characteristics?

    2 min answer weights-biasesml-platformtelemetryingestion
  2. ML Platform advanced

    A model performs well in evaluation and poorly in production. What are the likely causes?

    2 min answer ml-platformtraining-serving-skewdriftfeatures
  3. ML Platform advanced

    A recommendation platform's ML systems span data pipelines, training, evaluation and online serving. Where should the boundaries be, and what causes the most costly class of bug?

    2 min answer ml-platformfeature-storetraining-serving-skewboundaries
  4. ML Platform advanced

    A training run on thousands of GPUs loses several nodes mid-run. How should checkpointing frequency, elastic training, straggler detection and scheduling minimise wasted compute, and what does each checkpoint cost?

    3 min answer metallamatrainingcheckpointing
  5. Model Selection intermediate Multiple choice

    A client wants a model that follows their specific document formatting conventions and answers from their knowledge base. What do you recommend?

    2 min answer ragfine-tuningarchitecture
  6. Model Selection intermediate

    A content platform must choose between a large hosted model, a smaller hosted model and a self-hosted open model for a moderation task. What decides it?

    2 min answer sharechatmodel-selectioncostlatency
  7. Model Selection advanced

    A platform must choose which model serves which request class. What should drive the decision, and what changes over time?

    2 min answer model-selectionroutingevaluationcost
  8. Model Selection intermediate

    A recommendation service's model server is down at peak. Should the application fall back to an older model, precomputed candidates, popularity baselines or cached per-user results — and how is the fallback tested?

    2 min answer fallbackdegradationrecommendationsresilience
  9. Model Selection intermediate

    Your music recommender uses collaborative filtering. Newly released tracks are never recommended. Why, and what do you do?

    2 min answer spotifyrecommendationscold-startensemble
  10. Multi-Agent Systems advanced

    A research assistant runs a planner that fans out to four worker agents and a synthesiser that writes the final answer. One worker retrieves the wrong document and returns a fluent confident summary of it. What happens downstream and what stops it?

    2 min answer multi-agenterror-compoundingprovenancehandoff
  11. Multi-Agent Systems advanced

    A team proposes decomposing a workflow into several specialised agents that coordinate. What justifies this over a single agent with more tools, and what does it cost?

    2 min answer multi-agentdecompositioncoordinationcost
  12. Multi-Agent Systems advanced

    A team proposes five specialised agents that collaborate to handle a customer request. Assess.

    2 min answer agentsdesignpragmatism