1. ML Platform advanced

    A model performs well in evaluation and poorly in production. What are the likely causes?

    2 min answer ml-platformtraining-serving-skewdriftfeatures
  2. ML Platform advanced

    A recommendation platform's ML systems span data pipelines, training, evaluation and online serving. Where should the boundaries be, and what causes the most costly class of bug?

    2 min answer ml-platformfeature-storetraining-serving-skewboundaries
  3. ML Platform advanced

    A training run on thousands of GPUs loses several nodes mid-run. How should checkpointing frequency, elastic training, straggler detection and scheduling minimise wasted compute, and what does each checkpoint cost?

    3 min answer metallamatrainingcheckpointing
  4. Model Selection advanced

    A platform must choose which model serves which request class. What should drive the decision, and what changes over time?

    2 min answer model-selectionroutingevaluationcost
  5. Multi-Agent Systems advanced

    A research assistant runs a planner that fans out to four worker agents and a synthesiser that writes the final answer. One worker retrieves the wrong document and returns a fluent confident summary of it. What happens downstream and what stops it?

    2 min answer multi-agenterror-compoundingprovenancehandoff
  6. Multi-Agent Systems advanced

    A team proposes decomposing a workflow into several specialised agents that coordinate. What justifies this over a single agent with more tools, and what does it cost?

    2 min answer multi-agentdecompositioncoordinationcost
  7. Multi-Agent Systems advanced

    A team proposes five specialised agents that collaborate to handle a customer request. Assess.

    2 min answer agentsdesignpragmatism
  8. Multi-Agent Systems advanced

    Interview prompt. A supervisor agent fans out to five worker agents that each call tools with real side effects - creating tickets, sending emails, updating records. One worker fails at its fourth tool call, after two of those side effects have already happened. Tell me what the system does next.

    3 min answer multi-agentidempotencycompensationcheckpointing
  9. Prompt Injection Defence advanced

    An AI assistant reads code and issues from repositories, including untrusted ones, and can call tools. Why is prompt injection a structural problem, and what actually mitigates it?

    2 min answer prompt-injectionuntrusted-contenttoolsisolation
  10. Prompt Injection Defence advanced

    An agent reads customer emails and can issue refunds. What is the threat and how do you bound it?

    2 min answer prompt-injectionleast-privilegehuman-in-the-looptrust-boundaries
  11. Prompt Injection Defence advanced

    An agent with tool access processes untrusted content. Why is prompt injection not solvable at the prompt layer, and what architectural controls actually limit the damage?

    3 min answer prompt-injectionagentstool-callingleast-privilege
  12. Prompt Injection Defence advanced

    An application passes user input into a model that can call tools. What is the threat, and where must the control live?

    2 min answer hasuraprompt-injectiontoolsauthorisation