1. Microservices advanced

    Your organisation has grown to 2,000 microservices. Teams report that understanding and changing the system is now harder than before decomposition. What do you do?

    2 min answer uberdomaboundariesgovernance
  2. Migration Risk advanced

    A logistics platform is migrating a system that coordinates live vehicle dispatch. What controls reduce migration risk when the operation cannot pause?

    2 min answer portermigration-riskprogressiveshadowing
  3. Migration Risk advanced

    A platform is migrating a live service and must manage risk. Which risks dominate, and which controls actually reduce them?

    2 min answer migration-riskblast-radiusrollbackverification
  4. Migration Risk advanced

    Ant Group moved Alipay's core transaction data off a commercial relational database onto OceanBase, its own distributed database, beginning with a 1% slice in 2014 and reaching full replacement in 2019. What risk control is that ramp actually providing, and where would copying it be a mistake?

    3 min answer alibabaoceanbasealipayincremental-migration
  5. Migration Risk intermediate

    Review this migration plan. Cutover is staged by customer cohort at 2 percent then 10 then 50 then 100. Readiness for each step is "all automated tests green and no open P1 defects". The rehearsal ran against a 5 percent copy of production. The rollback step reads "restore from the pre-cutover snapshot if data integrity is compromised". What would you remove, what would you change, and what would you leave alone?

    3 min answer migration-riskcutoverabort-criteriarehearsal
  6. Migration Risk intermediate Multiple choice

    What is the single most important risk control in a migration, and why?

    2 min answer migration-riskreversibilityverificationblast-radius
  7. Migration Risk intermediate

    You want to replace a twelve-year-old monolith one capability at a time. The organisation releases that monolith once every six weeks on a Thursday evening, behind a change advisory board that meets fortnightly. Tell me what you would do first and why.

    3 min answer etsydeployment cadencestranglerchange advisory
  8. ML Platform intermediate

    A machine-learning platform must ingest telemetry from many concurrent training runs. What are the workload's distinguishing characteristics?

    2 min answer weights-biasesml-platformtelemetryingestion
  9. ML Platform advanced

    A model performs well in evaluation and poorly in production. What are the likely causes?

    2 min answer ml-platformtraining-serving-skewdriftfeatures
  10. ML Platform advanced

    A recommendation platform's ML systems span data pipelines, training, evaluation and online serving. Where should the boundaries be, and what causes the most costly class of bug?

    2 min answer ml-platformfeature-storetraining-serving-skewboundaries
  11. ML Platform advanced

    A training run on thousands of GPUs loses several nodes mid-run. How should checkpointing frequency, elastic training, straggler detection and scheduling minimise wasted compute, and what does each checkpoint cost?

    3 min answer metallamatrainingcheckpointing
  12. ML Platform advanced

    NVIDIA's NCCL implements the same all-reduce with several algorithms - ring, tree, and switch-assisted variants over NVLink and NVSwitch - and chooses between them at runtime. Why does one collective operation need several algorithms, and what does that force a cluster scheduler to know?

    3 min answer nvidiancclcollective-communicationscheduling