1. Leader Election advanced

    Your cluster fails over spuriously under load, but a real leader failure takes 45 seconds to detect. How do you resolve the tension?

    2 min answer leader-electionfailoverdetectiontuning
  2. Load Shedding advanced

    A fantasy-sports platform faces a synchronised traffic spike in the minutes before a major match starts. Workers are saturated. What should be shed, in what order, and what must never be shed?

    2 min answer dream11load-sheddingprioritisationburst
  3. Load Shedding advanced Multiple choice

    An AI inference service receives computationally expensive requests with unpredictable bursts, where a single request can occupy a GPU for many seconds. Which combination of controls should it use, and in what order do they apply?

    2 min answer load-sheddingadmission-controlbatchingopenai
  4. Load Shedding advanced

    One customer's batch job saturates a shared service and degrades everyone. Rate limiting them fixes it, until the next customer does the same. What is the structural answer?

    2 min answer multi-tenancyisolationadmission-controlfairness
  5. Load Shedding advanced

    Your service is at capacity and must reject some traffic. What do you shed, and how do you decide?

    2 min answer overloadprioritisationresilience
  6. Load Shedding advanced

    Your service will exceed capacity by 30% during a known peak. Do you shed load or brown out, and how do you decide what goes first?

    2 min answer sheddingbrownoutdegradationprioritisation
  7. Messaging & Queues advanced

    A flash sale causes an incident. Two hours later the fault is fixed but 4 million jobs are backed up and notifications are hours late. Walk through recovery and prevention.

    3 min answer queuesbackpressureshopifyincident
  8. Retries & Backoff advanced

    A dependency degrades to 2-second latency rather than failing outright, and retries triple the load on it. Why is slow worse than down, and which mechanisms actually prevent the amplification?

    2 min answer retriesdeadlinesadaptive-concurrencyload-shedding
  9. Retries & Backoff advanced

    A dependency degrades. Your service retries three times with exponential backoff. The dependency never recovers until you deploy a change. Why?

    2 min answer retriesoverloadjitter
  10. Retries & Backoff advanced

    During a regional degradation, a mobility platform's services perform millions of retries against a slow dependency, and the dependency never recovers until traffic is manually cut. Analyse the failure and identify which four controls would have prevented it.

    2 min answer retriesretry-stormmetastable-failuregrab
  11. Sagas & Compensation advanced

    A long-running customer workflow spanning payment, verification, notification and fulfilment must survive worker crashes and deploys mid-flow. Compare a hand-written saga with a durable-execution engine.

    2 min answer temporalsagadurable-executionworkflow
  12. Sagas & Compensation advanced

    A marketplace charges the customer, pays the merchant, pays the courier and takes a fee. Design the transaction, and say what happens if the courier payment fails.

    2 min answer marketplacesagapaymentscompensation