1. Failure Modes advanced

    One instance in a fleet of fifty is returning correct responses very slowly. Health checks pass and it stays in rotation. How do you detect and handle this?

    2 min answer grey-failurehealth-checksdetectionmitigation
  2. Leader Election advanced

    A nightly job occasionally runs twice, producing duplicate charges. The team proposes a distributed lock. What do you say?

    2 min answer leader-electionlockingidempotencyfencing
  3. Leader Election advanced

    Your cluster fails over spuriously under load, but a real leader failure takes 45 seconds to detect. How do you resolve the tension?

    2 min answer leader-electionfailoverdetectiontuning
  4. Load Shedding advanced

    One customer's batch job saturates a shared service and degrades everyone. Rate limiting them fixes it, until the next customer does the same. What is the structural answer?

    2 min answer multi-tenancyisolationadmission-controlfairness
  5. Load Shedding advanced

    Your service will exceed capacity by 30% during a known peak. Do you shed load or brown out, and how do you decide what goes first?

    2 min answer sheddingbrownoutdegradationprioritisation
  6. Sagas & Compensation advanced

    In an order saga, the shipment service permanently rejects an order after the card has been captured. What now?

    2 min answer sagacompensationpivotexceptions
  7. Service Discovery advanced

    Your service registry becomes unavailable. Every service is healthy. What happens, and what should happen?

    2 min answer discoverystatic-stabilityavailabilitycaching
  8. Timeouts & Deadlines advanced

    Design the timeout configuration for a request that passes through gateway, orders, pricing and inventory. What numbers, and what rule generates them?

    2 min answer timeoutsdeadlinescascadingretries