1. Capacity Planning advanced

    How should a platform with extreme scheduled peaks decide how much capacity headroom to hold, and what evidence should drive it?

    2 min answer dream11capacityheadroomforecasting
  2. Capacity Planning advanced

    Your platform runs across three availability zones. How would you determine whether it actually survives losing one?

    2 min answer capacityzone-failureheadroomtesting
  3. Reliability Culture advanced

    An airline's crew-scheduling system cannot re-plan fast enough during cascading weather disruption, forcing manual processes that cannot keep up. What signals indicated this risk earlier, and how should such a system be modernised?

    3 min answer southwestlegacytechnical-debtcapacity
  4. Reliability Culture intermediate

    An organisation adopts blameless postmortems but engineers still avoid admitting mistakes and incidents are under-reported. What is missing beyond the stated policy?

    2 min answer cultureblamelesspsychological-safetyincentives
  5. Reliability & Resilience advanced

    A super-app combines messaging, payments, social feeds, mini-programs and notifications in one product used by hundreds of millions daily. What is the dominant reliability risk, and what structural property addresses it?

    2 min answer reliabilityisolationsuper-appblast-radius
  6. Reliability & Resilience intermediate

    A team is launching a new service and asks what its SLO should be. How do you help them decide, and why is "99.99%" usually the wrong first answer?

    2 min answer slosreerror-budgetreliability
  7. Reliability & Resilience advanced

    A workflow executes over days or weeks while individual workers restart many times. How should workflow state, timers, retries, heartbeats and task ownership be designed?

    2 min answer temporaldurable-executionworkflowstimers
  8. Reliability & Resilience advanced

    In July 2019 a single regex deployed globally took Cloudflare's network to 100% CPU within seconds. What does this say about how configuration should be released?

    2 min answer deploymentblast-radiusconfigcase-study
  9. Reliability & Resilience advanced

    Maersk rebuilt roughly 4,000 servers and 45,000 PCs in about ten days after NotPetya in 2017, and recovered its directory only because one data centre had been offline during the attack. What does this say about DR design?

    2 min answer disaster-recoveryransomwarebackupcase-study
  10. Reliability & Resilience intermediate

    The business asks for "100% uptime" for a new customer portal. Walk me through the conversation that ends in an agreed SLO.

    3 min answer slisloslanegotiation
  11. Resilience Testing advanced

    A brief database slowdown caused a two-hour full outage. Explain the likely amplification chain and the fixes at each stage.

    2 min answer cascading-failureretriestimeoutsincident
  12. Resilience Testing advanced

    A travel platform depends on hundreds of external suppliers with varying reliability. How should resilience testing be designed when the failures originate outside the system?

    2 min answer resilience-testingfault-injectionexternal-dependenciesexpedia