1. Reliability Culture advanced

    An airline's crew-scheduling system cannot re-plan fast enough during cascading weather disruption, forcing manual processes that cannot keep up. What signals indicated this risk earlier, and how should such a system be modernised?

    3 min answer southwestlegacytechnical-debtcapacity
  2. Reliability & Resilience advanced

    A super-app combines messaging, payments, social feeds, mini-programs and notifications in one product used by hundreds of millions daily. What is the dominant reliability risk, and what structural property addresses it?

    2 min answer reliabilityisolationsuper-appblast-radius
  3. Reliability & Resilience advanced

    A workflow executes over days or weeks while individual workers restart many times. How should workflow state, timers, retries, heartbeats and task ownership be designed?

    2 min answer temporaldurable-executionworkflowstimers
  4. Reliability & Resilience advanced

    In July 2019 a single regex deployed globally took Cloudflare's network to 100% CPU within seconds. What does this say about how configuration should be released?

    2 min answer deploymentblast-radiusconfigcase-study
  5. Reliability & Resilience advanced

    Maersk rebuilt roughly 4,000 servers and 45,000 PCs in about ten days after NotPetya in 2017, and recovered its directory only because one data centre had been offline during the attack. What does this say about DR design?

    2 min answer disaster-recoveryransomwarebackupcase-study
  6. Resilience Testing advanced

    A brief database slowdown caused a two-hour full outage. Explain the likely amplification chain and the fixes at each stage.

    2 min answer cascading-failureretriestimeoutsincident
  7. Resilience Testing advanced

    A travel platform depends on hundreds of external suppliers with varying reliability. How should resilience testing be designed when the failures originate outside the system?

    2 min answer resilience-testingfault-injectionexternal-dependenciesexpedia
  8. Resilience Testing advanced

    How would you test that your service degrades correctly when a dependency's latency rises tenfold?

    2 min answer resilience-testinglatency-injectiontimeoutslittle's-law
  9. Resilience Testing advanced

    Review this resilience programme. Chaos experiments run weekly in staging at 03:00, they inject only instance termination, results are recorded in a spreadsheet, and there is a kill switch that has never been used. What would you change, and what would you keep?

    3 min answer chaos engineeringgame daysstagingsteady state
  10. RTO & RPO advanced

    A collaborative workspace product must define RTO and RPO. The product team says "we can never lose a user's work". What does that requirement actually mean, and what does it cost?

    2 min answer rportodurabilitycollaboration
  11. RTO & RPO advanced

    A stakeholder asks for zero RPO across the estate. What does that cost and what would you propose instead?

    2 min answer rporeplicationcostbusiness-continuity
  12. RTO & RPO advanced

    In April 2022 a maintenance script at Atlassian used the wrong identifiers and deleted sites belonging to about 775 customers. Restoration took up to about two weeks for some of them, even though backups existed and were working. What property of the recovery design accounts for that gap?

    3 min answer atlassianrestore granularitymulti-tenantblast radius