1. Resilience Testing advanced

    How would you test that your service degrades correctly when a dependency's latency rises tenfold?

    2 min answer resilience-testinglatency-injectiontimeoutslittle's-law
  2. Resilience Testing advanced

    Review this resilience programme. Chaos experiments run weekly in staging at 03:00, they inject only instance termination, results are recorded in a spreadsheet, and there is a kill switch that has never been used. What would you change, and what would you keep?

    3 min answer chaos engineeringgame daysstagingsteady state
  3. RTO & RPO advanced

    A collaborative workspace product must define RTO and RPO. The product team says "we can never lose a user's work". What does that requirement actually mean, and what does it cost?

    2 min answer rportodurabilitycollaboration
  4. RTO & RPO intermediate

    A financial-operations platform sets one RTO and RPO for the entire business. What goes wrong, and how should they be derived instead?

    2 min answer ramprtorpodr
  5. RTO & RPO advanced

    A stakeholder asks for zero RPO across the estate. What does that cost and what would you propose instead?

    2 min answer rporeplicationcostbusiness-continuity
  6. RTO & RPO beginner Multiple choice

    A team documents an RPO of 24 hours because backups run nightly at 02:00 and take about 40 minutes. The last restore was attempted at the time the system was built. What is the honest recovery point objective?

    3 min answer rpobackupsrestoreverification
  7. RTO & RPO advanced

    In April 2022 a maintenance script at Atlassian used the wrong identifiers and deleted sites belonging to about 775 customers. Restoration took up to about two weeks for some of them, even though backups existed and were working. What property of the recovery design accounts for that gap?

    3 min answer atlassianrestore granularitymulti-tenantblast radius
  8. SLI, SLO & SLA intermediate Multiple choice

    A brokerage sets a single availability target of 99.9% for its whole platform. Why is that wrong, and what should replace it?

    2 min answer zerodhaslojourneysavailability
  9. SLI, SLO & SLA advanced

    A live-streaming platform defines an SLO of "99.9% of API requests succeed". During a major event, the SLO is met while viewers report the product is unusable. What is wrong with the SLO?

    2 min answer slosliuser-journeysaggregation
  10. SLI, SLO & SLA advanced

    A service exhausts its error budget in week two of a 28-day window. What actually happens next?

    2 min answer sloerror-budgetsregoogle
  11. SLI, SLO & SLA advanced

    Leadership asks for "five nines" across the platform. Engineering says it is impossible. Design the response.

    2 min answer sloerror-budgetstakeholdersavailability
  12. Static Stability advanced Multiple choice

    A marketplace's failover mechanism requires calling a control plane to provision replacement capacity. Why is this a design flaw, and what is the alternative?

    2 min answer static-stabilitycontrol-planefailovercapacity