1. Availability Mathematics advanced

    A stakeholder wants a video streaming service at 99.99% availability. Walk through whether that is the right target and what it would take.

    2 min answer netflixslocostrequirements
  2. Chaos Engineering advanced

    A dependency fails only under high load, so normal testing never exposes it. How should load testing, fault injection, capacity experiments and traffic shadowing be combined to find it?

    2 min answer swiggychaosload-testingfault-injection
  3. Chaos Engineering advanced

    Leadership read about Chaos Monkey and wants chaos engineering in production next month. Your last three incidents were caused by known unfixed reliability issues. What do you say?

    2 min answer netflixchaosprerequisitesjudgement
  4. Chaos Engineering advanced

    What is the first chaos experiment you would run on a system you have just inherited, and what must be in place first?

    2 min answer chaoslatency-injectionnetflixprerequisites
  5. Degradation Modes advanced

    A crypto aggregator's upstream exchange feeds become unreliable during extreme market volatility - the moment users most need them. How should stale data, provider failover, breakers, caching and user-facing warnings interact?

    2 min answer coinswitchmarket-datastalenessfailover
  6. Degradation Modes advanced

    An e-commerce platform designs explicit degradation modes for extreme demand events. What must be true of the mode ladder for it to work on the day?

    2 min answer degradationload-sheddingflash-salerehearsal
  7. Degradation Modes advanced

    Define three degradation modes for a real-time marketplace. What triggers each, what stops working, and how does the system return to normal?

    2 min answer degradation-modesuberprioritisationload-shedding
  8. DR Testing advanced

    An auditor requires evidence that a four-hour RTO is achievable for a system that has never failed over. Production cannot be risked and the business will not accept an unplanned outage. Sequence the first real disaster-recovery test.

    3 min answer dr testingrtorehearsalevidence
  9. DR Testing advanced

    In January 2017 GitLab lost roughly six hours of database data after an engineer deleted a directory on the wrong host during replication troubleshooting, and then found that several backup and replication mechanisms had silently not been working. What signal would have revealed the broken backups beforehand, and why did nothing report them?

    3 min answer gitlabbackupsrestore testingsilent failure
  10. DR Testing advanced

    Your DR plan is pilot light with a 30-minute RTO. What would you test, and what will the test probably reveal?

    2 min answer drtestingcontrol-plane
  11. Error Budgets advanced

    A platform has an error budget policy stating that feature work stops when the budget is exhausted. The budget is exhausted, and product leadership wants a major launch to proceed. How should this be resolved?

    2 min answer error-budgetspolicygovernancereliability
  12. Error Budgets advanced

    A team consistently ends every window with 90% of its error budget unspent. What does that tell you?

    2 min answer error-budgetsloover-investmentrisk