1. DR Testing intermediate

    An enterprise runs an annual disaster recovery test that always succeeds, yet the team has low confidence in real recovery. What is likely wrong with the test?

    2 min answer dr-testingdrillsrealismenterprise
  2. DR Testing advanced

    In January 2017 GitLab lost roughly six hours of database data after an engineer deleted a directory on the wrong host during replication troubleshooting, and then found that several backup and replication mechanisms had silently not been working. What signal would have revealed the broken backups beforehand, and why did nothing report them?

    3 min answer gitlabbackupsrestore testingsilent failure
  3. DR Testing advanced

    Your DR plan is pilot light with a 30-minute RTO. What would you test, and what will the test probably reveal?

    2 min answer drtestingcontrol-plane
  4. DR Testing advanced

    Zoom's daily meeting participants rose from roughly 10 million in December 2019 to about 300 million by April 2020. Your platform has a plausible 10x surge ahead and a disaster-recovery plan last tested by tabletop review two years ago. Sequence the move from tabletop to a genuinely tested recovery capability, without an outage.

    4 min answer zoomdisaster recoverydr testingsurge
  5. Distributed Locking advanced

    A media platform uses a distributed lock to stop two workers rendering the same expensive export. Workers sometimes crash while holding the lock, and occasionally two workers process the same job anyway. Diagnose both problems and redesign.

    2 min answer distributed-lockingleasesfencingidempotency
  6. Distributed Locking advanced

    A team proposes a distributed lock to stop two workers processing the same transaction. Workers sometimes crash or pause while holding the lock. What is the safer design, and when is a lock genuinely necessary?

    2 min answer jupiterdistributed-lockfencingidempotency
  7. Distributed Locking advanced Multiple choice

    A team uses a distributed lock in a key-value store to ensure only one worker processes a job. Occasionally two workers process the same job. Explain.

    2 min answer lockingcorrectnessfencing
  8. Distributed Locking advanced

    Three designs need distributed locks: a nightly report, a per-customer state machine, and a global config reload. For each, is a lock the right answer?

    2 min answer lockingpartitioningidempotencydesign
  9. Distributed Systems advanced

    A 43-second network partition caused GitHub over 24 hours of degraded service in 2018. How does a 43-second event become a day-long incident?

    2 min answer failoversplit-brainconsistencycase-study
  10. Distributed Systems advanced Multiple choice

    A card payment authorisation service runs active-active across two regions. A network partition splits them. Do you keep accepting authorisations, and what breaks either way?

    2 min answer capconsistencypaymentsavailability
  11. Distributed Systems advanced

    A downstream service slows from 50 ms to 3 s. Within two minutes every service in the request path is down, including ones that do not call it. Explain the mechanism and how you would have prevented it.

    2 min answer cascading-failureretriestimeoutsresilience
  12. Distributed Systems advanced

    A payments platform sees a 20x increase in transaction attempts during a major commerce event. Idempotency, rate limiting, queueing, fraud checks, database contention and a downstream provider all interact. What is the correct ordering of these controls on the request path, and why does ordering matter more than any single control?

    2 min answer razorpaypaymentssurgeidempotency