1. Graceful Degradation advanced

    A payments platform's fraud service is degraded during a peak commerce event. Should the payment path wait, skip the check, use cached risk scores, fall back to simpler rules, or fail? What should determine the answer?

    2 min answer phonepedegradationfraudfallback
  2. Incident Management advanced

    A 43-second network blip triggers automated database failover across regions. Service is degraded for over 24 hours. Analyse.

    2 min answer githubsplit-brainfailoverreconciliation
  3. Incident Management advanced

    A configuration change disconnects a company's internal network. Engineers cannot access the tooling needed to revert it. Analyse.

    2 min answer metacircular-dependencyrecoveryout-of-band
  4. Incident Management advanced

    A post-mortem finds the telemetry system depended on the same infrastructure that failed, leaving engineers blind. What must be isolated, and what is the minimum set of break-glass signals?

    3 min answer observabilityblast-radiusbreak-glassincident
  5. Incident Management advanced

    Incidents at your company are chaotic: unclear ownership, no communication, and postmortems that produce nothing. Design the improvement.

    2 min answer incidentrolesseveritypostmortem
  6. Postmortems advanced

    An engineer runs a routine capacity-removal command with a typo. It removes far more than intended and a core service is down for hours. What does the postmortem conclude?

    2 min answer awsblamelesstoolingblast-radius
  7. Postmortems advanced

    An organisation writes thorough postmortems and keeps experiencing the same class of failure. What is missing?

    2 min answer growwpostmortemaction-itemssystemic
  8. Redundancy advanced

    A platform runs three redundant instances of a component and calculates its availability as extremely high. In practice, all three fail together during incidents. What is the flaw in the reasoning?

    2 min answer redundancycorrelated-failureindependenceconfiguration
  9. Redundancy advanced

    A platform runs three replicas of every service across three availability zones and still experiences total outages. What kinds of failure does that redundancy not address?

    2 min answer zeptoredundancycorrelated-failureblast-radius
  10. Capacity Planning advanced

    A grocery delivery platform experiences a sudden multi-week increase in demand well beyond any forecast. Which capacity constraints bind first, and which cannot be solved with autoscaling?

    2 min answer capacity-planningdemand-shockmarketplacesupply
  11. Capacity Planning advanced

    How should a platform with extreme scheduled peaks decide how much capacity headroom to hold, and what evidence should drive it?

    2 min answer dream11capacityheadroomforecasting
  12. Capacity Planning advanced

    Your platform runs across three availability zones. How would you determine whether it actually survives losing one?

    2 min answer capacityzone-failureheadroomtesting