1. On-Call intermediate

    A team is paged eleven times a week and morale is poor. What do you do first?

    3 min answer oncallalert-fatiguesustainabilityarchitecture
  2. On-Call intermediate

    An on-call rotation is producing burnout and slow responses. Alert volume is high and most pages are not actionable. What changes, and in what order?

    2 min answer sliceoncallalertingactionability
  3. On-Call intermediate

    You join a team of 12 engineers taking around 40 pages a week, of which perhaps 5 required action. The team is exhausted and two people have resigned. The director asks for a plan. What do you do, and in what order?

    3 min answer on-callalert fatiguetoilsustainability
  4. Postmortems intermediate

    A platform publishes detailed public postmortems for significant incidents. What does this practice cost, and what does it buy that internal postmortems do not?

    2 min answer postmortemstransparencylearningtrust
  5. Postmortems intermediate

    An engineer ran a command that deleted production data. What does the postmortem investigate?

    2 min answer postmortemblamelesssystemicgoogle
  6. Postmortems advanced

    An engineer runs a routine capacity-removal command with a typo. It removes far more than intended and a core service is down for hours. What does the postmortem conclude?

    2 min answer awsblamelesstoolingblast-radius
  7. Postmortems advanced

    An organisation writes thorough postmortems and keeps experiencing the same class of failure. What is missing?

    2 min answer growwpostmortemaction-itemssystemic
  8. Redundancy advanced

    A platform runs three redundant instances of a component and calculates its availability as extremely high. In practice, all three fail together during incidents. What is the flaw in the reasoning?

    2 min answer redundancycorrelated-failureindependenceconfiguration
  9. Redundancy advanced

    A platform runs three replicas of every service across three availability zones and still experiences total outages. What kinds of failure does that redundancy not address?

    2 min answer zeptoredundancycorrelated-failureblast-radius
  10. Redundancy intermediate Multiple choice

    Two service instances each run at 60% CPU behind a load balancer. Is that redundant?

    2 min answer redundancycapacitycorrelationfailure-domains
  11. Capacity Planning advanced

    A grocery delivery platform experiences a sudden multi-week increase in demand well beyond any forecast. Which capacity constraints bind first, and which cannot be solved with autoscaling?

    2 min answer capacity-planningdemand-shockmarketplacesupply
  12. Capacity Planning intermediate

    A service runs across three availability zones at 70% CPU during peak. The team wants to survive losing one zone at peak with no user impact. Roughly what does that require, and what does the same arithmetic say about running across two zones?

    3 min answer capacityheadroomavailability zonesstatic stability