1. Incident Management advanced

    A 43-second network blip triggers automated database failover across regions. Service is degraded for over 24 hours. Analyse.

    2 min answer githubsplit-brainfailoverreconciliation
  2. Incident Management advanced

    A configuration change disconnects a company's internal network. Engineers cannot access the tooling needed to revert it. Analyse.

    2 min answer metacircular-dependencyrecoveryout-of-band
  3. Incident Management intermediate

    A payment platform is mid-incident: transactions are partially failing, the cause is unclear, and several teams are investigating. What structure makes the response effective?

    2 min answer razorpayincident-commandrolescommunication
  4. Incident Management advanced

    A post-mortem finds the telemetry system depended on the same infrastructure that failed, leaving engineers blind. What must be isolated, and what is the minimum set of break-glass signals?

    3 min answer observabilityblast-radiusbreak-glassincident
  5. Incident Management intermediate Multiple choice

    An incident is underway and everyone is debugging. What role is missing?

    1 min answer incident-managementrolescoordinationcommunication
  6. Incident Management intermediate

    An observability platform's incident process treats all incidents with the same severity and process. What problems does this create, and how should severity be defined?

    2 min answer incident-managementseverityescalationdatadog
  7. Incident Management advanced

    Incidents at your company are chaotic: unclear ownership, no communication, and postmortems that produce nothing. Design the improvement.

    2 min answer incidentrolesseveritypostmortem
  8. Incident Management intermediate

    You are incident commander. Error rate is 8% and rising, cause unknown. What do you do in the first ten minutes?

    2 min answer incidentcommandmitigationcommunication