1. Alert Fatigue intermediate

    A team has 340 alerts. On-call receives roughly 40 pages a week and most are ignored. What would you change?

    2 min answer alertingfatigueslooncall
  2. Alert Fatigue intermediate Multiple choice

    A team receives hundreds of alerts a day and misses a genuine outage. What is the fastest structural fix, and what should replace the current approach?

    2 min answer olaalertingfatiguesymptoms
  3. Alert Fatigue intermediate

    An on-call engineer receives so many alerts that they routinely acknowledge without reading. What has failed, and what is the risk beyond tiredness?

    2 min answer alert-fatiguesignal-to-noisereliabilityinstacart
  4. Alert Fatigue intermediate

    An on-call team receives 60 alerts per shift and has stopped reading most of them. What is the structural fix, and what should the alerting rules actually be based on?

    3 min answer alertingoncallslosymptom-based
  5. Alert Fatigue intermediate

    Half your pages result in no action being taken. How do you fix that without reducing coverage?

    2 min answer alertingsymptom-basedsloburn-rate
  6. Alert Fatigue advanced

    Your team receives 200 pages a week and the on-call rotation has lost two engineers in six months. Fix it.

    2 min answer alertingon-callburn-rateculture
  7. Alerting intermediate

    A platform's alerts are defined per component - CPU, memory, queue depth, replica lag. During an incident, forty alerts fire simultaneously. What should the alerting strategy be instead?

    2 min answer alertingsymptomscausesslo-based
  8. Alerting advanced

    Error rate has been 1.5% for three days. No alert fired. Customers are complaining. What is happening and what failed?

    2 min answer gray-failurealertingoutliersdetection
  9. Alerting intermediate

    What is the test for whether an alert should exist, and what does applying it honestly do?

    2 min answer alertingactionabilitysymptomsburn-rate
  10. Alerting beginner

    Your team's on-call phone has not been paged in nine days and every dashboard is green. Why is that not yet evidence that anything is healthy, and what single mechanism turns silence into a signal?

    2 min answer alertingdead-mans-switchwatchdogheartbeat
  11. Application Performance Monitoring advanced

    In a multi-tenant platform, aggregate application performance metrics look healthy while specific tenants experience severe slowness. What must performance monitoring do differently?

    2 min answer apmmulti-tenancysegmentationoutliers
  12. Application Performance Monitoring advanced

    Review this instrumentation. A checkout service emits about 400 spans per request, every span carries roughly 30 attributes including the full request body, tracing runs at 100% with no sampling, and the team reports that they cannot find anything in the traces. What would you remove, what would you change, and what would you keep even though it looks excessive?

    3 min answer tracingapminstrumentationsampling