1. Alert Fatigue intermediate

    A team has 340 alerts. On-call receives roughly 40 pages a week and most are ignored. What would you change?

    2 min answer alertingfatigueslooncall
  2. Alert Fatigue intermediate Multiple choice

    A team receives hundreds of alerts a day and misses a genuine outage. What is the fastest structural fix, and what should replace the current approach?

    2 min answer olaalertingfatiguesymptoms
  3. Alert Fatigue intermediate

    An on-call engineer receives so many alerts that they routinely acknowledge without reading. What has failed, and what is the risk beyond tiredness?

    2 min answer alert-fatiguesignal-to-noisereliabilityinstacart
  4. Alert Fatigue intermediate

    An on-call team receives 60 alerts per shift and has stopped reading most of them. What is the structural fix, and what should the alerting rules actually be based on?

    3 min answer alertingoncallslosymptom-based
  5. Alert Fatigue intermediate

    Half your pages result in no action being taken. How do you fix that without reducing coverage?

    2 min answer alertingsymptom-basedsloburn-rate
  6. Alerting intermediate

    A platform's alerts are defined per component - CPU, memory, queue depth, replica lag. During an incident, forty alerts fire simultaneously. What should the alerting strategy be instead?

    2 min answer alertingsymptomscausesslo-based
  7. Alerting intermediate

    What is the test for whether an alert should exist, and what does applying it honestly do?

    2 min answer alertingactionabilitysymptomsburn-rate
  8. Application Performance Monitoring intermediate Multiple choice

    When does APM tell you something that distributed tracing and metrics cannot?

    2 min answer apmtracingprofilingtooling
  9. Business Metrics intermediate

    A marketplace's technical dashboards are all green during a checkout failure that costs significant revenue. What kind of monitoring was missing?

    2 min answer meeshobusiness-metricsdetectionfunnels
  10. Business Metrics intermediate

    A marketplace's technical metrics are all healthy during an incident in which sellers cannot list items. Why did technical monitoring miss it, and what should be measured?

    2 min answer business-metricsdetectionsilent-failureetsy
  11. Correlation IDs intermediate

    A logistics platform's request crosses synchronous services, message queues, scheduled batch jobs and third-party callbacks. Tracing works within services and breaks between them. What is missing?

    2 min answer delhiverycorrelationasynctracing
  12. Correlation IDs intermediate

    An API platform's requests trigger asynchronous work, webhook deliveries and retries, sometimes hours later. How should correlation identifiers be designed so a customer question can be answered end to end?

    2 min answer correlation-idscausalityasyncsupport