1. AI Risk Tiering advanced

    How should AI systems be risk-tiered so that governance is proportionate?

    1 min answer ai-governancerisk-tieringproportionalityautonomy
  2. AI Risk Tiering advanced

    The business wants to deploy a model that ranks loan applications, with a credit officer making the final decision. What must the architecture provide, and what will you insist on before go-live?

    2 min answer ai-governancemodel-riskfairnesshuman-oversight
  3. AI Risk Tiering advanced

    Your company is deploying twelve AI features. Compliance wants one governance process for all of them. What do you propose?

    2 min answer ai-governanceriskprocess
  4. Alert Fatigue intermediate

    A team has 340 alerts. On-call receives roughly 40 pages a week and most are ignored. What would you change?

    2 min answer alertingfatigueslooncall
  5. Alert Fatigue intermediate Multiple choice

    A team receives hundreds of alerts a day and misses a genuine outage. What is the fastest structural fix, and what should replace the current approach?

    2 min answer olaalertingfatiguesymptoms
  6. Alert Fatigue intermediate

    An on-call engineer receives so many alerts that they routinely acknowledge without reading. What has failed, and what is the risk beyond tiredness?

    2 min answer alert-fatiguesignal-to-noisereliabilityinstacart
  7. Alert Fatigue intermediate

    An on-call team receives 60 alerts per shift and has stopped reading most of them. What is the structural fix, and what should the alerting rules actually be based on?

    3 min answer alertingoncallslosymptom-based
  8. Alert Fatigue intermediate

    Half your pages result in no action being taken. How do you fix that without reducing coverage?

    2 min answer alertingsymptom-basedsloburn-rate
  9. Alert Fatigue advanced

    Your team receives 200 pages a week and the on-call rotation has lost two engineers in six months. Fix it.

    2 min answer alertingon-callburn-rateculture
  10. Alerting intermediate

    A platform's alerts are defined per component - CPU, memory, queue depth, replica lag. During an incident, forty alerts fire simultaneously. What should the alerting strategy be instead?

    2 min answer alertingsymptomscausesslo-based
  11. Alerting advanced

    Error rate has been 1.5% for three days. No alert fired. Customers are complaining. What is happening and what failed?

    2 min answer gray-failurealertingoutliersdetection
  12. Alerting intermediate

    What is the test for whether an alert should exist, and what does applying it honestly do?

    2 min answer alertingactionabilitysymptomsburn-rate