1. AI Gateways advanced

    An AI gateway now terminates streaming responses for six products and holds a semantic cache shared across tenants. Two properties of that design have caused real outages and real data exposure at large AI providers. What has the organisation taken on and when does the bill arrive?

    3 min answer ai-gatewayopenaiblast-radiussemantic-cache
  2. AI Observability advanced

    A fraud model has been in production for eight months. Ground truth arrives weeks later. How do you know if it is still working?

    2 min answer mlmonitoringdrift
  3. AI Observability advanced

    Anthropic reported in 2025 that a routing bug sent a share of Claude Sonnet 4 requests to servers configured for a different context length, peaking at 16% of those requests in the worst hour on 31 August, and that it took weeks to identify. Error rates never moved. Which design decision allowed the delay and what telemetry closes it?

    3 min answer anthropicpostmortemsilent-regressionrouting
  4. AI Observability advanced

    What must be observable in an AI feature that is not covered by conventional application monitoring?

    2 min answer ai-observabilitytracingqualitycost
  5. AI Observability advanced

    You are asked to log every prompt and response for an assistant handling 2 million interactions a day, so quality regressions are diagnosable. Roughly how much data does that commit you to per year, and does the number change the design?

    3 min answer observabilityloggingretentioncost
  6. AI Risk Tiering advanced

    An enterprise AI provider must apply governance proportionate to risk across many model deployments. How should tiering work?

    2 min answer coherescale-airisk-tieringgovernance
  7. AI Risk Tiering advanced

    An organisation deploys many AI systems with very different risk profiles. How should they be tiered, and which controls attach to each tier?

    2 min answer ai-governancerisk-tieringmodel-riskhuman-in-the-loop
  8. AI Risk Tiering advanced

    How should AI systems be risk-tiered so that governance is proportionate?

    1 min answer ai-governancerisk-tieringproportionalityautonomy
  9. AI Risk Tiering advanced

    The business wants to deploy a model that ranks loan applications, with a credit officer making the final decision. What must the architecture provide, and what will you insist on before go-live?

    2 min answer ai-governancemodel-riskfairnesshuman-oversight
  10. AI Risk Tiering advanced

    Your company is deploying twelve AI features. Compliance wants one governance process for all of them. What do you propose?

    2 min answer ai-governanceriskprocess
  11. Alert Fatigue advanced

    Your team receives 200 pages a week and the on-call rotation has lost two engineers in six months. Fix it.

    2 min answer alertingon-callburn-rateculture
  12. Alerting advanced

    Error rate has been 1.5% for three days. No alert fired. Customers are complaining. What is happening and what failed?

    2 min answer gray-failurealertingoutliersdetection