1. Alert Fatigue advanced

    Your team receives 200 pages a week and the on-call rotation has lost two engineers in six months. Fix it.

    2 min answer alertingon-callburn-rateculture
  2. Alerting advanced

    Error rate has been 1.5% for three days. No alert fired. Customers are complaining. What is happening and what failed?

    2 min answer gray-failurealertingoutliersdetection
  3. Application Performance Monitoring advanced

    In a multi-tenant platform, aggregate application performance metrics look healthy while specific tenants experience severe slowness. What must performance monitoring do differently?

    2 min answer apmmulti-tenancysegmentationoutliers
  4. Application Performance Monitoring advanced

    Review this instrumentation. A checkout service emits about 400 spans per request, every span carries roughly 30 attributes including the full request body, tracing runs at 100% with no sampling, and the team reports that they cannot find anything in the traces. What would you remove, what would you change, and what would you keep even though it looks excessive?

    3 min answer tracingapminstrumentationsampling
  5. Business Metrics advanced

    Customers reported an outage 40 minutes before your monitoring did. How do you close that gap?

    2 min answer detectionbusiness-metricssyntheticsclient-side
  6. Cardinality advanced Multiple choice

    A live-streaming platform of Twitch's shape needs usable latency quantiles for 300 API endpoints whose responses span 2 ms to 90 seconds. The current histograms use 12 fixed buckets topping out at 10 seconds and every p99 above that reads as the overflow bucket. Which change fits the problem?

    3 min answer cardinalityhistogramsopentelemetryquantiles
  7. Cardinality advanced

    A single deploy took down your monitoring platform. What happened, and how do you prevent a recurrence?

    2 min answer cardinalitymetricscostguardrails
  8. Cardinality advanced

    A team wants to debug failures nobody predicted, using wide high-cardinality events rather than pre-aggregated metrics. How does the storage design differ, and how must sampling preserve rare errors?

    3 min answer honeycombwide-eventscardinalitysampling
  9. Cardinality advanced

    An observability platform ingests billions of telemetry events daily, and some customers attach dimensions with unbounded distinct values. How should ingestion, storage, indexing and query be designed so cost and latency stay manageable?

    2 min answer cardinalitytime-seriesingestionmulti-tenancy
  10. Cardinality advanced

    An observability platform ingests enormous telemetry volume, but a small number of dimensions create extreme cardinality. How should ingestion, aggregation, indexing, sampling, retention and storage tiers be designed so cost and query performance stay predictable?

    2 min answer grafanacardinalitytelemetry-costretention
  11. Debugging Distributed Systems advanced

    A collaborative editing platform receives a report that one document occasionally shows different content to different users, but it cannot be reproduced. How do you investigate a rare, non-deterministic, distributed bug?

    2 min answer debuggingnon-determinismdistributedforensics
  12. Debugging Distributed Systems advanced

    A platform of 40 services has logs only, and incidents take hours to diagnose. Design the observability strategy and its rollout order.

    2 min answer observabilitytracingmetricsrollout