1. Metrics intermediate

    Etsy released StatsD in 2011: a small daemon that receives metric samples from applications over UDP, aggregates them in memory, and flushes to the metrics backend on a fixed interval. Why was UDP the right choice for that design, and what did it make permanently impossible?

    3 min answer etsystatsdmetricsudp
  2. Metrics intermediate

    Every service dashboard shows healthy metrics while users cannot complete a purchase. How is that possible and what would have detected it?

    2 min answer metricsbusiness-metricsdetectionblind-spots
  3. Observability intermediate

    A new service goes live next week. What observability must exist on day one?

    2 min answer observabilitylaunchoperations
  4. Observability intermediate

    You are designing a new service. What must be in place before it goes live so that whoever is paged at 3 AM can diagnose it without you?

    2 min answer observabilityoncallalertinglogging
  5. OpenTelemetry intermediate

    A platform with services in five languages and three monitoring vendors considers adopting OpenTelemetry. What does it solve, and what is the realistic migration cost?

    2 min answer opentelemetrystandardisationvendor-lock-inmigration
  6. OpenTelemetry intermediate

    An organisation with several existing telemetry systems considers adopting OpenTelemetry. What does it actually solve, and what does adopting it not fix?

    2 min answer segmentopentelemetrystandardsvendor-lock-in
  7. SLO Monitoring intermediate Multiple choice

    A checkout API has a 99.9% availability SLO and the team must decide where the indicator is computed from. The candidates are load-balancer access logs, in-process server metrics, the mobile client's own reporting, and synthetic probes. Which should be the primary source?

    3 min answer sloslimeasurementavailability
  8. SLO Monitoring intermediate Multiple choice

    A food-delivery platform in Zomato's mould pushes order-status events to restaurant tablets and to customers through a queue-backed webhook fleet. The complaints are that status arrives late rather than that it never arrives. Which indicator should the SLO be written on?

    3 min answer slofreshnessasynchronouswebhooks
  9. Structured Logging intermediate

    A platform moves from free-text logs to structured logs. What becomes possible, and what discipline must accompany it to avoid making things worse?

    2 min answer structured-loggingschemacardinalityquerying
  10. Structured Logging intermediate

    A ride-hailing platform of Grab's shape standardises on structured logs. After a release, the on-call dashboard that counts failed trips by city reads zero for three services and normal numbers for the rest. No errors are being reported anywhere. What broke, and what should have caught it?

    2 min answer grabstructured loggingschema driftlog management
  11. Telemetry Cost intermediate Multiple choice

    A design platform of the kind Canva runs exports 40 metrics per service. An engineer adds a `pod_name` label so a noisy pod can be identified. The service runs 600 pods and deploys twice a day, so pod names turn over completely every 12 hours. Retention is 30 days. Roughly how many distinct series does that one label create over the retention window?

    2 min answer canvacardinalitytelemetry costmetrics
  12. Telemetry Cost intermediate

    A platform's log volume has grown until log storage is one of its largest infrastructure costs. What should change, and what should not?

    2 min answer elasticlogscostindexing