1. OpenTelemetry advanced

    A platform team standardises every service on OTLP push to a collector and switches off Prometheus scraping. Metrics still arrive and dashboards still work. What has the team given up, and when does that bill arrive?

    3 min answer opentelemetryprometheuspushscrape
  2. OpenTelemetry intermediate

    A platform with services in five languages and three monitoring vendors considers adopting OpenTelemetry. What does it solve, and what is the realistic migration cost?

    2 min answer opentelemetrystandardisationvendor-lock-inmigration
  3. OpenTelemetry intermediate

    An organisation with several existing telemetry systems considers adopting OpenTelemetry. What does it actually solve, and what does adopting it not fix?

    2 min answer segmentopentelemetrystandardsvendor-lock-in
  4. Profiling advanced

    A data platform has good tracing and still cannot explain why a specific job is slow. What does tracing not tell you, and what does?

    2 min answer databricksprofilingtracingcpu
  5. Profiling advanced Multiple choice

    A developer-tools company needs to find a performance regression that only appears under real production workloads. What are the options for profiling in production, and what are their costs?

    2 min answer profilingcontinuous-profilingsamplingoverhead
  6. Profiling advanced

    A service is slow and CPU utilisation is 4%. What do you profile and what do you expect to find?

    2 min answer profilingwall-clockblockingio
  7. Profiling advanced

    Discord's 2020 post on rewriting its Read States service from Go to Rust described latency spikes on a roughly two-minute cadence, matching Go's forced garbage-collection interval, in a service that allocated very little. An engineer brings you a similar graph today and asks you to fund always-on profiling across 12000 containers. Walk me through what you would fund and what you would refuse.

    3 min answer discordprofilinggarbage-collectiontail-latency
  8. SLO Monitoring advanced

    A canary release looks healthy on p50 latency but a small set of enterprise tenants sees timeouts. Which metrics and gates should have caught it?

    2 min answer canarytenant-segmentationtail-latencyrollout-gates
  9. SLO Monitoring intermediate Multiple choice

    A checkout API has a 99.9% availability SLO and the team must decide where the indicator is computed from. The candidates are load-balancer access logs, in-process server metrics, the mobile client's own reporting, and synthetic probes. Which should be the primary source?

    3 min answer sloslimeasurementavailability
  10. SLO Monitoring advanced

    A communication platform sets a 99.9% availability SLO. How should alerting on that SLO be structured so it catches both sudden outages and slow degradation?

    2 min answer sloburn-ratemulti-windowalerting
  11. SLO Monitoring intermediate Multiple choice

    A food-delivery platform in Zomato's mould pushes order-status events to restaurant tablets and to customers through a queue-backed webhook fleet. The complaints are that status arrives late rather than that it never arrives. Which indicator should the SLO be written on?

    3 min answer slofreshnessasynchronouswebhooks
  12. Structured Logging intermediate

    A platform moves from free-text logs to structured logs. What becomes possible, and what discipline must accompany it to avoid making things worse?

    2 min answer structured-loggingschemacardinalityquerying