1. Observability advanced

    An error-tracking platform receives millions of events, many of them the same underlying problem with different stack details. How should grouping, deduplication, high-cardinality metadata and retention be designed?

    2 min answer sentrygroupingdeduplicationretention
  2. Observability advanced

    On 11 December 2024 OpenAI rolled a new telemetry service out across every Kubernetes cluster. Within about half an hour the API servers were saturated and services could no longer resolve one another; full recovery took until the evening. What turned a monitoring change into a total outage, and which of the contributing factors would you fix first?

    3 min answer openaikubernetescontrol planedns
  3. Observability advanced

    Your observability bill is now 40% of your compute bill. Leadership wants it cut in half without going blind. What do you cut?

    2 min answer observabilitycostcardinalitysampling
  4. Observability advanced

    p99 latency on checkout tripled overnight. Dashboards look normal, no deployment went out, and every service reports healthy. How do you find it?

    3 min answer debuggingtracinglatencydiagnosis
  5. OpenTelemetry advanced

    A fleet exports OTLP to a gateway collector deployment with memory_limiter first in the pipeline, then a batch processor, then an exporter with a sending queue. The telemetry backend starts answering in 8 seconds instead of 80 ms, and request volume doubles at the same time. What happens second by second, and what stops it?

    2 min answer opentelemetrycollectorbackpressureload-shedding
  6. OpenTelemetry advanced

    A platform team standardises every service on OTLP push to a collector and switches off Prometheus scraping. Metrics still arrive and dashboards still work. What has the team given up, and when does that bill arrive?

    3 min answer opentelemetryprometheuspushscrape
  7. Profiling advanced

    A data platform has good tracing and still cannot explain why a specific job is slow. What does tracing not tell you, and what does?

    2 min answer databricksprofilingtracingcpu
  8. Profiling advanced Multiple choice

    A developer-tools company needs to find a performance regression that only appears under real production workloads. What are the options for profiling in production, and what are their costs?

    2 min answer profilingcontinuous-profilingsamplingoverhead
  9. Profiling advanced

    A service is slow and CPU utilisation is 4%. What do you profile and what do you expect to find?

    2 min answer profilingwall-clockblockingio
  10. Profiling advanced

    Discord's 2020 post on rewriting its Read States service from Go to Rust described latency spikes on a roughly two-minute cadence, matching Go's forced garbage-collection interval, in a service that allocated very little. An engineer brings you a similar graph today and asks you to fund always-on profiling across 12000 containers. Walk me through what you would fund and what you would refuse.

    3 min answer discordprofilinggarbage-collectiontail-latency
  11. SLO Monitoring advanced

    A canary release looks healthy on p50 latency but a small set of enterprise tenants sees timeouts. Which metrics and gates should have caught it?

    2 min answer canarytenant-segmentationtail-latencyrollout-gates
  12. SLO Monitoring advanced

    A communication platform sets a 99.9% availability SLO. How should alerting on that SLO be structured so it catches both sudden outages and slow degradation?

    2 min answer sloburn-ratemulti-windowalerting