1. Metrics intermediate

    A multi-tenant SaaS platform cannot tell which customer caused a performance problem. What must be present in the telemetry, and what does that cost?

    2 min answer freshworksmulti-tenancyattributioncardinality
  2. Metrics intermediate

    Etsy released StatsD in 2011: a small daemon that receives metric samples from applications over UDP, aggregates them in memory, and flushes to the metrics backend on a fixed interval. Why was UDP the right choice for that design, and what did it make permanently impossible?

    3 min answer etsystatsdmetricsudp
  3. Metrics intermediate

    Every service dashboard shows healthy metrics while users cannot complete a purchase. How is that possible and what would have detected it?

    2 min answer metricsbusiness-metricsdetectionblind-spots
  4. Observability advanced

    A commerce platform can tell that its peak-day checkout success rate has dropped, but cannot tell which of forty services is responsible. What is missing, and in what order should it be added?

    2 min answer observabilitytracingattributionpeak-events
  5. Observability intermediate

    A new service goes live next week. What observability must exist on day one?

    2 min answer observabilitylaunchoperations
  6. Observability advanced

    An analytics platform ingests billions of events while customers run arbitrary segmentation queries. How should the two workloads be isolated, and when should pre-aggregation be introduced?

    2 min answer posthogamplitudeingestionquery-isolation
  7. Observability advanced

    An error-tracking platform receives millions of events, many of them the same underlying problem with different stack details. How should grouping, deduplication, high-cardinality metadata and retention be designed?

    2 min answer sentrygroupingdeduplicationretention
  8. Observability advanced

    On 11 December 2024 OpenAI rolled a new telemetry service out across every Kubernetes cluster. Within about half an hour the API servers were saturated and services could no longer resolve one another; full recovery took until the evening. What turned a monitoring change into a total outage, and which of the contributing factors would you fix first?

    3 min answer openaikubernetescontrol planedns
  9. Observability intermediate

    You are designing a new service. What must be in place before it goes live so that whoever is paged at 3 AM can diagnose it without you?

    2 min answer observabilityoncallalertinglogging
  10. Observability advanced

    Your observability bill is now 40% of your compute bill. Leadership wants it cut in half without going blind. What do you cut?

    2 min answer observabilitycostcardinalitysampling
  11. Observability advanced

    p99 latency on checkout tripled overnight. Dashboards look normal, no deployment went out, and every service reports healthy. How do you find it?

    3 min answer debuggingtracinglatencydiagnosis
  12. OpenTelemetry advanced

    A fleet exports OTLP to a gateway collector deployment with memory_limiter first in the pipeline, then a batch processor, then an exporter with a sending queue. The telemetry backend starts answering in 8 seconds instead of 80 ms, and request volume doubles at the same time. What happens second by second, and what stops it?

    2 min answer opentelemetrycollectorbackpressureload-shedding