1. Dashboards intermediate

    A platform has hundreds of dashboards and engineers cannot find the right one during an incident. How should dashboards be structured?

    2 min answer dashboardsincident-responsehierarchybooking
  2. Dashboards intermediate

    Design the one dashboard your team opens first during an incident. What is on it and what is deliberately not?

    2 min answer dashboardsincident-responsedesignobservability
  3. Dashboards intermediate

    Review this incident dashboard. One page holds 28 panels, the default range is 6 hours, auto-refresh is 10 seconds, the scrape interval is 30 seconds, and every panel plots a 5-minute rate over per-pod series for a 900-pod fleet. During the last incident the page took 40 seconds to load and two panels timed out. What would you remove, what would you change, and what would you leave alone?

    3 min answer dashboardsquery-costresolutionincident-response
  4. Dashboards intermediate

    Your metrics pipeline normally makes data queryable about 20 seconds after it is emitted. During a large incident the ingestion tier falls behind and the lag grows to four minutes. Nothing is lost. Walk through what happens to the on-call engineer's decisions.

    3 min answer telemetry lagincident responserollbackalerting
  5. Debugging Distributed Systems advanced

    A collaborative editing platform receives a report that one document occasionally shows different content to different users, but it cannot be reproduced. How do you investigate a rare, non-deterministic, distributed bug?

    2 min answer debuggingnon-determinismdistributedforensics
  6. Debugging Distributed Systems intermediate

    A deployment completes successfully. Twenty minutes later latency has tripled, but only on new instances. Diagnose.

    2 min answer deploymentcold-startdiagnosiscaching
  7. Debugging Distributed Systems advanced

    A platform of 40 services has logs only, and incidents take hours to diagnose. Design the observability strategy and its rollout order.

    2 min answer observabilitytracingmetricsrollout
  8. Debugging Distributed Systems advanced

    A streaming consumer's lag is growing steadily while workers show low CPU and no errors. What are the candidate causes and how do you distinguish them?

    2 min answer confluentconsumer-lagdiagnosisbackpressure
  9. Distributed Tracing advanced

    A marketplace adopts distributed tracing but engineers rarely use it during incidents. What typically causes low adoption, and what makes tracing actually useful?

    2 min answer tracingsamplingadoptionexemplars
  10. Distributed Tracing intermediate

    A request takes 3 seconds. Every individual service reports healthy latency. How does tracing resolve this and what must have been instrumented?

    2 min answer tracinglatencysamplinginstrumentation
  11. Distributed Tracing advanced

    A team enabling distributed tracing must decide between head-based and tail-based sampling. Compare them, and explain what makes traces useful beyond a single request's timeline.

    3 min answer tracingsamplingopentelemetrytail-based
  12. Distributed Tracing advanced

    You are introducing distributed tracing across 40 services owned by 12 teams. Plan the adoption.

    2 min answer tracingopentelemetryadoptionsampling