1. Dashboards intermediate

    A platform has hundreds of dashboards and engineers cannot find the right one during an incident. How should dashboards be structured?

    2 min answer dashboardsincident-responsehierarchybooking
  2. Dashboards intermediate

    Design the one dashboard your team opens first during an incident. What is on it and what is deliberately not?

    2 min answer dashboardsincident-responsedesignobservability
  3. Dashboards intermediate

    Review this incident dashboard. One page holds 28 panels, the default range is 6 hours, auto-refresh is 10 seconds, the scrape interval is 30 seconds, and every panel plots a 5-minute rate over per-pod series for a 900-pod fleet. During the last incident the page took 40 seconds to load and two panels timed out. What would you remove, what would you change, and what would you leave alone?

    3 min answer dashboardsquery-costresolutionincident-response
  4. Dashboards intermediate

    Your metrics pipeline normally makes data queryable about 20 seconds after it is emitted. During a large incident the ingestion tier falls behind and the lag grows to four minutes. Nothing is lost. Walk through what happens to the on-call engineer's decisions.

    3 min answer telemetry lagincident responserollbackalerting
  5. Debugging Distributed Systems intermediate

    A deployment completes successfully. Twenty minutes later latency has tripled, but only on new instances. Diagnose.

    2 min answer deploymentcold-startdiagnosiscaching
  6. Distributed Tracing intermediate

    A request takes 3 seconds. Every individual service reports healthy latency. How does tracing resolve this and what must have been instrumented?

    2 min answer tracinglatencysamplinginstrumentation
  7. Health Checks intermediate

    A platform's health check returns 200 while the service cannot serve real traffic. What should a health check actually verify, and what should it deliberately not?

    2 min answer vercelhealth-checksreadinessliveness
  8. Logging intermediate

    A discussion platform's log volume grows faster than its traffic and now costs more than the compute generating it. What is driving the growth, and what should change?

    2 min answer loggingcostverbosityretention
  9. Logging intermediate

    A platform team stops shipping logs from inside each application process to the log vendor and instead writes JSON to stdout for a node agent to collect. Crash-time logs now survive and the request path no longer touches the network. What has the team given up, and when does that bill arrive?

    2 min answer loggingkuberneteslog-rotationtruncation
  10. Logging intermediate

    Your logging bill has tripled alongside traffic growth. Name four changes in order of effectiveness.

    2 min answer loggingcostsamplingretention
  11. Metrics intermediate

    A live platform's dashboards show mean latency, which stays flat during an incident where many users experience severe delays. Why do averages hide this, and what should be measured?

    2 min answer metricspercentilesaggregationtail-latency
  12. Metrics intermediate

    A multi-tenant SaaS platform cannot tell which customer caused a performance problem. What must be present in the telemetry, and what does that cost?

    2 min answer freshworksmulti-tenancyattributioncardinality