1. Structured Logging advanced

    At 09:40 a release begins logging the whole inbound request as nested JSON, including a customer-supplied metadata map. At 10:05 the log store starts rejecting documents for that index with an illegal-argument error. By 10:20 no service writing to that index can log at all, including services that changed nothing. What happened, and which design decision allowed it?

    2 min answer structured-loggingschemacardinalityblast-radius
  2. Telemetry Cost advanced

    A platform's observability spend has grown to a substantial fraction of its infrastructure budget with no single team responsible. What are the drivers, and how should this be controlled without losing visibility?

    2 min answer telemetry-costcardinalityretentionattribution
  3. Telemetry Cost advanced

    An observability bill has grown to a significant fraction of infrastructure spend. Diagnose it systematically and describe the reductions that do not lose diagnostic power.

    3 min answer costcardinalitysamplingretention
  4. Sampling advanced

    A commerce API serves 2000 requests per second and keeps 1% of traces. The team wants to watch a 0.5% checkout error rate from the trace store and page when it moves by a tenth of itself. Roughly how much traffic must the sample cover before that reading is trustworthy, and what does the answer say about the alert?

    3 min answer samplingstatisticstracingalerting
  5. Sampling advanced

    A platform with extreme traffic spikes needs distributed tracing. Head-based sampling loses the interesting traces and 100% sampling is unaffordable. What strategy resolves this?

    2 min answer dream11tracingtail-samplingsampling
  6. Sampling advanced Multiple choice

    An edge platform handles enormous request volume and cannot retain telemetry for every request. Which sampling strategy should it use, and what breaks under naive uniform sampling?

    2 min answer samplingtail-basedhead-basedstatistics