1. Debugging Distributed Systems advanced

    A streaming consumer's lag is growing steadily while workers show low CPU and no errors. What are the candidate causes and how do you distinguish them?

    2 min answer confluentconsumer-lagdiagnosisbackpressure
  2. Distributed Tracing advanced

    A marketplace adopts distributed tracing but engineers rarely use it during incidents. What typically causes low adoption, and what makes tracing actually useful?

    2 min answer tracingsamplingadoptionexemplars
  3. Distributed Tracing advanced

    A team enabling distributed tracing must decide between head-based and tail-based sampling. Compare them, and explain what makes traces useful beyond a single request's timeline.

    3 min answer tracingsamplingopentelemetrytail-based
  4. Distributed Tracing advanced

    You are introducing distributed tracing across 40 services owned by 12 teams. Plan the adoption.

    2 min answer tracingopentelemetryadoptionsampling
  5. Health Checks advanced

    A liveness probe checks the database. The database slows down. Describe what happens.

    2 min answer health-checkscascading-failurelivenessreadiness
  6. Health Checks advanced

    A platform's gateway processes pass their health checks while holding connections they can no longer serve. What is wrong with the health check, and what should it assert?

    2 min answer health-checkslivenessreadinessprogress
  7. Log Management advanced

    A globally distributed platform generates logs at every edge location. Should logs be centralised, kept local, or something else? Analyse the trade-off.

    2 min answer log-managementedgeaggregationcost
  8. Log Management advanced

    An observability vendor indexes logs by labels only, not full text, to control cost. Which queries become cheap, which become expensive, and how should teams structure logs and labels to fit?

    2 min answer grafanalokiindexingcardinality
  9. Log Management advanced

    Datadog described Husky in 2022 as its third-generation event store: writers consume from a queue, write schemaless columnar files to object storage, and commit the presence of those files to a transactional metadata store, with ingestion, storage and query scaled independently. What forced that shape rather than a conventional search cluster on local disks, and where would copying it be a mistake?

    2 min answer datadoglog-managementobject-storagecompaction
  10. Log Management advanced

    You own the logging platform for a travel company of Expedia's shape. Legal requires that records evidencing a booking and a payment be retrievable for seven years. Engineering wants 30 days of everything, searchable in seconds. Finance has capped the platform's spend. Walk me through the design, and tell me what you would refuse.

    3 min answer expediaretentioncompliancelog management
  11. Observability advanced

    A commerce platform can tell that its peak-day checkout success rate has dropped, but cannot tell which of forty services is responsible. What is missing, and in what order should it be added?

    2 min answer observabilitytracingattributionpeak-events
  12. Observability advanced

    An analytics platform ingests billions of events while customers run arbitrary segmentation queries. How should the two workloads be isolated, and when should pre-aggregation be introduced?

    2 min answer posthogamplitudeingestionquery-isolation