1. Health Checks advanced

    A liveness probe checks the database. The database slows down. Describe what happens.

    2 min answer health-checkscascading-failurelivenessreadiness
  2. Health Checks advanced

    A platform's gateway processes pass their health checks while holding connections they can no longer serve. What is wrong with the health check, and what should it assert?

    2 min answer health-checkslivenessreadinessprogress
  3. Health Checks intermediate

    A platform's health check returns 200 while the service cannot serve real traffic. What should a health check actually verify, and what should it deliberately not?

    2 min answer vercelhealth-checksreadinessliveness
  4. Log Management advanced

    A globally distributed platform generates logs at every edge location. Should logs be centralised, kept local, or something else? Analyse the trade-off.

    2 min answer log-managementedgeaggregationcost
  5. Log Management advanced

    An observability vendor indexes logs by labels only, not full text, to control cost. Which queries become cheap, which become expensive, and how should teams structure logs and labels to fit?

    2 min answer grafanalokiindexingcardinality
  6. Log Management advanced

    Datadog described Husky in 2022 as its third-generation event store: writers consume from a queue, write schemaless columnar files to object storage, and commit the presence of those files to a transactional metadata store, with ingestion, storage and query scaled independently. What forced that shape rather than a conventional search cluster on local disks, and where would copying it be a mistake?

    2 min answer datadoglog-managementobject-storagecompaction
  7. Log Management advanced

    You own the logging platform for a travel company of Expedia's shape. Legal requires that records evidencing a booking and a payment be retrievable for seven years. Engineering wants 30 days of everything, searchable in seconds. Finance has capped the platform's spend. Walk me through the design, and tell me what you would refuse.

    3 min answer expediaretentioncompliancelog management
  8. Logging intermediate

    A discussion platform's log volume grows faster than its traffic and now costs more than the compute generating it. What is driving the growth, and what should change?

    2 min answer loggingcostverbosityretention
  9. Logging intermediate

    A platform team stops shipping logs from inside each application process to the log vendor and instead writes JSON to stdout for a node agent to collect. Crash-time logs now survive and the request path no longer touches the network. What has the team given up, and when does that bill arrive?

    2 min answer loggingkuberneteslog-rotationtruncation
  10. Logging beginner

    A service emits one INFO log line per request and increments one counter per request. At 3,000 requests per second the two describe exactly the same traffic. Why does keeping the log cost orders of magnitude more than keeping the counter, and when is the log still the right thing to emit?

    3 min answer loggingmetricstelemetry costcardinality
  11. Logging intermediate

    Your logging bill has tripled alongside traffic growth. Name four changes in order of effectiveness.

    2 min answer loggingcostsamplingretention
  12. Metrics intermediate

    A live platform's dashboards show mean latency, which stays flat during an incident where many users experience severe delays. Why do averages hide this, and what should be measured?

    2 min answer metricspercentilesaggregationtail-latency