Observability

General material on understanding a system from its outputs.

8Questions
18Flashcards
17Terms
1Deliverables
Questions

8 to work through

  1. intermediate

    A new service goes live next week. What observability must exist on day one?

    2 min answer
  2. intermediate

    You are designing a new service. What must be in place before it goes live so that whoever is paged at 3 AM can diagnose it without you?

    2 min answer
  3. advanced

    A commerce platform can tell that its peak-day checkout success rate has dropped, but cannot tell which of forty services is responsible. What is missing, and in what order should it be added?

    2 min answer
  4. advanced

    An analytics platform ingests billions of events while customers run arbitrary segmentation queries. How should the two workloads be isolated, and when should pre-aggregation be introduced?

    2 min answer
  5. advanced

    An error-tracking platform receives millions of events, many of them the same underlying problem with different stack details. How should grouping, deduplication, high-cardinality metadata and retention be designed?

    2 min answer
  6. advanced

    On 11 December 2024 OpenAI rolled a new telemetry service out across every Kubernetes cluster. Within about half an hour the API servers were saturated and services could no longer resolve one another; full recovery took until the evening. What turned a monitoring change into a total outage, and which of the contributing factors would you fix first?

    3 min answer
  7. advanced

    Your observability bill is now 40% of your compute bill. Leadership wants it cut in half without going blind. What do you cut?

    2 min answer
  8. advanced

    p99 latency on checkout tripled overnight. Dashboards look normal, no deployment went out, and every service reports healthy. How do you find it?

    3 min answer
Terminology

17 terms in this topic

concept

Alert Fatigue

The desensitisation that follows from alerts that are frequent, non-actionable, or not tied to user impact — after which real alerts are missed too.

tool

Application Performance Monitoring

Instrumentation inside the application that attributes latency and errors to specific code paths, queries and dependencies.

concept

Cardinality

The number of distinct time series produced by a metric, which is the product of the distinct values of all its labels — and the main driver of monit…

practice

Correlation ID

A single identifier attached to one logical operation and included in every log line it produces, anywhere in the system.

tool

Distributed Tracing

Following one logical request across every service it touches by propagating a shared trace identifier and recording timed spans.

pattern

Error Fingerprinting

Deriving a stable identifier from an error's invariant attributes so that many occurrences collapse into one actionable issue - and the two direction…

metric

Golden Signals

The four measurements that cover most of what matters for a request-driven service: latency, traffic, errors and saturation.

practice

Health Check

An endpoint the platform polls to decide whether an instance should be restarted or should receive traffic — two different questions needing two diff…

pattern

Ingestion-Query Isolation

Separating a continuous high-volume write path from a spiky, arbitrary query path so that an expensive query cannot stall ingestion - because a query…

concept

Observability

The property of being able to answer new questions about a system's internal state from its external outputs, without shipping new code.

pattern

Observability in Practice

The three signals, what each is actually for, and why the links between them matter more than any of them individually.

practice

RED Method

A minimal per-service dashboard: Rate, Errors, Duration — the request-centric view of whether users are being served.

practice

Runbook

A short, actionable document telling an on-call engineer what an alert means, what to check, and what the safe mitigations are.

metric

Telemetry Pipeline Lag

The delay between an event being emitted and being queryable, which bounds how quickly any decision made from telemetry can respond to reality.

practice

Telemetry Sampling

Keeping a subset of traces or events to bound observability cost, chosen so the ones that matter survive.

practice

USE Method

For every resource, track Utilisation, Saturation and Errors — the resource-centric complement to request-centric monitoring.

concept

Wide Event

A single structured record per unit of work carrying every field that might matter, aggregated only at query time - so questions nobody anticipated r…

Observability

Neighbouring topics

Logging

What to log, at what level, and what must never appear in a log.

4 quiz 8 cards 3 terms

Structured Logging

Machine-parseable events with stable names and consistent fields.

3 quiz 8 cards 3 terms

Metrics

Counters, gauges and histograms, and percentiles rather than means.

4 quiz 14 cards 4 terms

Cardinality

The label that multiplies series count and the bill with it.

5 quiz 11 cards 2 terms

Distributed Tracing

Reconstructing one request's path across every service it touched.

4 quiz 10 cards 5 terms

Correlation IDs

One identifier propagated through every hop and every log line.

3 quiz 9 cards 3 terms

Sampling

Head-based versus tail-based, and keeping the traces that matter.

3 quiz 11 cards 2 terms

Health Checks

Liveness versus readiness, and the check that causes the outage.

3 quiz 10 cards 3 terms

Alerting

Symptom-based, actionable, user-impacting — and linked to a runbook.

4 quiz 9 cards 3 terms

Alert Fatigue

How noise makes the real page invisible, and the structural fix.

6 quiz 12 cards 3 terms

Dashboards

Answering 'is it us' in under a minute, for someone who was asleep.

4 quiz 8 cards 2 terms

Application Performance Monitoring

Attributing latency to code paths, queries and dependencies.

3 quiz 7 cards 3 terms

Profiling

Continuous CPU and memory attribution in production.

4 quiz 12 cards 3 terms

Business Metrics

Orders per minute alongside error rate, because healthy is not enough.

3 quiz 9 cards 3 terms

SLO Monitoring

Burn-rate alerting that fires on user impact rather than on thresholds.

4 quiz 9 cards 2 terms

Log Management

Aggregation, retention tiering, search and the cost of keeping everything.

4 quiz 11 cards 2 terms

Telemetry Cost

Observability bills that rival compute, and where to cut without going blind.

4 quiz 10 cards 2 terms

Debugging Distributed Systems

Localising a regression when every service reports healthy.

4 quiz 10 cards 2 terms

OpenTelemetry

Instrumenting once against an open standard rather than a vendor agent.

4 quiz 9 cards 3 terms