A food-delivery platform in Zomato's mould pushes order-status events to restaurant tablets and to customers through a queue-backed webhook fleet. The complaints are that status arrives late rather than that it never arrives. Which indicator should the SLO be written on?
Show the full answer Hide the answer
The deciding property
The user-visible failure is the latency of an asynchronous outcome, and the clock starts when the order state changed, not when the delivery system decided to try. Any indicator whose clock starts inside the delivery system is structurally unable to see the queueing that produces the complaint.
Why the freshness indicator
Emit delivered_at minus state_changed_at as a histogram and report the fraction under the deadline. One detail makes it honest: an event that is never delivered counts as infinitely late, so it lands in the failure bucket rather than falling out of the denominator. Without that, a system that silently drops 2% of events and delivers the rest in 300 ms scores 100%, and dropping events becomes a way to improve the number.
Express the budget in events rather than in minutes, because the work is discrete. A 99% objective at four million events a day leaves about 1.2 million late events a month, which is a figure a product owner can actually argue about.
Why the other options fail
The 2xx rate measures the last hop. A message that waited eleven minutes in the queue and then got a 200 is counted as a success, and events never attempted are invisible. It is the right indicator for whether tablets are reachable, and it is blind to the reported problem.
Consumer lag under 30 seconds at p99 is a cause metric promoted to an objective. Lag can look healthy while a poison-message retry loop delays one subset by minutes, and lag is per-partition, so a single hot restaurant's partition hides inside the p99 of two hundred partitions. Keep it as a diagnostic and as a paging signal for the pipeline; do not make it the promise.
Queue depth under 10000 has no fixed relationship to impact: 10000 messages is a second of backlog at 50000 messages per second and half an hour at 5. It also reads healthy when the consumer is dead and nothing is being enqueued because upstream broke too.
Synthetic probe availability proves the service accepts work. The probe usually bypasses the queue, it runs at a trickle compared with real load, and it cannot see the tail of real deliveries. It belongs in the black-box layer that tells you the endpoint exists.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| Complaints become "it never arrived" | A completeness indicator on delivered-versus-emitted | The failure is loss rather than delay |
| Deadlines differ sharply by consumer | Per-journey SLOs with separate budgets | A 5-second tablet and an hourly accounting export cannot share one threshold |
| The path becomes synchronous | Request availability and latency percentiles | There is no queue to hide the wait in |
When not to use a freshness indicator
For a synchronous API, availability and latency percentiles on the request already carry the user's experience, and a freshness indicator adds instrumentation and a joined timestamp for nothing. Freshness earns its cost only when work is deferred, because deferral is the mechanism that decouples "the system responded" from "the user knows".