pattern

Long-Running Operation SLO

also called Operation Completion Objective, Terminal-State Objective

A reliability objective shaped for multi-minute platform operations like deploys and provisioning - share of runs reaching a terminal state within a deadline with no manual repair - replacing availability percentages that cannot describe them.

platform-slosslisprovisioningdeploysmeasurement

The platform's dashboard says the deploy service was 99.98% available last month. Teams say deploys are unreliable. Both are true: the API answered nearly every request, and roughly one deploy in twenty hung mid-run and needed an engineer to unstick it. A request-availability metric cannot see that, because the failure is not in a request.

Deploys, provisioning runs, cluster upgrades and secret rotations are operations, not requests. They are accepted in milliseconds and completed in minutes, they have intermediate states, they can partially apply, and their worst outcome is not an error - it is a run that neither finishes nor fails.

Why it matters

A platform's consumers experience operations. If the published objectives describe components instead, the platform optimises what it measures and the gap becomes the argument that recurs every quarter: dashboards green, users unhappy.

The stuck state is the one worth instrumenting, because it is invisible to every other metric and expensive in a way that compounds: it consumes a human, it blocks the team's pipeline, and it is usually repaired silently, which hides the frequency from everyone who could fix the cause.

Implementation patterns

  • Define success as three conditions together: terminal state, within the deadline, with no manual intervention. Dropping any one produces a number that can be met while the experience stays bad.
  • Report three rates, not one: completion-within-deadline, failure, and stuck. A mature deploy path looks like a completion rate in the high nineties with a stuck rate under about 0.5%.
  • Persist a durable operation record with an identifier, a desired state, an observed state and a monotonic revision, so a caller can distinguish accepted from achieved.
  • Run a timeout sweeper. Anything past its deadline becomes a declared terminal failure, which converts silent stuckness into a counted event.
  • Count manual repair explicitly, even when the run then succeeds, otherwise the platform makes its number with human effort.
  • Make operations idempotent and resumable by identifier, so a retry after a timeout does not create a second database or a second rollout.
  • Report p50 and p95 duration beside the rate, because a 15-minute objective met at p95 of 14 minutes is a different system from one met at p95 of 4 minutes.

Industry example

The pattern is codified in the Kubernetes API conventions, where a controller records observedGeneration against the spec it has actually acted on and publishes conditions with reasons rather than a single boolean. That is what lets a caller tell "accepted" from "achieved" and what lets a platform compute a completion rate at all for work that takes minutes in production. Platforms that expose provisioning as a synchronous call and then wrap it in retries have no equivalent, which is why their only available metric is request availability - the metric that cannot describe the thing their users are complaining about.

Failure scenarios

  • Quiet repair. An engineer finishes hung runs by hand every few days, the completion rate looks healthy, and the cause is never funded.
  • Duplicate resources. A client times out and retries a non-idempotent operation, so two clusters exist and one is paid for and unmonitored.
  • Terminal states that are not terminal. A run marked failed is still mutating cloud resources, and cleanup races the retry.
  • Measured from the platform's side, so an operation that succeeded internally but never became usable counts as success.
  • A deadline set from the median, so half of all runs breach it and the error budget is meaningless from day one.

Trade-offs

Choose Gains Pays
Operation-shaped SLO A number that matches what consumers feel; stuck runs become visible Durable operation state, a sweeper, and an operations API for callers to poll
Request availability only Trivial to measure from existing proxy logs Blind to the dominant failure mode of multi-minute work

When not to use it

Operations that finish inside a few seconds should stay request-shaped; adding durable state and a sweeper to a 200-millisecond call is pure overhead. Under roughly ten consumer teams, a status page, a support rota and a stated response time do most of the work of any published objective. And an objective with no error-budget policy attached is a reporting exercise: if nothing changes when the number is missed, do not publish the number.

Interview question

Q: Your platform promises a 15-minute deploy. Define the SLI precisely, and tell me what it would look like if the platform were gaming it.

What a strong answer covers: terminal state within the deadline with no manual repair; stuck counted separately from failed; measurement from the consumer's side; durable operation records with an observed revision; a timeout sweeper that forces a terminal state; a manual-repair counter; p95 duration alongside the rate; and the gaming modes - operators quietly unsticking runs, cancelling slow runs before the deadline, or excluding "known-slow" classes from the denominator.

Quick check

Quiz: Why is availability the wrong metric for a 15-minute provisioning operation? Because the operation's dominant failure is a run that neither completes nor errors, which produces no failed request and therefore no availability signal.

Flashcard: What three rates describe a long-running platform operation? Completion within the deadline, declared failure, and stuck - with manual repair counted even when the run eventually succeeds.