intermediate 4 min answer Multiple choice

Every service in a fleet exposes `/health`, and the check verifies the service can reach its database, its cache and two downstream APIs, returning 503 if any is unreachable. The shared cache has a 90-second degradation. What happens to the fleet during those 90 seconds?

health checkscorrelated failureload balancergraceful degradationblast radius
Pick one
Show the full answer Hide the answer

Second by second, what happens

Second 0. The shared cache degrades. It is shared, so this is not a per-instance event — every instance in the fleet sees it at once.

Second 1–5. Each instance's next health probe runs, attempts the cache, and fails. The check returns 503. Because the probe interval and the failure are independent, the instances do not fail in a staggered way; they fail within one probe interval of each other, which is to say simultaneously, for all operational purposes.

Second 5–15. The load balancer applies its unhealthy threshold — commonly two or three consecutive failures — and begins removing targets. It removes all of them, because all of them are failing. The fleet goes from N healthy targets to zero.

Second 15 onwards. Either the load balancer returns 503 to every client because it has no target, or it enters "fail-open" and routes to all of them anyway, having concluded that zero healthy is a measurement error. The safety valve is then the only thing preventing a total outage, and whether it exists is usually unknown until this moment.

Where it amplifies. Clients retry into a fleet that is refusing traffic or serving on a degraded cache, and the retry volume extends the 90 seconds. Meanwhile any orchestrator with a liveness probe on the same endpoint starts killing and restarting containers, so when the cache recovers the fleet is cold and restarting rather than serving.

What the user sees

Total unavailability for a degradation that, handled correctly, should have been slightly slower pages. The cache is a cache; most of these requests could have been served from the database at higher latency. The system had a working path to a correct answer and the health check forbade it.

Why the other options fail

  • "Only the instances that lost connectivity are removed." This is the mental model everyone has, and it is correct for a dependency private to the instance — a local disk, a sidecar, a per-instance connection. The stem says shared. A shared dependency produces a correlated failure, and health checks that test shared dependencies convert any shared degradation into a fleet-wide one. Correlation is the entire mechanism, and the instinct to reason about one instance at a time is what hides it.
  • "The load balancer keeps the last known healthy set." Some do, under names like fail-open or panic threshold. It is a backstop, not a design: usually off by default, rarely tested, and relying on it makes correctness depend on a flag nobody has looked at. If it saves you, you were lucky.
  • "Requests slow down but no instance is removed." This is what should happen and what a correctly scoped health check would produce. The stem specifies that the check returns 503 when the cache is unreachable, which is precisely the instruction to remove the instance.

What stops it

A mechanism, not vigilance:

  • Separate liveness from readiness from dependency health. Liveness answers "is this process wedged" and must depend on nothing external. Readiness answers "should this instance receive traffic" and may test only dependencies private to this instance. Dependency health is a third thing: useful telemetry, never wired to a decision that removes capacity.
  • Degrade rather than refuse. If the cache is down, serve from the origin and emit a metric. The service is not healthy, and it is useful, and useful beats healthy during an incident.
  • A minimum healthy percentage on the target group. Configure the load balancer never to remove more than, say, 50% of targets regardless of check results. This converts the catastrophic case into the merely degraded one and costs nothing when things are normal.

When not to separate them

The separation above is right for a fleet behind a load balancer. For a single-instance service with a private dependency it is overhead for nothing: if the process cannot reach the one database it exists to serve, removing it from rotation is correct, because there is no degraded mode to fall back to and nothing else to route to. The rule is about shared dependencies and replicated fleets, and stating it as "health checks must never test dependencies" is the over-correction that leaves a genuinely broken instance in rotation serving errors. Ask instead: if this check fails, will it fail on every instance at once, and is there a useful answer the service could still return? Two yeses mean degrade; two noes mean remove.

What would have to be true for it to self-heal

The cache recovers, probes pass, and instances return to rotation — which does happen, and only after the orchestrator has finished restarting the containers it killed, and after the retry backlog has been absorbed by a cold fleet. Self-healing in the narrow sense, with a recovery time several times longer than the original 90-second fault. The lesson is that the recovery time of a system is set by what its automation did during the fault, not by the duration of the fault.