Your service is 99.9% available. It calls four dependencies, each also 99.9% available, and it cannot answer without all of them. What is your actual availability, and what does that arithmetic imply about architecture?
Show the full answer Hide the answer
The mechanism
Availability through a chain multiplies. If every link must work, the probability that all of them do is the product of their individual probabilities:
0.999 × 0.999 × 0.999 × 0.999 × 0.999 = 0.9950
So 99.5%, not 99.9% — and that is before anything specific to your system, like a shared dependency failing two links at once.
Translated into time, which is where it becomes concrete:
| Availability | Downtime per 30-day month |
|---|---|
| 99.9% | about 43 minutes |
| 99.5% | about 3 hours 36 minutes |
| 99% | about 7 hours 12 minutes |
You promised 43 minutes and the arithmetic delivers over three and a half hours. Nobody made a mistake; the structure produced it.
The consequence people miss
Each new hard dependency subtracts from your availability budget, permanently. This is the honest cost of adding a synchronous call, and it is the cost nobody puts in the design document. A fifth dependency at 99.9% takes you to roughly 99.4%; a dependency at 99% takes you to 98.5% on its own, whatever else you do.
The corollary is the useful one: you cannot be more available than the weakest thing you cannot answer without. A team targeting 99.99% while synchronously calling a 99.9% third-party API has set a target the arithmetic forbids, and no amount of redundancy in their own tier changes it.
What to do about it
Three moves, in descending order of value:
- Make the dependency non-hard. If you can answer — degraded, cached, with a default — when it is unavailable, it leaves the product. This is why graceful degradation is an availability technique and not a user-experience nicety: it changes the arithmetic rather than improving a term in it.
- Make the call asynchronous. Accept the request, queue the work, respond. The dependency's availability now affects latency rather than availability, which is a far cheaper currency.
- Add redundancy only on the links you control. Two independent instances at 99.9% give 1 − 0.001² = 99.9999% — but only if the failures are genuinely independent, and they usually are not. Shared power, shared network, shared configuration push and shared deploy pipeline correlate failures, so the real figure is far lower than the formula suggests.
Where the arithmetic stops being the answer
It is a planning tool, not a prediction. Two caveats that matter in practice:
- Published availability figures are not probabilities of independent events. Real outages are correlated: one bad config push, one expired certificate, one region event takes several dependencies together. The product formula assumes independence and is therefore optimistic about correlated failure and pessimistic about the common case.
- Most downtime comes from change, not from random failure. Deploys, config, migrations. A system whose arithmetic says 99.5% may do considerably better if it changes rarely, and considerably worse if it deploys forty times a day without a fast rollback.
So use the multiplication to reject targets that are impossible and to price each new synchronous call. Do not use it to predict next quarter's uptime — that number is set by your change process.
Common weak answers
- "99.9%, the same as the weakest link." That is the ceiling for one dependency, not the result of five in series.
- "Add retries." Retries help with transient blips and do nothing for the dependency being down, which is the case the figure describes. They also consume the caller's latency budget.
- "Add more replicas of my own service." Improves your term and leaves the four multiplicands untouched.
- Quoting the formula and stopping. The point is the design consequence — remove the dependency from the critical path — not the number.