Dependency Chain Availability
also called Serial Availability, Compound Availability
The product of the availabilities of every component a request must pass through, which is always lower than the weakest link and is the number a team's SLO is actually constrained by.
A team commits to 99.9% availability. Its service runs at 99.9%, and it makes synchronous calls to four internal services, each also at 99.9%, and cannot answer without all of them. Everyone has done their part and the promise is already broken.
0.999 × 0.999 × 0.999 × 0.999 × 0.999 = 0.9950
99.5%, which is about 3 hours 36 minutes of monthly downtime rather than 43 minutes. Nobody made a mistake. The structure produced it, and no individual team's dashboard shows it.
Dependency chain availability is that product: the probability that every component in the serial path is available at once. It is the number the SLO is constrained by, and it is almost never the number anyone quotes.
Why it matters
It converts "add a dependency" from a free architectural decision into a priced one. A new synchronous call to a 99.9% service costs roughly 0.1 percentage points of availability, permanently, and that is the line missing from every design document that adds one.
Two consequences follow, and both are worth having to hand in a review:
- You cannot be more available than the weakest thing you cannot answer without. A team targeting 99.99% behind a 99.9% third-party API has chosen a target the arithmetic forbids, and no redundancy in their own tier changes it.
- Depth is worse than breadth. Ten dependencies at 99.99% give 99.9%; three at 99.9% give 99.7%. Reducing the number of hard dependencies beats improving any one of them, which is the opposite of where effort usually goes.
Useful reference points, for a 30-day month: 99.99% ≈ 4 minutes · 99.9% ≈ 43 minutes · 99.5% ≈ 3 h 36 min · 99% ≈ 7 h 12 min.
Implementation patterns
- Draw the hard-dependency graph and multiply. Hard means "cannot answer without". The exercise usually finds two or three dependencies nobody considered hard, and that list is the deliverable.
- Convert hard dependencies to soft ones. If you can answer degraded, cached or with a default, the dependency leaves the product entirely. This is why graceful degradation is an availability technique rather than a user-experience nicety: it changes the arithmetic instead of improving a term in it.
- Make the call asynchronous. Accept, enqueue, respond. The dependency's availability now affects latency, which is a far cheaper currency than availability.
- Budget the chain explicitly. If the target is 99.9%, allocate: 99.97% for your own tier and the rest across dependencies, with each owner holding a stated share. Unallocated chains are where the gap hides.
- Add redundancy only where failures are genuinely independent. Two instances at 99.9% give 99.9999% on paper and far less in reality, because they share a config push, a deploy pipeline and an image.
- Track it as a derived metric, recomputed when the graph changes, so adding a dependency visibly moves a number someone owns.
Industry example
The arithmetic is why cloud providers publish per-service SLAs rather than a platform figure, and why composite-SLA calculation appears in their architecture guidance: a customer assembling six services at 99.95% has built something with a roughly 99.7% availability ceiling, and the provider is explicit that the composite is theirs to compute. The same reasoning drives the standard advice to prefer asynchronous and cached paths at the edges of a system — not for latency, but because each synchronous hop removed is a multiplicand removed, and that is the only lever with a reliable effect on the product.
Failure scenarios
- An SLO promised that the graph forbids, discovered at the first quarterly review when the budget was exhausted in week two.
- The invisible dependency. A library that calls a service, a sidecar that resolves a feature flag, an auth check on every request. These are in the chain and absent from the diagram.
- Correlated failure making the real figure worse than the product. One bad config push, one expired certificate or one regional event takes several multiplicands together, so the independence the formula assumes is the best case rather than the typical one.
- Shared sub-dependencies double-counted as independent. Four services that all call the same database are one dependency wearing four hats, and the product formula flatters it.
- Optimising the wrong term. A team spends a quarter taking its own tier from 99.9% to 99.99% while five dependencies at 99.9% set the ceiling, and the measured availability does not move.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Fewer hard dependencies | The product improves directly and permanently | Duplicated logic or data, or a less normalised design |
| Degrade instead of failing | The dependency leaves the availability chain | A degraded answer must be correct enough to serve, and must be built and tested |
| Asynchronous calls | Availability cost becomes a latency cost | Eventual consistency, and a queue to operate |
When not to use it
It is a planning tool, not a prediction. Two limits matter in practice. Published availability figures are not probabilities of independent events, so the product is optimistic about correlated failure. And most downtime comes from change, not from random failure — a system whose arithmetic says 99.5% may do much better if it changes rarely, and much worse if it deploys forty times a day with no fast rollback.
So use the multiplication to reject impossible targets and to price each new synchronous call. Do not use it to forecast next quarter's uptime; that number is set by your change process, and a team that models availability carefully while deploying without canaries has measured the wrong thing.
It is also the wrong lens for a system with few dependencies and a high change rate. A two-component system deploying continuously should spend its attention on rollback speed and staged rollout, where the actual downtime originates, rather than on an arithmetic that will not move.
Interview question
Q: Your team commits to 99.95%. You call six internal services synchronously, each committed to 99.95%. What do you say in the review?
What a strong answer covers: doing the multiplication aloud — 0.9995⁷ ≈ 99.65%, roughly 1 hour 50 minutes a month against the 22 minutes promised · stating that the commitment is arithmetically unavailable before any engineering is discussed · identifying which of the six are genuinely hard and proposing degradation or async for the rest, since removing a multiplicand is the only reliable lever · noting that shared sub-dependencies make the real figure worse than the product · and observing that if the change process has no fast rollback, the arithmetic is not where the downtime will come from anyway.
Quick check
Quiz: Five components in series, each 99.9%. What is the chain availability? — 0.999⁵ ≈ 99.5%, about 3 h 36 min of monthly downtime.
Flashcard: Which improves chain availability more — raising one dependency from 99.9% to 99.99%, or removing a 99.9% dependency entirely? — Removing it. It deletes a multiplicand; raising one improves a term, and depth hurts more than any single term's value.