advanced 3 min answer

A payments company moves to serving live traffic from two cloud providers at once so that no single provider outage can stop payments. Twelve months later measured availability is lower than it was on one provider and there has been no provider outage. What did they buy and what did they pay?

multi-cloudavailability-mathentity-homingfailoveroperability
Show the full answer Hide the answer

What was gained

Exactly one failure class: a provider-wide or region-wide event the provider cannot recover inside the company's recovery-time objective. These are real and rare, a handful of multi-hour events per major provider per decade. The ask is reasonable on its face, because against a 99.99% annual target — about 52 minutes of error budget — a single four-hour provider event spends five years of budget at once.

What was paid

  • The data layer, permanently. Cross-provider replication runs over the public internet or a dedicated interconnect. Inside one metro a cross-provider round trip is a few milliseconds; across US coasts it is 60 to 80 ms. Synchronous commit is therefore unavailable, so the recovery point is greater than zero forever and something needs conflict resolution. For payments the only safe resolution is never writing the same entity on both sides, which is entity homing — and entity homing does not require two providers.
  • Two of everything that is not the application. Two identity models and policy languages, two network models, two secret stores, two observability pipelines, two quota systems, two support contracts. The team did not double.
  • A ceiling at the lowest common denominator. Every capability must exist on both sides, so the managed services that would have removed operational work are excluded by construction.

The availability arithmetic

This is where the surprise comes from. Redundancy divides failure probability only when either side can serve the request alone with no shared coordination. The moment one request touches a component in both providers — a global lock, a cross-provider sequence number, a shared ledger write — the two systems are in series and their failure probabilities add rather than multiply down. Two paths at 99.95% in series give 99.90%, worse than either alone.

That is the usual explanation for availability falling after a dual-provider migration with no provider outage: the coordination added to keep the two sides consistent became a new single point of failure, and it is the component nobody drew on the diagram.

When the bill arrives

At the switch. A dual-provider posture is rehearsed rarely, because a rehearsal means running production on half the estate during a payments peak. So the failover path accumulates unexercised configuration, stale capacity assumptions and quotas nobody raised on the standby side, and the outage it was built for is the first time it runs end to end. An unrehearsed failover is a hypothesis, and the measured availability already reflects the hypotheses that turned out false in the small failures.

How to keep the option to reverse: one provider primary and authoritative, the second holding a continuously restored copy with a timed scheduled promotion drill, and a costed exit document. That keeps most of the resilience, keeps one identity and network model, and can be abandoned without unwinding a replicated data layer.

When this is the wrong answer

A request path with no shared state genuinely does divide failure probability across providers, because nothing is in series: stateless edge delivery, a read-only content path, batch inference against an immutable model, DNS. Classify your request paths first and run dual-provider only for the ones with nothing in the middle. Serving the read path from two providers while writes stay homed to one is a defensible design; the version in the stem is not.