metric

Cost of Downtime

The quantified business impact of unavailability per unit of time, which converts reliability investment from an argument into an arithmetic comparison.

Reliability discussions stall because one side argues for availability and the other for cost, with no common unit. The cost of downtime supplies one.

Components to include: lost revenue during the outage, adjusted for what is merely deferred rather than lost — a retailer recovers some orders, a payments processor does not; operational cost of the response and recovery; contractual penalties or service credits; customer churn, which is the largest and hardest component; and reputational and regulatory consequences for severe events.

The figure differs enormously by system and by time. An hour of checkout unavailability on a peak trading day and an hour of internal reporting downtime on a Sunday differ by orders of magnitude — which is exactly why a single platform-wide availability target is the wrong shape, and why tiered SLOs follow directly from this calculation.

The comparison it enables: each additional nine costs roughly an order of magnitude more and changes the mechanisms — 99.9% needs redundancy and competent operations, 99.99% needs automated failover with no manual step, 99.999% needs multi-region active-active. Price the next nine and set it against the downtime it prevents.

The honest outcome is often that the current target is already correct, and that the money is better spent on reducing time to restore.