advanced 2 min answer

A board's risk appetite statement allows at most four hours of customer-visible unavailability per year for the payments API. The service currently runs as two VMs in a single Azure availability set behind a load balancer, with a managed database in the same region. Roughly what availability does the appetite imply and does the design fit?

risk-appetiteavailabilityslaazureerror-budget
Show the full answer Hide the answer

The assumptions, stated

  • A year is 525,600 minutes. Four hours is 240 minutes.
  • "Customer-visible unavailability" includes planned change. Boards mean total; engineers usually hear "unplanned".
  • The dependency chain in the request path is serial: load balancer, compute, database, plus identity and DNS.

The arithmetic

240 / 525,600 = 0.000457, so the appetite is about 99.954% availability. The reference points that matter:

Target Minutes per year
99.9% 526
99.95% 263
99.954% 242
99.99% 53

Microsoft's published SLA for virtual machines gives 99.9% for a single VM using premium SSD, 99.95% for two or more instances in an availability set, and 99.99% for two or more instances spread across two or more availability zones in the same region. The current design's SLA floor of 99.95% is 263 minutes, which is already outside a 240-minute appetite before anything else in the path is counted.

Serial composition makes it worse. Five independent components each at 99.99% multiply to 0.9999^5 = 99.95%, another 263 minutes. Getting a whole path to 99.954% requires every component above it, so the honest reading is that a single-zone design cannot meet this appetite and a zonal design can only meet it if change is nearly free.

Which assumption dominates the error

Not the infrastructure SLA. It is planned change and human error, which the SLA excludes entirely and which in most organisations contributes more downtime than hardware does. A team deploying twice a week with a two-minute degradation per deploy spends 208 minutes a year — 87% of the appetite — on deployments alone. That single number usually decides the design: it forces zero-downtime deployment and connection draining before it forces multi-region anything.

What the number rules in and out

  • Ruled out: single availability set, in-place deployments, maintenance windows treated as free, a database with a failover that takes minutes.
  • Ruled in: zonal redundancy, deployment that does not interrupt in-flight requests, and an error budget that turns the appetite into a monthly number engineers can see.
  • Not settled: multi-region. Cross-region failover buys availability against a whole-region event and pays with data-consistency complexity on every write. At 99.954% the maths does not force it; at 99.99% with a regional risk in scope it does.

An SLA is also not a prediction. It is a credit — a discount on the bill when the provider misses. No credit returns the four hours to the board, so treat published numbers as a floor for design comparison and measure your own achieved availability against the appetite.

When this is over-thinking it

If the appetite were eight hours (99.9%), the current design plus fast restore is fine and the money belongs in recovery time rather than redundancy. Do the composition arithmetic before proposing architecture: it takes ten minutes and it is the difference between an appetite statement that constrains design and one that is decoration.