To cut 45% from database spend a team moved the primary from a synchronous two-zone pair to a single zone, with snapshots every 30 minutes and log shipping every 5. The zone never fails. It gets slow - storage latency goes from 1 ms to 80 ms and stays there for two hours. What happens, minute by minute?
Show the full answer Hide the answer
Minute by minute
Minutes 0 to 2. Commit latency follows storage latency. A transaction that took 4 ms takes 120 ms. Throughput per connection collapses by roughly the same factor because each connection is serialised on its own commits.
Minutes 2 to 5. Application connection pools fill. Requests queue for a connection rather than for the database, so application latency rises faster than database latency, which is the first confusing signal.
Minutes 5 to 10. Request latency crosses the caller's timeout. Clients retry. Offered load roughly doubles at the worst possible moment, and the retries arrive at a system whose capacity has already fallen. Error rate goes from zero to significant in about a minute once the pool is exhausted.
Minutes 10 to 120. The system is up, answering some requests, failing others, and nothing changes on its own.
Where it amplifies
The failover that was removed was also the detector. With two zones, a slow zone shows up as replication lag between two things that should agree, and that comparison is a signal available before any user notices. With one zone there is nothing to compare against, so the only evidence is application latency, which has a dozen other causes. Expect the first twenty minutes of the incident to be spent on the application tier.
What the user sees, and who gets paged
Not an error page. Slow checkout, then timeouts on writes while reads from any cache still work. The incident is routed to the wrong team because the symptom is "the app is slow", and the storage metric that would settle it is a percentile on a device nobody has an alert on.
What stops it
Nothing automatic, and that is the point. A zone that is slow is not a zone that fails health checks, so no orchestrator moves anything. The options are to wait, or to restore. Restore means up to 30 minutes of lost writes plus the restore and replay time. At 200 writes per second, a 30-minute recovery point objective is on the order of 360000 lost writes, and somebody has to decide whether those can be reconstructed from upstream logs before choosing.
What would have to be true for it to self-heal
An independent replica outside the failing zone, promotion driven by a latency signal rather than a liveness signal, and an application that tolerates promotion without manual intervention. That is precisely the machinery the saving removed. The 45% bought a change from a recovery point of near zero to 30 minutes, and a change in failure handling from automatic to human. Both can be correct choices. The error is making them without writing them down, so the trade is discovered during the incident rather than agreed in advance.
When this is the wrong answer
Single zone is the right call for a store that can be rebuilt from an authoritative source inside the tolerance of its consumers: a derived index, a cache tier, an analytics replica, a development environment. The test is not "how important does this feel" but two numbers: what is the cost of an hour of unavailability, and what is the time to rebuild from the upstream truth? Where rebuild time is under the tolerance, take the 45% without hesitation. Where the store is the only copy, the saving is a loan against an incident that has not happened yet.