intermediate 3 min answer Multiple choice

A grocery-delivery marketplace peaks for three hours on Sunday morning and keeps orders in a single-region managed Postgres with multi-zone failover. After a competitor's regional outage the board asks for a 15-minute recovery time and zero data loss against the loss of the whole region. Which posture should the team adopt?

cloud-disaster-recoveryrpo-rtowarm-standbyentity-homingnegotiation
Pick one
Show the full answer Hide the answer

The deciding property

Two workload facts settle it. Order writes are concentrated in a three-hour Sunday peak, and inventory reservations must be unique. A recovery point of zero across regions requires a synchronous commit across the inter-region round trip, which is 25 to 40 ms for a continental pair and 60 to 80 ms coast to coast, paid on every write during the peak. A zero recovery point across regions converts an availability requirement into an availability risk, because a synchronous commit cannot complete during a partition: writes stop rather than fail over.

Why the warm standby wins

Asynchronous replication keeps the standby within tens to hundreds of milliseconds of the primary when healthy, so the data genuinely at risk in a regional loss is the replication lag at that instant, usually under a second. Converting "zero" into "30 seconds, measured and alarmed" is the honest negotiation: 30 seconds of Sunday peak orders is a reconciliation job with a known cost, and most of those orders are reconstructable from payment-provider records and client retries. A warm standby — scaled down, running, replicating — meets 15 minutes because the only failover steps are promotion, traffic shift and scale-up.

Why the other options fail

Multi-zone plus hourly snapshots. The posture most teams already have, correct for the zone case and wrong for the region case. Recovery point is up to an hour of peak ordering and recovery time is the restore of the full dataset plus the time to recreate everything around it: identity, secrets, queues, caches, quotas. It also dodges the question the board is really asking, which is whether the capability has ever been exercised.

Active-active writes. Attractive because it removes failover time entirely, and specifically wrong for this workload. Inventory reservation is the one entity that cannot tolerate concurrent regional writes — two regions each selling the last unit is an oversell, and the compensation is a cancelled order on a Sunday morning, which is the most expensive customer moment this business has. Active-active is viable here only with entities homed to one region, which is a different design from writes in both.

Synchronous cross-region commit. It does deliver a zero recovery point, and it pays on every write, during the peak, forever. The team would meet the stated requirement and lose more availability to inter-region latency and partitions than the regional outage they were protecting against.

What would flip the decision

If this changes Choose Because
Orders become legally non-repudiable with no tolerance for loss Synchronous commit across zones in a nearby region pair A 10 to 15 ms metro-pair round trip is affordable where a continental one is not
Inventory stops being scarce (digital goods or unlimited stock) Active-active with entity homing The oversell failure mode disappears
Peak flattens and write rate halves Keep multi-zone and add cross-region snapshot copies The regional posture is not yet worth its standing cost

When not to build any of this

If the business has never lost a region, and the modelled cost of a four-hour regional outage on a Sunday is smaller than a year of operating a second region, the right answer to the board is the number, not the architecture. Price the standby — replica compute, storage, cross-region replication traffic, and the engineering hours the drills consume — against the modelled loss, and present both. A board that asked for zero data loss after reading about somebody else's outage will usually accept a measured recovery point and a rehearsed drill, and the team that produced the number keeps control of the roadmap.