A nightly close job commits 120000 single-row transactions one at a time and finishes in about 2 minutes. After the team enables synchronous multi-zone replication on the managed database the same job takes about 5 minutes with CPU unchanged on both primary and standby. Estimate where the extra time went and say what the same arithmetic implies about synchronous replication across regions.
Show the full answer Hide the answer
The assumptions, stated
Zones in a region sit up to roughly 100 km apart - the published ceiling for zones inside a cloud region, current in 2026. Light in fibre travels at about 200,000 km per second, so propagation alone costs about 5 microseconds per kilometre: 0.5 ms one way at the full 100 km, 1 ms for a round trip, before a single switch hop. Measured cross-zone round trips land between 0.3 ms and 2 ms once switching and host networking are added. Take 1.5 ms as the working number.
A synchronous commit cannot return until the standby has acknowledged, so every commit pays at least one cross-zone round trip. Nothing else in the job changed.
The arithmetic
- Before: 120,000 commits in 120 s, so roughly 1 ms per commit, which is a local fsync-bound write.
- Added per commit: ~1.5 ms of round trip.
- Added in total: 120,000 × 1.5 ms = 180 s = 3 minutes.
- Predicted: 2 + 3 = 5 minutes. That matches the observation, so the model is the right one.
Range: at 0.5 ms the job gains 1 minute, at 2 ms it gains 4. The round trip dominates the error and the row size is irrelevant — a 200-byte row at 1 Gbps serialises in about 1.6 microseconds, three orders of magnitude under the round trip. Synchronous replication prices commits, not bytes. The derived ceiling is useful on its own: a single connection committing serially cannot exceed 1/RTT, so roughly 500 to 2,000 commits per second per connection, whatever the hardware underneath.
What the number rules out
Run the same arithmetic across regions. A US coast-to-coast round trip is 60 to 80 ms. The same job with synchronous cross-region commit is 120,000 × 70 ms ≈ 8,400 s, about 2 hours 20 minutes, a 70× regression. This is the whole reason multi-zone synchronous commit is a sensible default and cross-region synchronous commit is not, and why an honest cross-region posture is asynchronous replication with a stated and measured recovery point rather than a promise of zero data loss.
The fix for the job is not a faster network. Batch it: 120,000 single-row transactions become 1,200 transactions of 100 rows, the round trip is paid 1,200 times instead of 120,000, and the job returns to roughly 2 minutes with the replication guarantee intact. The lever on synchronous replication cost is always transaction size and write concurrency.
When this is the wrong estimate
The model breaks when the commit is not the serialised step. A job that commits from 32 parallel connections hides the round trip behind concurrency and shows almost no regression, so measuring one thread and extrapolating overstates the cost. It also breaks where the standby is in the same zone, where the replication is semi-synchronous (acknowledged on receipt rather than on apply), or where the engine groups commits — group commit amortises the round trip across whatever arrives in the same window, which is why a busy database degrades less than a single-threaded batch job. Measure commits per second, not the wall clock of one job, before concluding that the topology is the problem.