Your deployment diagram shows three availability zones with every service replicated into each and a mesh spreading calls evenly across zones. Finance asks why network charges are $14k of a $90k monthly cloud bill. Estimate the cross-zone traffic implied by that number and say what you would annotate on the diagram.
Show the full answer Hide the answer
The assumptions stated
- Inter-zone transfer inside one region is charged at roughly \(0.01 per GB in each direction** on the major clouds at 2025 list prices, so a gigabyte that crosses a zone boundary costs about **\)0.02 in total.
- Three zones with placement-blind routing means two of every three internal calls cross a zone boundary.
- A front-end request fans out to about 8 internal calls, each returning on the order of 20 KB.
The arithmetic
$14,000 a month at $0.02 per crossed gigabyte is 700 TB of cross-zone transfer a month. That is about 270 MB/s averaged over the month, or roughly 2 Gbit/s of steady cross-zone chatter — a number worth sanity-checking against interface counters before going further.
Cross-zone bytes per front-end request: 8 calls × ⅔ crossing × 20 KB ≈ 107 KB. So 700 TB implies about 6.5 billion front-end requests a month, near 2,500 requests per second sustained. If the real traffic is 2,500 rps, the estimate holds together and the cost is structural rather than a leak. If the real figure is 250 rps, something is transferring far more than the user-facing path — replication, a chatty cache or a log shipper is the likely culprit and the diagram will not show it until you annotate it.
Which assumption dominates the error
Payload size, by a wide margin. Internal response sizes are long-tailed: a handful of endpoints returning a 400 KB object can account for most bytes while being a few percent of calls. The ⅔ crossing fraction is nearly exact under random placement, and the price is published. So the honest range on the request estimate is about 3× in either direction, and narrowing it means measuring bytes per route, not guessing harder.
What to annotate on the diagram
Put GB per month on every edge that crosses a zone boundary, and mark each crossing as mandatory or incidental. Quorum writes and replication are mandatory: the crossing is what buys zone survivability. Service-to-service reads are incidental, and zone-affinity routing — prefer the local replica, fail over to a remote one on health or latency — removes most of them. On the numbers above, pushing the crossing fraction from ⅔ to under 0.1 saves on the order of $10k a month, and costs you an uneven load profile per zone plus a harder failover test.
When this is the wrong answer
If the 700 TB turns out to be synchronous database replication, zone affinity cannot touch it without changing the recovery-point objective, and nobody should trade RPO for $10k. Then the levers are compression, batching and moving read replicas rather than routing. Equally, at a $9k bill rather than $90k, the two weeks of engineering time this work costs is worth more than the saving: annotate the diagram, do not re-architect.
Common weak answers
- "Collapse to one zone." It removes the charge and the availability the three zones were bought for.
- "Turn on the mesh's locality setting." It is the right lever and it is not free: locality routing concentrates load, so a zone losing capacity now shifts a 50% step change onto its neighbours. Test the failover before trusting the saving.
- "Ask the provider for a discount." Committed-use discounts rarely apply to inter-zone transfer, and the diagram still cannot tell anyone which edges are the expensive ones.