intermediate 3 min answer

A checkout path makes 40 calls across module boundaries to render one page. Roughly what does that path cost in a modular monolith versus the same modules extracted as services in one availability zone, and what does the number tell you about where to draw boundaries?

modular-monolithmicroserviceslatency-budgetcall-amplificationestimation
Show the full answer Hide the answer

The assumptions, stated

  • An in-process call across a module boundary: a virtual dispatch and perhaps an argument copy. On the order of 100 nanoseconds, and not more than a microsecond unless the arguments are large.
  • A same-zone RPC: serialise, write to a socket, traverse the kernel and the network, deserialise, and the same again on the way back. On the order of 0.5 to 2 ms end to end for a small payload on a healthy network, taking 1 ms as the planning figure.
  • 40 boundary crossings on the path, sequential unless deliberately batched.

The arithmetic

In-process: 40 × 100 ns ≈ 4 microseconds. Invisible. The page's latency is whatever the database and the template engine cost.

Extracted: 40 × 1 ms ≈ 40 ms added at the median, before any of the work. At p99, where each hop may cost 5 to 10 ms because of garbage collection, connection churn or a retry, the same path is 200 to 400 ms of pure overhead. The tail is the number that matters, and it is four orders of magnitude worse than the in-process case rather than merely larger.

Availability compounds too. Forty sequential hops each at 99.95% gives roughly 0.9995^40 ≈ 98% for the path, which is about 14 hours of monthly failure budget consumed by topology alone.

Which assumption dominates the error

The call count, not the per-call cost. Halving RPC latency saves 20 ms; reducing the path to 5 boundary crossings saves 35 ms and most of the tail. This is why the first question about a service boundary is how many times a request crosses it, and why co-change analysis predicts boundary quality better than any domain diagram: modules that are called 40 times per request are one module that has been drawn as several.

What the number rules in or out

  • It rules out extraction along a chatty seam, whatever the domain argument for it. If the seam carries 40 calls per request, the only way to extract it is to change the interface first so it carries one or two coarse calls, and that interface change is the actual work.
  • It rules in extraction along a seam crossed once or twice per request, where 1 to 2 ms is a rounding error against a 300 ms budget.
  • It explains why "we will add caching" is not a plan. A cache turns some of the 40 hops into local lookups and leaves you with a cache invalidation problem and a cold-start cliff, in exchange for a latency you could have had by not splitting.
  • It sets a fitness function: assert the number of cross-module calls on the critical path in a test, and fail the build when it grows. In a monolith this costs nothing to measure and prevents the boundary rotting before anyone proposes extracting it.

When not to trust this arithmetic

If the 40 calls are independent, they can be issued concurrently, and the cost falls towards the slowest one plus coordination, which changes the answer entirely. If each call does substantial work, 1 ms of overhead on 50 ms of work is noise. And none of this arithmetic touches the reasons to extract that are not about latency: an independent scaling profile, a separate failure domain, a different compliance boundary, or a team that needs to ship without coordinating. Extraction is bought with latency and operational cost, and those reasons are worth the price. "Microservices are modern" is not.