beginner 2 min answer

A mobile screen issues eight backend calls one after another because each needs the previous result. Server time is 25 ms per call and the network round trip is 40 ms. Why does the screen take over half a second, and why does making the servers twice as fast barely help?

latencyround-tripswaterfallmobilebeginner
Show the full answer Hide the answer

The arithmetic first

Each call costs one round trip plus the server's work: 40 + 25 = 65 ms. Eight of them in sequence is 520 ms, of which 320 ms is network round trips and only 200 ms is server work.

Latency adds along a serial chain, and the unit of serial cost is the round trip, not the function call. Nothing in the backend's own metrics shows this. Each service honestly reports a p50 of 25 ms and a healthy dashboard, because a server cannot measure the time its own response spends in flight or the time the client spent waiting to be allowed to ask.

Why doubling server speed disappoints

Halve the server work and each hop becomes 40 + 12.5 = 52.5 ms, so the screen takes 420 ms. Weeks of optimisation, 19% better, and the user still waits. The 320 ms of round trips is untouched, because it is a property of the call graph's shape.

Removing one layer of the chain beats optimising all of it:

  • Collapse the chain server-side. One request to an endpoint that performs the eight dependent steps inside the data centre costs 40 ms of round trip plus 200 ms of work: 240 ms, a 54% cut, with no code made faster.
  • Parallelise the calls that are not actually dependent. Teams usually find that only three of the eight genuinely need their predecessor. Three serial hops of 65 ms is 195 ms.
  • Overlap the first paint with the chain so the user sees something after one round trip.

On a mobile network the case is stronger, not weaker: a round trip of 60–150 ms is ordinary, so the same eight hops cost 0.9–1.6 s, and a cold start adds TLS and connection setup on top.

The decision rule

Compare hops × round trip against the sum of server time before optimising anything. If round trips dominate, the work is the call graph. If server time dominates, profile the servers. In this scenario the ratio is 320:200 and the answer is obvious once it is written down.

When this is the wrong answer

If the chain is two calls and each server takes 400 ms, the servers are the problem and collapsing the chain saves one round trip out of 840 ms. Chain collapsing is not free either: a composite endpoint couples the client's screen to one server-side aggregate, becomes harder to cache (its response is per-user and changes whenever any input changes), and turns eight independently deployable calls into one release train. The flip condition is the ratio above, measured at the client.

Common weak answers

  • "Put a CDN in front." The data is per-user and request-specific, so every call is a cache miss with an extra hop added.
  • "Switch to HTTP/2." Multiplexing removes head-of-line blocking between concurrent requests. A dependent chain has no concurrency to multiplex; the eight round trips remain eight round trips.
  • "The backend is fast, so this is a frontend problem." Both measurements are honest. The failure is that nobody measured the composition, and that is the lesson.