metric

Serial Round-Trip Depth

also called Critical-Path Hop Count, Request Waterfall Depth

The number of dependent network round trips on a request's critical path, which multiplies by the round-trip time to set a latency floor that no amount of server optimisation removes.

latencyround-tripswaterfallmobilecritical-path

A screen takes 520 ms to render. Every backend involved reports a p50 of 25 ms and a green dashboard. Both measurements are honest, and the gap between them is structural: the screen makes eight calls one after another because each needs the previous result, and each dependent call costs a full round trip before any server work begins.

Eight hops at a 40 ms round trip is 320 ms of waiting that no server can see, measure or optimise. Serial round-trip depth is the count of those hops: a property of the call graph, and the first number to establish when a system is slow and every component looks fast.

Why it matters

Latency adds along a serial chain while throughput does not, so the instincts that work for throughput mislead here. Doubling every server's speed above takes the screen from 520 ms to 420 ms: a 19% gain for the hardest work available. Collapsing the eight hops into one server-side call takes it to 240 ms with nothing made faster.

The floor is physical. Light in fibre travels at roughly 200,000 km/s, so 6,000 km is about 30 ms one way and 60 ms round trip before any queueing. On a mobile network a round trip of 60–150 ms is ordinary and a cold start adds connection and TLS setup. Depth × round-trip time is a budget item you cannot optimise away, only restructure — and protocol work does not rescue it, because HTTP/2 (RFC 7540, 2015) removes head-of-line blocking between concurrent requests and a dependent chain has no concurrency to multiplex.

Implementation patterns

  • Measure depth at the client, not the server. A browser waterfall or a mobile trace shows the chain; server dashboards structurally cannot.
  • Separate genuine dependencies from habitual ones. Teams routinely find that three of eight calls need their predecessor and five were sequential by convention. Those five become one parallel batch.
  • Collapse remaining chains server-side, where the hop cost is sub-millisecond rather than tens of milliseconds, via a composite endpoint or a graph query executed in the data centre.
  • Budget depth explicitly: "no screen exceeds three dependent round trips" is enforceable in review because a reviewer can count hops, which a latency target never is.

Industry example

The pattern is most visible in mobile clients talking to a microservice estate. A booking screen in a ride-hailing or food-delivery app needs the user, saved addresses, pricing, promotions, availability and payment methods. Implemented naively each is a separate call, several of them dependent, and the app is slow on a good network and unusable on a weak one. The industry's answer has been the aggregation layer — a backend-for-frontend or a graph gateway — whose primary justification is reducing client round trips, not reducing server work.

Failure scenarios

  • A refactor that splits one service into three and silently turns one hop into three, with every service's own latency metric improving.
  • An authentication or feature-flag lookup added at the front of the chain, adding a round trip to every screen at once.
  • A composite endpoint that becomes the dependency bottleneck, because it fails whenever any of its eight inputs fails and the error probability compounds across them.

Trade-offs

Collapsing a chain buys round trips and pays in coupling. A composite endpoint ties the screen to one server-side aggregate, becomes harder to cache (its response changes when any input changes), merges several independently deployable services into one release path, and needs a partial-failure policy for each input.

Parallelising instead keeps the services independent and shifts the cost to tail latency: the screen waits for the slowest of five calls, so its p99 approaches the worst component's p99 rather than the average. Usually the better trade, and one to make knowingly.

When not to use it

When server time dominates, depth is the wrong focus. Two hops with 400 ms of server work each means 80 ms of round trips inside 880 ms, and the work is profiling, not restructuring. Compare depth × round-trip time against the sum of server time, and let that ratio choose the project.

Inside a data centre the same arithmetic applies with a round trip of 0.2–1 ms, so a 20-hop internal chain costs single-digit milliseconds of network and is rarely worth collapsing for latency, though it may be for failure probability.

Interview question

Q: "A screen takes 600 ms. Every service in the path reports a p50 under 30 ms and no error budget is burning. Where is the time, and what is the first thing you would change?"

What a strong answer covers: measuring at the client to count dependent hops; multiplying depth by the observed round-trip time to get the structural floor; separating real dependencies from conventional ones and parallelising the rest; collapsing server-side only where the coupling is acceptable; and the tail-latency cost that parallelising introduces.

Quick check

Quiz: Eight dependent calls, 40 ms round trip, 25 ms of server work each. How much of the 520 ms can server optimisation ever remove? At most 200 ms, because 320 ms is round trips.

Flashcard: Why does a backend dashboard never show a waterfall problem? — A server cannot measure the flight time of its own reply or the client's wait before asking, so every component looks fast while the composition is slow.