concept

Client Request Waterfall

also called Fetch Waterfall, Client-Side N Plus One

A chain of requests where each one can only start after the previous response arrives, so the screen's load time is set by the depth of the chain times the round trip rather than by bandwidth.

latencyround-tripsbffapi-shapesanti-pattern

A profile screen takes 1.4 seconds on a connection that downloads the whole payload in 80 ms. The server traces show every endpoint answering in under 30 ms. Nothing is slow, and the screen is slow.

The cause is structural: the client fetches the user, then that user's orders, then each order's line items. Each level can only begin once the previous one has returned, so the floor on load time is depth multiplied by round-trip time. Three levels at 80 ms is 240 ms before any server work; add a per-item fan-out issued serially and it is a second. This is an anti-pattern rather than a technique, and it is the frontend form of the N+1 query.

Why it matters

Bandwidth has improved far faster than latency, and latency is the thing a waterfall spends. Compressing the responses, upgrading to a newer HTTP version or adding a CDN leaves the depth untouched, which is why teams who optimise bytes first report no improvement and conclude the problem is elsewhere.

It also hides inside component trees. When a component fetches its own data, the render tree becomes the request graph, and nobody can predict the waterfall from reading either one. A list of 24 rows each fetching an author is 24 requests that no single file mentions.

Implementation patterns

  • Collapse depth on the server. One screen-shaped endpoint, a backend-for-frontend, or one graph query. The fan-out still happens, inside a data centre where a hop is sub-millisecond rather than 80 ms.
  • Parallelise what is not genuinely dependent. Promise.all over independent resources converts depth into width, so the cost becomes the slowest single request.
  • Hoist fetching out of leaf components to the route or page level, so the full set of requests is visible in one place and can be issued at once.
  • Declare data requirements statically — a route loader, a server component, or a query declared alongside the component and executed by the router — so the framework can start every request before rendering.
  • Push the first level with the document. Inline the initial payload into the server-rendered HTML so the browser does not have to parse a bundle to learn what to fetch.
  • Deduplicate in-flight requests by key in the data layer, which removes the accidental width that follows from component-level fetching.

Industry example

The documented context for progressive web app work aimed at emerging markets makes the arithmetic plain: the published Flipkart Lite case study on web.dev reports that 63% of those users reached the experience over 2G, where round-trip time runs into the hundreds of milliseconds. At a 400 ms round trip a four-level waterfall is 1.6 seconds of nothing, which is the kind of number that explains why that rebuild reported time on site rising from about 70 seconds to about 3.5 minutes and data usage falling threefold. The fix in those conditions is structural, not incremental.

Failure scenarios

  • Component-level fetching in a list, producing one request per row with no code path that says so.
  • Authentication or configuration as level one, so every screen inherits one extra serial round trip.
  • A sequential loop with await inside it, which looks parallel and is not.
  • A waterfall discovered only on mobile, because at 20 ms round trip on fixed lines the same chain costs 80 ms and nobody notices.
  • An aggregation layer that itself waterfalls internally, moving the problem rather than removing it.
  • Suspense boundaries that make a waterfall render pleasantly, so it is never fixed.

Trade-offs

Collapsing a waterfall into one endpoint buys round trips and pays coupling: the endpoint now knows the screen's shape, so a screen change needs a server change unless the client team owns that endpoint. A graph query buys flexibility and pays a query-cost problem — depth limits, persisted queries and resolver fan-out become your concern.

Parallelising is almost free and limited: it cannot remove a dependency that genuinely exists, and a browser caps concurrent connections per origin, so 40 parallel requests queue anyway.

When not to use it

Do not build an aggregation layer for a two-level chain on a fast network. Two levels at 30 ms is 60 ms, and an extra service to operate is worse than 60 ms. Measure the depth and the round-trip distribution of real users first: if the p75 round trip is 30 ms and the chain is two deep, the waterfall is not your problem and the bundle probably is.

Interview question

Q: A screen makes 11 requests and takes 1.8 seconds on mobile. How do you find out whether the problem is depth, width or payload, and what would you change first in each case?

What a strong answer covers: read the network panel as a dependency graph rather than a list — depth shows as staircasing, width as a wall of parallel bars, payload as long bars. Depth is fixed by aggregation or hoisting; width by connection limits and prioritisation; payload by field selection. A strong answer also asks who owns the endpoints, because that decides whether aggregation is a one-week or one-quarter change.

Quick check

Quiz: Why does halving every response size not help a waterfall? Because the cost is depth times round-trip latency, and payload size is not in that product.

Flashcard: What sets the floor on a waterfall's load time? — The number of dependent levels multiplied by the round-trip time, independent of bandwidth.