intermediate 3 min answer

A travel platform with Booking.com's catalogue breadth replaces nine client-side calls with one backend-for-frontend call that fans out to the same nine services. Mobile p99 improves by 40% on the first day. Each of the nine services holds 99.9% availability. What has the team given up, and when does the bill arrive?

bffaggregationavailabilitypartial renderingfallbacks
Show the full answer Hide the answer

What is gained, quantified

One round trip instead of nine, which on a poor mobile network is the dominant term: connection setup alone runs 100 to 300 ms there, and nine serial-ish round trips is where the 40% came from. A payload shaped for the screen rather than for every client, which typically takes a generic aggregate of a hundred-odd kilobytes down to around ten. And composition logic that ships on the server's release cadence rather than waiting for an app-store review, which is the benefit teams undervalue and keep the longest.

What is paid

Availability multiplies where it used to be independent. Nine dependencies at 99.9% each, all required for the response, gives 0.999^9 ≈ 99.1%. That is roughly 6.5 hours a month during which the screen renders nothing. Before the change, one dependency's bad 43 minutes cost one section of the screen and nothing else.

The team converted nine independent partial failures into one correlated total failure, and the latency metric improved while it happened. That is the whole trade in one sentence, and it is why the day-one dashboard looks like a win.

The latency tail moves the same way, and that part is well known: nine parallel calls each 99% under their p99 means only 0.99^9 ≈ 91% of responses have all nine inside it, so the aggregate's p90 sits near the dependencies' p99 and you need p99.9 from each to hold the aggregate's p99.

The quieter loss is the one users feel: the client gave up progressive rendering. Nine calls let the screen paint prices while reviews were still loading. One aggregated call is all-or-nothing unless the BFF was deliberately built not to be.

When the cost becomes visible

On the first afternoon that one of the nine has a problem, which for nine dependencies is most months. Nothing in the BFF's own metrics predicts it, because on a good day the aggregate behaves like its slowest healthy dependency.

The second bill arrives with the tenth client. A BFF belongs to one surface by definition, so the web team builds a second one and the TV app a third, and a cross-cutting change is now three changes in three repositories on three release cadences.

How to keep the option to reverse

  • Classify all nine as required or optional at design time, and write the classification into the code rather than a diagram. Optional calls get their own deadline and a declared fallback - cached, empty, or a default - and can never fail the response.
  • Keep required calls to two or fewer. If a screen genuinely needs five services to be up simultaneously, the screen is the thing to redesign, not the BFF.
  • Return per-section status in the response so the client can paint what arrived. This is the single change that gives back progressive rendering.
  • Alert on response completeness, the share of responses carrying all nine sections, not on the BFF's error rate. A BFF that returns 200 with three sections missing is invisible to an error-rate alert and perfectly visible to users.
  • Keep per-call deadlines summing to less than the screen's budget, not equal to it, so the BFF has time left to serialise a degraded answer.

When not to build the BFF

Three or fewer dependencies, a client that can render progressively, and no payload problem: leave it on the client. A BFF earns its place when round trips dominate the latency budget, when the payload is genuinely wrong for the device, or when composition logic needs to change faster than the app can ship. It does not earn its place for being tidier, and the availability arithmetic above is the price of tidiness.