advanced 3 min answer

Review this. An API gateway terminates TLS and validates tokens. Behind it a GraphQL gateway stitches 23 services. Behind that, four backend-for-frontend services - web, iOS, Android and partner - each call the GraphQL gateway, and the web client also runs a normalised client cache. Two clients carry 96% of traffic. Adding one field to one screen takes three pull requests across three repositories. What would you remove, what is the one change that matters, and what would you leave alone?

bffgraphqlaggregationarchitecture-reviewownership
Show the full answer Hide the answer

What is actually required

Three responsibilities, and they need two layers, not four.

  • One place that terminates TLS, validates tokens and enforces rate limits. That is the gateway.
  • One place per client surface that shapes a screen's payload and owns its own release cadence. That is a BFF.
  • Twenty-three services that own domain logic and nothing about screens.

The three-pull-request cost of one field is the measurement of the extra layer, not a culture problem. Every aggregation layer a field has to pass through is a repository, a review and a deploy.

What I would remove, and why it is safe

The GraphQL gateway, as a layer beneath the BFFs. It is a second aggregation point whose schema nobody owns end to end, it adds a network hop to every screen, and it makes a new field a change in the stitched schema and in the BFF that consumes it. Removing it is safe precisely because the BFFs already aggregate; they can call the 23 services directly with the same client libraries the gateway used.

The separate iOS and Android BFFs, if the two clients ship the same screens. One mobile BFF with a device parameter is less code and drifts less than two that are 90% identical. The condition matters: if the iOS surface genuinely diverges, split again, and expect to.

That leaves gateway → BFF → services. One field, one pull request.

The one change that matters

Give each BFF to the client team that consumes it, and make its contract screen-shaped. The layer count is a symptom; the cause is that the aggregation layer is owned by a platform team, so every client change becomes a cross-team request and the queue is the real latency. A BFF owned by someone other than its client is just another service with an extra hop.

What I would leave, even though it looks wrong

The normalised client cache, which looks like a third aggregation layer and is doing a different job: it deduplicates and invalidates data across screens within one session, which no server-side layer can do because no server knows which screens this tab has open. Removing it trades a measurable number of requests for an imagined simplification.

The partner BFF, despite partner traffic being 2%. Partner contracts carry deprecation windows measured in quarters and consumers you cannot redeploy. Coupling that surface to the web client's weekly release is how you end up unable to change the web client. The asymmetry in change cadence is the reason the boundary exists, and traffic share is the wrong metric for it.

How I would argue this in the review

With three numbers and no adjectives: pull requests per field change today versus after, the p99 contributed by the extra hop measured by removing it on one route behind a flag, and the number of engineers who can explain the stitched schema without opening it. Elegance arguments lose; a count of deploys per field wins.

When this review is wrong

If the 23 services have wildly inconsistent transport and auth conventions, the stitching layer is doing real normalisation work and removing it pushes 23 integration quirks into four BFFs. Then the correct order is different: standardise the service contracts first, and remove the gateway afterwards. Check before recommending removal — read two BFF handlers and count how much of their code is translation.

Choose removal only if the BFFs can call services with the libraries the gateway already uses, and keep the layer unless you can name its owner. The failure mode of getting this wrong has been visible in stitched-schema deployments since about 2018: the gateway becomes the slowest hop and the least understood artefact, so when a screen degrades nobody can say which layer broke.