beginner 3 min answer

An API gateway validates every external token and strips the caller's credentials before forwarding, so the order service is configured to trust any request arriving on the internal network. Six weeks later a batch job owned by another team creates refunds against live accounts. Why did the gateway not prevent that, and what is the smallest change that does?

api-gatewayauthenticationauthorisationdefence-in-depthtoken-exchange
Show the full answer Hide the answer

The mechanism

A gateway can only make a statement about the traffic that goes through it, and the order service's port accepts traffic that does not. The gateway sat on one path into the service. The pod network, the batch cluster, a misconfigured job and any future internal service are other paths, and none of them pass the check.

There is a second, subtler gap. Even for requests that do arrive through the gateway, the gateway proved only that a valid token existed at the edge. It did not decide whether that principal may create a refund on that account, because it does not know what a refund is or which accounts the caller owns. Authentication is a property of the request; authorisation is a property of the request and the object, and only the service holds the object.

This is the confused-deputy shape, and in production it fails silently: the refunds were well-formed, authorised by the service's own rules, and indistinguishable in the logs from legitimate ones.

What the gateway can and cannot assert

  • Can assert: this credential was valid at 09:14, it belongs to tenant 4471, it has not exceeded 1,000 calls a minute, the body matches the published schema.
  • Cannot assert: that nothing else can reach the service · that this principal may act on this particular order · that an internal caller is who its header claims.

Stripping the credential made it worse. Once the gateway removes the token and the service trusts the network, the request carries no identity at all, so there is nothing for the service to check and nothing in its logs to attribute the refund to afterwards.

The smallest change that works

At the gateway, exchange the external token for an internal, signed, short-lived token — a 5 minute lifetime is normal — carrying the caller identity, the tenant and the scopes the edge actually granted. Every service verifies the signature and the audience before doing work, and rejects an unauthenticated internal call rather than assuming good intent. With the signing keys cached, a verification costs tens of microseconds, so this is not a latency decision.

The service then makes the one decision only it can make: does this principal have the right to refund this order. That check lives next to the data, not at the edge.

The division of labour that follows: the gateway owns the concerns that are the same for every service (terminate TLS, authenticate, rate limit, validate shape, exchange tokens); the service owns every decision that depends on its own data.

When this is the wrong answer

For a genuinely single-caller internal tool on a closed network — one client, one service, no multi-tenancy — a shared mutual-TLS identity plus a network policy that permits exactly that one source is sufficient, and issuing per-request tokens is ceremony. The thing that flips it is the second caller. As soon as more than one workload can open a connection to the port, "the network is trusted" has become an assumption nobody is testing, and the refund job is the test.

Common weak answers

  • "Put the gateway in front of internal traffic too." It helps, and it does not hold: every service still has a reachable port, and a mesh or gateway in the path proves the call came from somewhere, not that the caller may do this.
  • "Add an allow-list of internal IP ranges." The batch job was inside the range. Network location is not identity, and in a shared cluster it is barely even a hint.
  • "The batch team should have been more careful." It had the access the design gave it. A permission that exists only in a convention will be used.