practice

Edge Policy Ordering

also called Gateway Filter Order, Admission Sequence

The sequence in which a gateway applies connection limits, authentication, quotas and validation, which decides whether a flood of unauthenticated traffic consumes your cheapest resource or your most expensive one.

api gatewayrate limitingauthenticationload sheddingfilter chain

Two gateways have the same filters configured and behave completely differently under attack. One authenticates the request and then applies the tenant's rate limit. The other applies a per-address request limit, then authenticates, then applies the tenant quota.

Under a flood of unauthenticated requests the first gateway spends a signature verification, and possibly a key lookup, on every junk request before rejecting it. The second rejects most of the flood for the cost of a counter increment. The filter list is identical. The order is the design.

Why it matters

A gateway's job is admission, and admission control only works if the cheap decisions come first. The costs differ by orders of magnitude: a local counter increment is sub-microsecond, a signature verification is tens to hundreds of microseconds, a token introspection call is a network round trip of several milliseconds, and body validation scales with the body.

Ordering is also forced by information. A per-tenant quota cannot run before authentication, because the tenant identity comes from the credential. So the sequence is not a free choice: identity-free controls first, authentication next, everything needing identity after. Teams who learn this during an incident record it as "our auth service fell over during a traffic spike", which is the correct diagnosis of the wrong component.

Implementation patterns

  • Connection and request-rate limits by source address first, in the proxy, with no lookup. This is the only control that works against traffic carrying no valid credential.
  • Request size and shape checks next, so a 40 MB body is rejected before anything parses it.
  • Authentication after that, with signing keys cached in memory and refreshed in the background. A key fetch on the request path makes the key service a dependency of every request.
  • Per-tenant and per-endpoint quotas after authentication, with counters in a shared store behind a local pre-check so the common case pays no network hop.
  • Schema validation last, being the most expensive per byte and the least likely thing an attacker trips.
  • Make the order explicit in configuration and test it, because most gateways apply filters in declaration order and a refactor can reorder them with no functional test failing.
  • Label the rejection reason on the metric. Shedding at the auth stage rather than the connection stage means the order is wrong.

Industry example

The layering is visible in any public API platform's documented edge: network-tier connection limits, then token validation, then plan-level quotas keyed to the authenticated account. A ticketing platform running a scheduled on-sale is the archetype where the order bites hardest, because the spike arrives from a large population of clients with stale or absent credentials, and a gateway that authenticates first turns a traffic spike into an identity-service outage that then fails every legitimate request too. The same shape recurs in incident write-ups from 2019 onward.

Failure scenarios

  • Authentication before shedding, so the identity service becomes the capacity ceiling of the platform.
  • Key material fetched on the request path, so a slow key endpoint stalls every request and exhausts the worker pool within one cache TTL.
  • Quota checks that each make a network call, adding a round trip to every request to protect against a minority of them.
  • Validation before authentication, so an unauthenticated attacker spends your CPU on parsing.
  • One global rate limit shared across routes, letting a failing route eat the healthy routes' allowance.

Trade-offs

Choose Gains Pays
Cheap identity-free controls first Floods are rejected for the price of a counter Legitimate bursts from one address are hit before anyone looks at their plan
Authenticate early Every policy can be per-tenant and precise The most expensive operation runs on traffic you were going to reject

The real cost of ordering correctly is that per-address limits are blunt: they punish a large corporate NAT or a mobile carrier gateway alongside an attacker. The mitigation is a generous identity-free limit sized to shed only clear abuse, with precise limits applied after authentication.

When not to use it

On an internal gateway inside a trusted network where every caller is authenticated by mTLS at the transport layer, the question largely disappears - identity arrives with the connection and costs nothing per request.

It also does not help when the bottleneck is downstream: if the gateway is cheap and the backend is the constraint, reordering filters changes nothing and the control you need is a concurrency limit per upstream. And for one client and modest traffic, a single authentication middleware is the right amount of machinery.

Interview question

Q: Your gateway validates a JWT and then applies a per-customer rate limit. During a traffic spike the auth path saturates and legitimate customers get errors. What do you change, and what do you deliberately leave in place?

What a strong answer covers: moving an identity-free connection and request-rate limit ahead of authentication; caching signing keys with background refresh; keeping the per-customer quota after authentication because it needs the identity; sizing the identity-free limit generously to avoid punishing shared egress addresses; and a rejection-reason label so the shedding stage is observable.

Quick check

Quiz: Why can a per-tenant quota never be the first filter in a gateway chain? Answer: The tenant identity comes from the credential, so the quota cannot be evaluated until authentication has run - which is why the only controls available earlier are identity-free ones.

Flashcard: What is the cost of authenticating before shedding? - Every junk request buys a signature verification and possibly a key lookup, so a flood saturates your most expensive edge operation instead of your cheapest.