A team moves request signing from one region to an edge platform. p50 falls from 140 ms to 35 ms and p99 rises from 220 ms to 900 ms. Error rate is unchanged and the platform reports no throttling. Which mechanism most likely explains the tail?
Show the full answer Hide the answer
The first three things to look at
- Origin requests divided by edge invocations. If the ratio is near 1 the edge is a proxy with extra hops rather than a decision point.
- Edge-to-origin latency reported separately from end-to-end latency, so the two distances are not added together in one histogram.
- Hit ratio on whatever the function reads, per point of presence rather than globally.
The diagnosis
Moving code to the edge removes the user's distance to the code. It does not remove the code's distance to the data. p50 improved because most invocations found the signing key already cached in the local point of presence. The tail is the misses: a miss pays the user-to-PoP hop plus the full PoP-to-origin round trip, which is roughly the latency the entire request used to cost, and it is paid by whichever request arrives first in each location after the cached value expires.
Distribution makes this worse rather than better. Running in 100 points of presence multiplies the number of cold caches by 100. With a 60-second key lifetime you generate up to 100 origin fetches a minute and up to 100 slow requests a minute. At a few thousand requests per second that is well under 1% of traffic, which is exactly where a p99 lives and exactly why p50 looks like a success.
Why the other options fail
Cold starts. The right instinct from container-based functions, where a cold start is hundreds of milliseconds, and wrong at the edge. Edge runtimes are built on isolates or ahead-of-time-compiled WebAssembly specifically to remove this: Fastly published a figure of 35.4 microseconds for Compute@Edge instance startup, four orders of magnitude below the 700 ms being explained. A start-up cost measured in microseconds cannot produce a 900 ms tail.
Silent throttling. A genuine failure mode on every edge platform, which surfaces as 429s or a platform-reported concurrency metric. The stem states error rate is unchanged and no throttling is reported, so chasing this means ignoring the telemetry you already have.
Users further from a PoP. This would have raised p50 as well, and p50 fell by 105 ms. A distance problem moves the whole distribution, not only its tail, which is the general rule that separates topology bugs from cache bugs.
The alert that would have caught it earlier
Alarm on origin fetches per edge invocation, not on edge latency. For a function that is supposed to decide locally, anything above a few per cent means the data did not move with the code. The fix here is a deploy-time secret pushed to every location instead of a runtime fetch, after which the tail disappears.
When not to move the workload at all
Keep the work in the region that owns the data when the decision needs a value that must be fresh and globally unique: an account balance, a rate-limit counter that must not over-admit, a monotonic sequence. The edge is the right home for a decision computable from the request itself plus data that tolerates being seconds stale — a signed token, a cookie, a routing rule, a feature flag. Push a key to the edge and you remove a round trip; push a balance to the edge and you have built a correctness bug that only appears under concurrency.