Edge Capability Gradient
also called Point-of-Presence Capability Ladder, Edge Tiering
The systematic decline in what a location can do as it moves closer to the user - thousands of points of presence with constrained runtimes and no writable store at one end, a few dozen full regions at the other - which is what decides whether a workload can move to the edge, rather than the latency saving.
An executive reads that a delivery network has thousands of locations and asks for the product to run "at the edge". The number is real: Akamai's FY2025 annual report describes more than 4,300 edge points of presence in over 130 countries and roughly 700 cities as of 31 December 2025, across roughly 1,200 partner networks. A hyperscaler runs on the order of 30 to 40 regions.
The two numbers are not comparable, because what a location can do falls as its count rises. A point of presence terminates TLS one hop from the user and runs a small function in a constrained runtime. It has no primary database, little memory per request, no shell to debug on, and a deploy model you do not control. A region has the full service catalogue and a writable store. That gradient, not the latency saving, decides which workloads can move.
Why it matters
Teams estimate the benefit and not the constraint. Removing 100 ms of network distance is visible in user metrics; the constraint is where projects fail, and it is almost always data.
Moving code to the edge removes the user's distance to the code and leaves the code's distance to the data unchanged. A function that must read central state pays the full point-of-presence-to-origin round trip whenever its local cache misses, which is roughly what the whole request used to cost. Distribution makes this worse: running in 100 locations multiplies the number of cold caches by 100, so a 60-second cached value can generate up to 100 origin fetches and 100 slow requests per minute. At a few thousand requests per second that is under 1% of traffic, which is exactly where p99 lives.
Implementation patterns
- Three tiers, written down. At the point of presence: TLS termination, caching, routing, header and token work, bot checks, and any decision computable from the request plus a small slowly-changing dataset. In a region: the writable database, transactional work, accelerator inference. In one home region: anything that must be globally unique or strongly consistent.
- Carry the decision in the request. Compute a derived value centrally and ship it to the edge in a signed cookie or token — a cohort identifier, a flag set, an entitlement bitmap. Browsers cap a cookie at about 4 KB, a generous budget for a decision and no budget at all for a profile.
- Push secrets and configuration at deploy time, not at request time, so the common path has no origin dependency.
- Alarm on origin fetches per edge invocation. For a function that should decide locally, anything above a few per cent means the data did not travel with the code.
- Stage activation per location, because edge configuration reaches thousands of sites in seconds and the edge is the one tier where a bad push has no natural blast-radius boundary unless one is built.
Industry example
Cold start is the constraint people expect at the edge and it is the one the platforms removed. Fastly published a figure of 35.4 microseconds for Compute@Edge instance startup, achieved by compiling WebAssembly ahead of time so there is nothing to compile when a request arrives — four orders of magnitude below a container-based function's cold start of hundreds of milliseconds.
The lesson is about where to look. On a platform with microsecond start-up, a tail of hundreds of milliseconds cannot be start-up cost, so an edge function with a bad p99 is almost always fetching something. Teams arriving from container-based edge functions carry the cold-start mental model across and spend weeks optimising the wrong thing.
Failure scenarios
- The profile replication proposal. A 2 KB record for 40 million users is 80 GB per location before indexes, times thousands of locations, plus write fan-out on every change.
- Eventual consistency discovered in production. A globally replicated edge key-value store converges in seconds, and a feature that assumed read-your-writes produces complaints that never reproduce locally.
- Residency breach by footprint, where a function deployed across 130 countries processes personal data in jurisdictions the assessment never considered.
- Debuggability collapse, with the incident in a location that has sampled logs and no shell.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Decide at the point of presence | One hop to the user; origin load removed; attack traffic absorbed far away | Constrained runtime; no writable store; weak debugging; a wide residency surface |
| Decide in a region | Full capability; strong consistency available; ordinary operability | The user's round trip on every request that cannot be cached |
When not to use it
Keep the work in the region that owns the data when the decision needs a value that must be fresh and globally unique: an account balance, a rate-limit counter that must not over-admit, a monotonic sequence, a seat reservation. Pushing those to the edge trades a visible latency gain for a correctness bug that appears only under concurrency.
Skip the edge tier entirely when the response is already cacheable, because a cache hit gives the same latency with none of the compute surface; and when the user population sits near one region, because a national product served from an in-country region has only 10 to 30 ms of network distance to remove and a point of presence will not recover much of it.
Interview question
Q: An executive wants personalisation moved to the edge this quarter after reading that a delivery network has thousands of locations. You have fifteen minutes. What do you establish, what do you commit to, and what do you refuse?
What a strong answer covers: convert the location count into the capability question immediately, then ask what the decision reads, how stale it may be, which jurisdictions may hold it, and who debugs it at 03:00. Commit to one slice — cohort assignment at the edge from a signed cookie, with the cohort computed in-region nightly — and a measurement that decides whether more follows. Refuse the profile store at the edge and say why in arithmetic. Raise the operational asymmetry: a bad configuration push at the edge has no blast-radius boundary unless staged activation is built, and that cost belongs in the estimate.
Quick check
Quiz: An edge function's p50 improves and its p99 triples. Why is cold start the wrong first hypothesis on a WebAssembly edge platform? Answer: instance start-up there is measured in microseconds — Fastly published 35.4 microseconds for Compute@Edge — so it cannot account for a tail of hundreds of milliseconds; the cause is almost always an origin fetch on a cache miss.
Flashcard: What does moving code to the edge not move? — The data. Proximity to the user is gained and proximity to central state is unchanged, so every per-request origin dependency puts the original latency straight back into the tail.