A video platform serves a catalogue of 4 million objects where roughly 90% are requested less than once an hour in any one location. The per-location edge cache hit ratio is reported as 93% and origin egress is still the largest line in the infrastructure bill. Akamai documents Tiered Distribution - edge servers fetching from a parent tier near the origin rather than from the origin itself - with a map setting that is explicitly a trade-off. What forced that design and what does a narrower map cost?
Show the full answer Hide the answer
The situation a long tail puts a CDN in
A per-location hit ratio hides a multiplier. Each point of presence maintains its own cache and fills it independently, so a cold object requested once per hour in each of 300 locations produces up to 300 origin fetches per hour for one object. Every one of those locations can truthfully report a 93% hit ratio while the origin sees the sum of all their misses.
The arithmetic is the whole story. At 10,000 requests per second spread over 300 locations with a 7% miss rate, origin sees roughly 700 requests per second. Doubling the number of locations improves user latency and makes the origin problem worse, because the same miss traffic is now split across twice as many independent caches and the chance that a given location already holds a long-tail object falls.
What Akamai documents
Akamai's Tiered Distribution behaviour lets edge servers fetch cacheable content from parent servers that sit along the path close to the origin, instead of going to the origin directly (Akamai techdocs, accessed 2026). The parent tier aggregates the misses of many edge locations, so an object fetched once into the parent serves every subsequent edge miss from inside the network.
The documented knob is the tiered distribution map, and Akamai states the trade-off directly: a narrower map reduces origin load and makes a parent cache hit more likely, while a wider map lowers end-user latency and makes a hit at any given parent less likely. That sentence is the general law of cache hierarchies, written as a configuration option: concentration buys offload, spread buys latency, and you cannot have both from one tier.
What it costs
- One extra hop on every miss. An edge miss now pays edge → parent → origin. If the parent is near the origin that added leg is cheap; if your map puts the parent on another continent from the origin, cold requests get slower while the bill improves.
- A concentration point. The parent tier carries the aggregated miss traffic of many locations, which makes it a hot spot and a correlated failure domain. A parent that is evicting under pressure converts into origin load that arrives all at once.
- Two tiers to invalidate. A purge that clears the edge but not the parent re-populates the edge from a stale parent, which is one of the two classic causes of "the purge returned 200 and customers still see the old price".
- Harder cache reasoning. Age, TTL and stale-while-revalidate now compose across two layers, so the oldest content a user can receive is the sum of the tiers rather than the edge TTL.
When copying this is the wrong answer
Shielding pays only for cacheable content with a long tail. Three cases where it does not:
| If this is true | Do this instead | Because |
|---|---|---|
| Responses are personalised or uncacheable | Nothing at the cache layer | A shield adds a hop and offloads zero bytes |
| Audience and origin are in one region | Fewer larger locations | Few caches means little miss multiplication to collapse |
| One object goes viral in minutes | Request collapsing at the edge | Collapsing solves concurrent misses for one key; shielding solves spread misses across many keys |
Request collapsing and shielding are different mechanisms and teams routinely buy one expecting the other. If your symptom is thousands of simultaneous origin requests for the same new object, you need collapsing. If your symptom is steady origin load for millions of rarely-requested objects, you need a parent tier.
What a strong answer adds
Measure origin requests per object per TTL, not hit ratio. Hit ratio is a per-location statistic and the origin bill is a fleet-wide one, so the two can move in opposite directions while a dashboard says everything improved.