advanced 4 min answer

A checkout read API is deployed as an edge function to 300 points of presence. Global traffic is 60 requests per second and follows population: the busiest location takes about 8% of it and the 150th-busiest takes about 0.1%. The runtime evicts an idle isolate after roughly 30 seconds. Roughly what share of requests at that 150th location pays an initialisation path and what does the number rule in or out?

edge-functionscold-startpoissoncapacity-modellinglatency
Show the full answer Hide the answer

The assumptions, stated

  • 60 rps globally; the 150th location takes 0.1%, so 0.06 requests per second there, one request every 17 seconds on average.
  • Arrivals are roughly independent at that rate, which is a fair model for low-traffic locations.
  • Idle eviction after 30 seconds. Treat this as an order of magnitude and measure your own runtime, because it is a platform behaviour a vendor can change.
  • The expensive part of a cold path is not the isolate. A V8 isolate starts in single-digit milliseconds; the cost is what the first request does: fetch configuration or a secret, build a TLS session to a regional database or API, warm a JWT verification key. Budget 100–300 ms for that first-request work.

The arithmetic

A request is cold if no request arrived at that location in the preceding eviction window.

λ  = 0.06 req/s
T  = 30 s idle eviction
λT = 1.8
P(cold) = e^(-λT) = e^(-1.8) ≈ 0.165

About one request in six at that location pays the initialisation path, so with a 200 ms cold path the location's p50 is fine and its p85 is roughly 200 ms worse than its warm latency. At the busiest location, λ = 4.8 rps gives λT = 144 and the cold share is effectively zero.

Then the floor: keeping a location warm needs roughly one request per eviction window, that is 1/30 ≈ 0.033 rps each. Across 300 locations that is about 10 rps of traffic spent purely on staying warm, and the traffic has to be distributed the way the locations are, which population-weighted traffic never is.

Which assumption dominates the error

The traffic distribution, by a wide margin. Concentrating the same 60 rps into 80 locations instead of 300 raises a mid-tail location's share by roughly 4x, which moves λT from 1.8 to 7 and the cold share from 16% to under 0.1%. The eviction window matters less: doubling it to 60 s only takes 16% down to 3%.

The second-order error is the cold path's own cost. If the function opens a connection to a database in one region, the cold path includes a cross-continent TLS handshake, which is three round trips and can be 400 ms from a distant location. The cold-start problem at the edge is usually a connection-establishment problem wearing a different name.

What the number rules in or out

  • Ruled in: pure edge work with no origin dependency at all. Signature verification, header rewriting, redirects, A/B bucket assignment, geo routing. These have no cold path worth measuring.
  • Ruled out at 60 rps across 300 locations: anything whose p99 is a product requirement and whose cold path touches a database. Restrict deployment to the 10–20 locations that carry most traffic, or keep the route in one or two regions behind a pooled connection.
  • A plain server round trip wins when the request must talk to a central store anyway. A user 150 ms from your region pays 150 ms once if the route is regional, and pays 150 ms plus the edge hop plus a cold connection if the route is at the edge and the edge is cold. Moving code closer to users does not move the data, and if every invocation reads the data, the edge adds a hop to the same journey.

The decision rule, and when this is the wrong analysis

Prefer the fewest locations that meet your latency target, unless the route has no origin dependency at all. Breadth is free only for stateless work; for anything that opens a connection, every extra location divides the traffic that keeps it warm.

The failure this arithmetic predicts is specific: on a cold path, the first request after an idle period or a deploy fails its own client timeout while the connection is established, so the symptom is a thin band of timeouts concentrated in low-traffic countries, invisible in global percentiles and reported as "the app does not work here".

If traffic is a spiky 10x-at-peak shape rather than a steady trickle, cold share at the trough is irrelevant and what matters is behaviour at the ramp. Model the first 30 seconds of a peak instead, where every location is cold simultaneously and the cold path hits a database that is also receiving the peak. Eviction windows and isolate behaviour are platform choices as of 2026 and vendors change them, so treat the published figure as a starting point and measure the one you actually have.