Edge Rendering
also called Edge-Side Rendering, POP Rendering
Producing a page's HTML in a lightweight runtime at a CDN point of presence rather than in a regional data centre - which removes network distance from the response and adds distance to the data.
A page's time to first byte is 420 ms for users in Singapore and 90 ms for users near the origin in Virginia. The server renders in 40 ms. Almost all of that 420 ms is distance: TLS handshake plus the request and response crossing an ocean.
Edge rendering moves HTML production into a small runtime running at the CDN's points of presence, typically dozens to hundreds of locations. The response is produced a few milliseconds from the user, so the distance cost disappears from the response path — and reappears on every call the renderer makes back to a database that still lives in one region.
Why it matters
For a page whose content depends on nothing but the URL and a cookie, edge rendering converts a cross-ocean round trip into a local one, commonly cutting time to first byte by 100-300 ms for distant users. That is a real improvement in the metric that gates largest contentful paint, and it is available without changing the application's shape.
The decision rests entirely on what the renderer needs to read. A renderer that makes two sequential calls to a single-region database has not moved the work closer to the user at all: it has added a hop and left the distance in place, usually making the page slower than rendering in the region where the data is.
Implementation patterns
- Render only what is URL-derivable and cookie-derivable at the edge: layout, navigation, locale and currency selection, A/B assignment, redirects, auth gating, request rewriting.
- Stream the shell and defer the data. Flush the head and layout immediately, then stream the data-dependent regions as the central calls return, so the user sees progress during the round trip you could not remove.
- Replicate the read path, not the write path. Regional read replicas, an edge key-value store for configuration, or a per-POP cache of rarely changing documents. Writes continue to go to the primary region.
- Keep the cacheable version cacheable. Render the shared shell once, cache it, and vary a small number of segments at the edge rather than producing a per-user document.
- Respect the runtime's limits as a design constraint, not as an obstacle: isolate-based runtimes typically offer no filesystem, no raw TCP sockets, a subset of Node APIs and a bounded CPU-time allowance per request — so a library that shells out, or spends 300 ms transforming an image, does not belong there.
- Keep the edge function thin enough to reason about, because it runs in hundreds of places and is the hardest part of the stack to debug.
Industry example
The commercial edge runtimes — the isolate-based platform a large CDN shipped in 2017, and the frontend hosting platforms that later layered their edge runtimes on similar technology — all made the same architectural bargain: V8 isolates instead of containers, so start-up is on the order of a few milliseconds rather than hundreds, in exchange for a restricted API surface. That start-up number is what makes per-request rendering at the edge viable at all; a container-based function at 200-500 ms of cold start would spend more than it saves for exactly the distant users it targets.
Failure scenarios
- Data-dependent rendering at the edge, where each page makes two or three central calls and total latency rises.
- Personalising the document so cache-key cardinality explodes and the hit ratio goes to zero.
- A bug that is regional, reproducible only at one POP, with logs scattered across a hundred locations.
- A deploy that rolls out to the network progressively, so two versions of the renderer serve simultaneously and a client asset hashed by one is requested from the other.
- An exceeded CPU allowance appearing as a truncated response rather than a clean error.
- Secrets distributed to hundreds of locations, widening the surface for a configuration mistake.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Edge rendering | 100-300 ms off time to first byte for distant users | data round trips from the edge; constrained runtime |
| Regional rendering | one place to debug; full runtime; data is local | distance for every distant user |
| Static plus client fetch | cheapest and most cacheable | a visible loading state and client CPU |
When not to use it
If the page needs the database, render next to the database. A single-region product whose users are mostly in that region gains nothing: the distance it removes is already small, and it inherits a second runtime, a second observability story and a second deployment. For a page that is identical for everyone, a statically generated document on a CDN is strictly better — no compute per request, no cold start, nothing to debug. Edge rendering earns its place in the narrow band between those two: globally distributed users, a response that must vary per request, and variation that can be computed without central data.
Interview question
Q: Your team proposes moving server rendering to the edge to cut time to first byte. Which pages would you move, which would you refuse, and what measurement would you take first to make that call?
What a strong answer covers: measure the split of time to first byte between network distance and server time, per region, before moving anything; classify routes by whether their output depends on central data; move the URL-and-cookie-derivable routes; refuse the data-dependent ones or pair the move with a read replica; and name the operational cost — debugging across POPs, progressive deploys, and a restricted runtime — as part of the price rather than a detail.
Quick check
Quiz: Why can edge rendering make a personalised dashboard slower? The renderer still has to reach a single-region database, so the central round trip remains and an extra hop has been added.
Flashcard: What makes per-request rendering at the edge viable where container functions are not? — Isolate-based runtimes start in a few milliseconds instead of hundreds, so cold start does not exceed the latency being saved.