practice

Chunk Retention Window

also called Asset Retention Window, Stale Chunk Horizon

The period after a deployment during which the previous build's code-split files must stay fetchable, because tabs opened before the deploy still resolve lazy imports against filenames the current build no longer contains.

code-splittingdeploymentversion-skewcdnsingle-page-application

A user opens the application at 09:10. The team deploys at 11:40, replacing the asset directory with the new build's content-hashed files. At 14:00 the user clicks through to a settings route whose code was never loaded in this session. The browser requests /assets/settings-a91f3c.js, the object no longer exists, and the dynamic import rejects.

The symptom is not an error page. It is a blank region, or an unhandled promise rejection nobody reads, because a rejected import() happens outside the render path that error boundaries cover. The user sees a button that does nothing.

The retention window is the explicit answer to "how long must the old build stay downloadable?" It is a number, it belongs in the deployment pipeline, and in most organisations it is zero by accident.

Why it matters

Exposure is session length against deploy frequency. If the median session is 20 minutes and the 99th percentile is 8 hours, and the team deploys six times a day, then every session in that tail crosses at least one deploy boundary. At 2 million sessions a day, 1% is 20,000 sessions that can hit a missing file.

The failure is indistinguishable from a product bug in every signal a team normally watches, and it correlates with deploys, so it gets dismissed as deploy noise. For products whose unit of use is a session rather than a page view — a meeting client, a design canvas, a trading screen — the share is far higher.

Implementation patterns

  • Content-hashed filenames and additive deploys. Write the new build alongside the old; never delete as part of a deploy. A 2 MB build, six deploys a day, 30 days of history is roughly 360 MB of object storage.
  • State the retention number. At or above the 99.9th percentile session length plus the longest plausible tab sleep. Seven to 30 days is the normal range. Write it as a lifecycle rule so it cannot drift.
  • Put the build id in the document and check it, so a tab can offer a quiet "new version available" affordance before the user reaches a missing file.
  • Catch the rejection and recover. Retry the import once against a freshly fetched manifest, then reload while preserving route and form state. A bare reload that loses a half-written message is its own incident.

Industry example

This is the mechanism behind the ChunkLoadError class of reports that accumulates in production wherever route-level code splitting meets continuous deployment, standard practice from roughly 2017. It is rarely written up as an outage because it never takes a whole site down; it degrades a slice of long sessions continuously.

Zoom's reported growth from about 10 million daily meeting participants in December 2019 to a peak above 300 million in April 2020 is the shape of the problem rather than a report of it: sessions measured in hours, a client nobody can redeploy mid-session, and a deploy cadence nobody wants to slow down.

Failure scenarios

  • The blank panel. The import rejects, nothing renders, and no error reaches the user or monitoring.
  • The silent purge. A lifecycle rule deletes objects older than seven days while a pinned tab has been open for two weeks.
  • Reverse skew after a rollback. The manifest reverts while some clients booted on the newer build, so the missing files are now the new ones.

Trade-offs

Choose Gains Pays
Long retention (30 days) almost no broken sessions, no forced reloads storage, and a wide surface of live code versions
Short retention plus a reload prompt one version effectively live interrupted users, lost in-progress state, a prompt some will dismiss

The security point is the one teams miss: retention keeps old JavaScript reachable, so a client-side fix is complete only once the window has passed or the old assets are purged.

When not to use it

If the application is server-rendered per navigation with no dynamic imports, there is no window to manage. If deploys are weekly and sessions last minutes, a version prompt alone is proportionate.

Retention also does not fix API skew: the old bundle calls the endpoints it was built against, and that is a separate contract with its own compatibility window.

Interview question

Q: "Your single-page application deploys six times a day and your p99 session is eight hours. A product manager reports that the settings page 'sometimes does nothing'. Walk me from that report to a fix, and tell me what number you would write into the deployment pipeline and how you would choose it."

What a strong answer covers: the dynamic-import rejection as mechanism; import failures instrumented apart from deploy noise; the retention number derived from the session-length distribution; recovery that preserves user state; and an owner for purging old bundles after a security fix.

Quick check

Quiz: A tab has been open nine days across fourteen deploys. Which failure arrives first, and why does monitoring miss it? — A lazy route's chunk 404s and the import rejects outside the render path, so no error boundary catches it.

Flashcard: How long must a build's code-split files stay fetchable? — At least the 99.9th percentile session length plus tab sleep, commonly 7 to 30 days; it costs storage, and old JavaScript stays reachable throughout.