concept

Working Set

also called Hot Data Set

The subset of data actually touched in a given window, whose size against the cache's capacity determines the hit rate and therefore how much load reaches the origin.

caching-performancecache-sizinghit-rateevictionorigin-load

A cache's hit rate is not a property of the cache. It is the answer to one question: does the data being touched right now fit? Everything else — eviction policy, TTL choice, serialisation format — moves the hit rate by a few points. Fitting or not fitting moves it by tens.

The working set is the set of distinct items accessed in a window relevant to your traffic: the last five minutes for a session store, the last day for a product catalogue, the last hour for a feed. Its size is measurable, it grows, and it is almost never on a dashboard.

Why it matters

The relationship between cache size and hit rate has a knee rather than a slope. While the cache holds the hot region, nearly every request hits. Once the hot region no longer fits, eviction begins removing entries that are about to be requested again, each miss inserts and evicts something else, and the cache churns — a regime where adding 10% more memory can recover far more than 10% of the hit rate, and losing 10% can cost far more.

What makes this urgent is that origin load is the miss rate, so a hit-rate movement that reads as a rounding error on the cache dashboard arrives behind the cache as a multiple. The working set is therefore the leading indicator and the hit rate is the lagging one — by the time the hit rate has moved, the origin is already carrying the extra load and there is no time to provision for it.

Implementation patterns

  • Measure it directly. Count distinct keys per window with a sketch such as HyperLogLog, and multiply by the mean serialised entry size. Two counters and a daily report are enough.
  • Track its growth rate, and size the cache for the forecast working set plus growth over your procurement interval, not for today's.
  • Cache the hot subset explicitly with an allow-list or an admission filter, so one large scan cannot evict the hot region. Admission policies that require a second access before caching (TinyLFU and similar) exist for this.
  • Separate workloads with different working sets into different caches. A nightly analytics scan and a user request path sharing one cache means the scan evicts the users' data every night.
  • Alert on working-set size against capacity, because hit rate is the lagging indicator: it tells you after the origin is already hot.

Industry example

Tiered CDN caching is this relationship made into an architecture, which is why a network in the mould of Akamai is layered rather than flat. An edge location can only hold the working set of the population it serves, so a long tail of unpopular objects would evict the popular ones; a mid-tier or origin shield exists to hold the larger working set that no single edge can, and to collapse the misses from many edges into one origin fetch.

Failure scenarios

  • A scan or a backfill that walks the whole table, inflating the working set for an hour and evicting everything hot.
  • A new feature adding a dimension to the cache key — locale, device, experiment arm — which multiplies the working set without changing the data.
  • Organic growth crossing the knee, so the origin load rises by a factor nobody can attribute to any change.
  • A cache flush at peak, where the recovery time is however long it takes to re-fetch the entire working set from an origin now receiving all of the traffic.

Trade-offs

Holding the forecast working set costs memory you do not need today, in a tier whose cost is linear in capacity. Holding less costs origin load that cannot be provisioned in minutes. The honest framing is insurance: cache memory is cheap and continuous, origin overload is expensive and sudden.

When not to use it

A near-uniform access distribution has no hot region, so there is nothing a cache can hold that will be asked for again, and no affordable size produces a useful hit rate — the fix belongs in the data model or the query.

Interview question

Q: "Origin load has doubled over a quarter with no release and no traffic growth. The cache reports a 96% hit rate, down from 98%. What happened, and what would you put on the dashboard so this is predicted rather than discovered?"

What a strong answer covers: origin load as one minus hit rate, so two points is a doubling; the working set crossing cache capacity as the mechanism; candidate causes such as an added cache-key dimension or organic growth; measuring distinct keys per window and entry size; and alerting on working set against capacity because hit rate lags.

Quick check

Quiz: A cache holds 80% of the working set and the working set doubles. Why is the hit rate loss worse than proportional? Eviction starts removing entries that are about to be re-requested, so the cache churns and origin load rises non-linearly.

Flashcard: What should you monitor instead of hit rate to predict origin load? — Working-set size against cache capacity, plus the working set's growth rate; hit rate only reports the problem after the origin is hot.