A service adds an in-process cache in front of its shared Redis tier to remove a 1 ms round trip. What has the team given up, and at what scale does that bill arrive?
Show the full answer Hide the answer
What is gained, quantified
Two things, and the smaller one is the stated reason. The round trip to a shared cache on a local network is roughly 0.3–1 ms plus serialisation.
The larger gain is load removed from the shared tier. A single-threaded cache server handles on the order of 100,000 operations per second per instance in production. A hot key taking 200,000 reads per second cannot be served by one shard at all — but an in-process cache with a 90% hit rate on that key cuts the shared tier's traffic tenfold, which is the difference between a working system and a resharding project.
What is paid
Staleness stops being a property of the cache and becomes a property of each instance. With 200 instances each holding a 30-second local entry, a committed write is visible somewhere between 0 and 30 seconds later depending on which instance answers. Two consecutive requests from one user can land on different instances and go backwards in time, which is a bug class the single shared cache did not have: there, invalidating once invalidated for everyone.
Recovering read-your-writes now costs one of three things:
- A broadcast invalidation channel to all 200 instances, with its own delivery failures, its own fan-out at write rate, and no way to know an instance missed a message.
- Accepting the local TTL as a published staleness budget, which is honest and often correct, and which must then appear in the API contract rather than in someone's head.
- Routing a user's reads to one instance, which couples cache correctness to load-balancer behaviour and breaks the moment that instance is replaced.
The memory arithmetic is the part that is never budgeted. The local cache is duplicated once per instance. A 500 MB working set held locally on 200 instances is 100 GB of RAM across the fleet for data that cost 500 MB once in the shared tier. That ratio makes the pattern defensible only for a small hot subset, never for the whole set.
When the bill arrives
At deploy time and at scale-out. A new instance starts with an empty local cache, so every rolling deploy sends a load spike to the shared tier and the origin proportional to the fraction of the fleet replaced at once: at 200 instances and 25% surge replacement, 50 cold caches miss simultaneously. A deploy becomes a traffic event.
And during an incident. One shared cache has one hit rate you can watch. Two hundred local caches have 200, and the instance you attach a profiler to is not the one with the cold cache.
How to keep the option to reverse
Keep the local tier a strict read-through cache of the shared tier with a short TTL and a hard entry cap; cache only keys on a published allow-list; and never let correctness depend on an invalidation reaching every instance. Under those rules the local tier can be switched off with a flag and the system is merely slower.
When the simpler alternative wins
If the shared tier is not hot and 1 ms is not inside your latency budget, one shared cache is strictly better: one hit rate, one invalidation path, one memory bill, and no staleness that varies by instance. The local tier is justified by shared-tier shard load, or by a p99 budget where 1 ms of 20 ms is material — not by an average that nobody notices.
Common weak answers
- "Local caches are always faster." Faster per hit and slower per deploy, and the fleet memory cost can exceed the cost of the shard you were protecting.
- "Just use a short TTL." A short TTL shrinks the staleness window and raises the miss rate on the hot key the local cache was added to protect, which undoes the gain.