Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
5 to work through
-
intermediate
A read-heavy content platform must scale. In what order should caching, replicas, denormalisation, CDN and query optimisation be applied, and what determines the order?
2 min answer -
intermediate
A service adds an in-process cache in front of its shared Redis tier to remove a 1 ms round trip. What has the team given up, and at what scale does that bill arrive?
3 min answer -
advanced
A search platform adds a cache and sees only a small latency improvement. What are the likely reasons, and what should be measured?
2 min answer -
advanced
What happens to your system if the entire cache tier is flushed at peak traffic?
2 min answer -
advanced
Which caching failure appears only under load, and what are the fixes in order of deployability?
2 min answer
4 terms in this topic
Access Pattern Skew
The uneven distribution of requests across keys - which determines whether caching helps at all, and which synthetic load tests systematically fail t…
conceptCache Hit Ratio Economics
The non-linear relationship between cache hit rate and backend load, which determines whether a cache improvement is worth making.
conceptCache Stampede
The surge of identical expensive requests to the origin when a popular cache entry expires and many concurrent callers all miss at once.
conceptWorking Set
The subset of data actually touched in a given window, whose size against the cache's capacity determines the hit rate and therefore how much load re…
Neighbouring topics
Performance & Capacity
General material on performance and capacity engineering.
Latency
Distributions rather than averages, and the floors physics imposes.
Throughput
Work completed per unit time, and why it trades against latency.
Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
Queueing Theory
Why latency explodes as utilisation approaches capacity.
Little's Law
L = λW, and the pool sizes it computes directly.
Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.