Latency Revenue Elasticity
The measured change in a business outcome per unit of added or removed latency, which turns "faster is better" into an amount of money and tells you when to stop buying milliseconds.
An architect proposes an in-memory cache tier that takes 40 ms off checkout p99 for about 9000 dollars a month. Finance asks what the 40 ms is worth. Nobody knows, so the argument is settled by seniority, and it recurs every quarter with a different component.
Latency elasticity is the number that ends that pattern: the measured change in conversion, order value or task completion per 100 ms. It is not a constant and not transferable between products, so the widely quoted figures from retailer conference talks around 2006 to 2010 should not be borrowed — they were measured on other people's customers at other people's latencies.
Why it matters
Without the number, performance work is funded by narrative and defunded by the next cost review. With it, both directions become arithmetic. A measured 0.4% conversion loss per 100 ms on a checkout doing 60 million dollars a year makes 100 ms worth about 240000 dollars a year, so a 9000-dollar monthly cache removing 40 ms is an easy purchase. On an internal admin tool the same 40 ms is worth nothing and the cache is waste.
The number also shows where the curve flattens. Elasticity is non-linear: 3 s to 1 s moves behaviour, 300 ms to 260 ms usually does not, because attention has already been retained.
Implementation patterns
- Measure by injection, not by correlation. Add 100 to 300 ms of deliberate delay to a small randomised share of sessions on one surface and compare outcomes. Correlational analysis of naturally slow sessions confounds device quality and network with speed.
- Measure per surface and segment. Search, product page and checkout differ, and mobile users on weak networks are usually several times more sensitive than desktop.
- Use the percentile that matches the experience. A session touching 30 requests meets your p99 about a quarter of the time, so p99 describes the user's experience even when p50 looks excellent.
- Express the result as money per 100 ms per year on that surface, and put it in the decision record beside the infrastructure quote.
Industry example
Public figures belong in the anecdote column rather than the planning column: the most repeated claims come from talks by large retail and search platforms between 2006 and 2010, describing losses of a fraction of a percent of revenue per 100 ms on their own properties. The transferable part is the method, not the coefficient. Teams that run injection experiments report elasticities an order of magnitude away from the folklore in both directions, and often find a flat region where speed buys nothing.
Failure scenarios
- Borrowed coefficients. A B2B product with a weekly login cadence funds a caching programme on a consumer retail figure, and the effect after launch is indistinguishable from zero.
- Correlational measurement. Slow sessions convert worse because they come from older devices in weaker markets; the analysis blames latency and the fix targets the wrong thing.
- Ignoring the reverse. The same curve says where to remove capacity, and teams that measure only to justify spending never bank that saving.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Measure elasticity by injection | Performance spend becomes defensible in both directions | A deliberately degraded experience for a small share of users during the experiment |
| Assume faster is better | No measurement work | Endless unwinnable arguments with finance and money spent on flat parts of the curve |
| Borrow a published figure | Free | A number that is probably wrong by an order of magnitude for your product |
When not to use it
Do not measure elasticity where latency is a contractual or safety property rather than a behavioural one. If an integration partner's timeout is 500 ms, the value of staying under it is not elastic: it is the contract. The same holds for a trading path or a control system, where the requirement comes from the domain. And do not run the injection experiment at all on a surface where deliberate slowdown is unethical or materially harmful — an emergency contact flow, a medical record lookup. Use the safe floor and spend without the arithmetic.
Interview question
Q: You can halve p99 on your search path by doubling its infrastructure cost. How do you decide, and what do you tell the CFO?
What a strong answer covers: the elasticity for that surface, measured by injecting delay on a small share of traffic rather than borrowed from a talk; money per 100 ms per year against the incremental spend; and the position on the curve, since 2 s to 1 s and 200 ms to 100 ms are not the same purchase. A strong answer names the cheaper levers that are not trade-offs at all — a missing index, a chatty call pattern, an oversized payload — and notes that a flat measurement is an argument for cutting capacity.
Quick check
Quiz: Why is correlational analysis of slow sessions a poor estimator of latency elasticity? Device and network quality are confounded with latency; only randomised injection isolates the effect.
Flashcard: What single number turns "faster is better" into a purchasing decision? — Money per 100 ms per year on the specific surface, measured by randomised latency injection and compared against the incremental infrastructure cost.