A B2B product has 30 enterprise customers. The API target of p99 under 300 ms has been met every month for a year and availability sits at 99.95%. In the last four renewal calls three customers raised the same complaint: a reported bug takes three weeks to reach production. You have one engineer-quarter. Where does it go?
Show the full answer Hide the answer
What is gained and what is paid
Spend the quarter on the quality attribute that is currently limiting a business outcome, and the evidence here points at exactly one: three of the last four renewal conversations named fix speed, while every technical target is green.
What it buys, in the customer's units: a reported bug reaches production in about two days instead of about 22. What it pays: a quarter with no features in it, spent on test runtime, a deploy one person can run unaccompanied, and a revert path. There is also a second-order gain worth naming in the funding conversation - a release of one day's changes has a candidate set small enough that attribution during an incident is free, so the same work shortens outages.
Why lead time is the limiting attribute here
A met target has no customers behind it. 99.95% availability and a p99 inside budget are not achievements to improve; they are constraints that are currently satisfied, and work on a satisfied constraint produces no change in any outcome anyone is paying for. The renewal calls are the measurement that identifies the binding constraint, and they are cheaper to collect than any telemetry.
One check before committing: ask whether the three weeks is engineering or contractual. If enterprise customers have a change-notification window in the contract, part of the three weeks is not yours to remove, and the work changes shape.
Why the other options fail
- Halving p99. The target is met and no customer named latency. A read cache in front of a catalogue also introduces a staleness window on data that sales teams quote from, so this option spends a quarter and adds a correctness question to get a number nobody is buying. It becomes correct the moment a customer names latency in a renewal, or a new workload makes 300 ms the binding limit.
- A second region. Regions are bought for a stated availability requirement, a residency rule or a disaster-recovery objective. With none of those written down, a second region doubles the operational surface and usually produces an untested failover, which lowers availability. Buy it when an RTO or a residency obligation exists on paper.
- Coverage from 40% to 80%. Coverage is a proxy. Driving the number up means writing tests for the code that is easiest to test, which is rarely the code that breaks. If the three weeks is actually a two-week manual regression phase, then the right framing is "replace the manual regression phase", and coverage is a consequence of that work rather than the goal of it.
When this answer is wrong
If the fix delay is caused by a customer-mandated change window, lead-time work caps out early and the quarter should go into making a fix safe to ship inside the window: flags, dark launches, cohort ramps. If a renewal has already been lost on latency rather than on fix speed, the latency option wins, because the limiting attribute has moved. Re-run the same test every quarter: which attribute is a customer currently paying attention to, and is it green?