Legacy Headroom Budget
also called Legacy Capacity Allocation, Integration Throughput Quota
The measured spare throughput of a legacy system after retry amplification, batch overlap and organic growth are deducted, allocated to new consumers as enforced quotas rather than published as available capacity.
A fourteen-year-old policy system sustains roughly 350 transactions per second at its measured peak, and the business peak already consumes about 260. Four new services want to call it. The obvious answer is that 90 transactions per second are available, and it is wrong by a factor of three.
Measured headroom is a ceiling, not a budget. It was measured while the system was healthy, on a day when no client was retrying, no batch overlapped the online window, and the business was smaller than it will be when the integration finishes. A headroom budget is the discipline of deducting those three before anyone is allowed to plan against the number.
Why it matters
Legacy systems are the one place in an architecture where capacity genuinely cannot be added. The licence model, the hardware generation or a single-writer design usually forbids it, and where it is possible the lead time exceeds the integration deadline. That makes throughput a fixed resource to be allocated rather than a target to be scaled, which is an unfamiliar mode for teams used to autoscaling.
The second reason is that overload on a legacy system rarely presents as a clean rejection. It presents as growing response times, then connection-pool exhaustion in every client at once, then a batch window that misses its deadline, which is a business incident rather than a technical one.
Implementation patterns
- Deduct retry amplification first. A default three-attempt client policy turns an offered 10 TPS into up to 30 during a partial failure, so halve anything you allocate. This is the deduction that dominates the error bar.
- Reserve for batch overlap, typically 10 to 15% of measured capacity, because the nightly and month-end jobs on a system this age run into the edges of the online window.
- Reserve for organic growth at the business's own rate, 5 to 10% a year, over the integration's full duration.
- Publish the remainder as per-consumer quotas enforced at the boundary, in a gateway or the integration layer, not as a number in a design document. An unenforced quota is a forecast.
- Measure the retry multiplier rather than assuming it: replay a controlled failure against a staging copy and watch what the client fleet actually sends.
- Give each consumer a circuit breaker with a fallback that is defined, so that hitting the quota degrades one consumer instead of queueing on the shared system.
Industry example
The pattern has been standard practice in banking and insurance integration layers since roughly 2015, where a core system's transaction rate is a contractual figure and new digital channels are onboarded against an allocated quota measured in production rather than forecast. The archetype: a mainframe-backed customer API where each new consuming service was admitted with a named ceiling and a mandatory fallback, and the onboarding review asked for the consumer's retry policy before it asked for its expected volume. Read-heavy consumers were routed to a change-data-capture copy from the start rather than being given a quota at all.
Failure scenarios
- Allocating nominal headroom. Ninety TPS of ceiling becomes 90 TPS of promises, and the first partial failure turns that into 270 TPS of offered load against a system with none to give.
- Quotas as guidance. Documented and unenforced, they hold until the first launch that exceeds its forecast.
- A per-page-render consumer admitted at all. A web tier at 200 requests per second calling once per render is six times the entire budget and cannot be throttled into fitting.
- Caching without a staleness contract, which moves the failure from overload to stale answers nobody can bound.
- The budget computed once and never revisited as the legacy system's own load grows through the programme.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Enforced quotas per consumer | The legacy system survives; failures are isolated to one consumer | Consumers must design a degraded mode; onboarding is slower |
| Replicated read copy | Removes the consumer from the legacy system entirely | A pipeline plus reconciliation, and a second thing to decommission |
| Unmanaged direct access | Fastest integration; no platform work | The first traffic spike is a core-system incident |
The decision rule: if a consumer's steady-state demand exceeds roughly half the measured headroom, it is a data-replication problem rather than an API problem.
When not to use it
If the legacy system is being switched off within about six months, do not build replication for it. The change-data-capture pipeline, its schema mapping and its reconciliation cost more than the throttling and a period of degraded response times, and the pipeline becomes a second system to retire. Buy the time with quotas and queueing instead.
The practice is also overhead where the legacy system has genuine elasticity — a system already running on scalable infrastructure with a horizontally scalable data tier does not need a rationed budget, and rationing it slows integration for nothing.
Interview question
Q: A core system peaks at 350 transactions per second and the business already uses 260. Four teams want to integrate. How much do you give them, how do you enforce it, and which of the four would you refuse outright?
What a strong answer covers: the three deductions with rough numbers and which one dominates the error; quotas enforced at the boundary rather than documented; the refusal of any consumer whose demand is a multiple of the whole budget, redirected to a replicated copy with a staleness contract; and the point at which building that copy stops being worth it because the system is close to retirement.
Quick check
Quiz: Why is nominal headroom roughly triple the allocatable figure? Because retry amplification alone can double or triple offered load during a partial failure, before batch overlap and the business's own growth are reserved.
Flashcard: A legacy system peaks at 350 TPS with 260 in use — how much can you allocate? On the order of 30 to 35 TPS in total, after halving for retry amplification, reserving 10 to 15% for batch overlap and setting aside 5 to 10% a year of growth.