A fourteen-year-old policy system sustains roughly 350 transactions per second at its measured peak and the business peak already consumes about 260 of them. Four new services want to call it directly. Roughly how much of that system can you sell them, and what does the number decide?
Show the full answer Hide the answer
The assumptions, stated
Three assumptions carry the estimate. The measured peak is a ceiling, not a budget, because it was measured while the system was healthy. Client retries multiply offered load during exactly the conditions where headroom disappears. And the legacy system's own load keeps growing, at whatever rate the business grows, while the migration runs.
The arithmetic
Nominal headroom is 350 minus 260, so 90 transactions per second. Now spend it honestly.
- Retry amplification. A default client policy of three attempts with backoff turns an offered 10 TPS into up to 30 TPS during a partial failure. Budget a factor of two on anything you allocate, which halves the usable figure to roughly 45 TPS.
- Batch overlap. The nightly and month-end jobs on a system this age overlap the online window at the edges. Reserve 10 to 15% of measured capacity, taking it to about 40 TPS.
- Organic growth. An eighteen-month integration at 5 to 10% annual growth in the legacy system's own load consumes another 8 to 12 TPS of the ceiling before the migration finishes.
That leaves on the order of 30 to 35 TPS to allocate, with a plausible range of 20 to 45. Four services get about 8 TPS each as a hard quota enforced at the integration boundary, not as a guideline in a design document.
Which assumption dominates the error
The retry multiplier. Growth and batch overlap move the answer by tens of percent; a client that retries aggressively on timeout, in a fleet that scales out under load, can move offered load by a factor of five and does so precisely when the system is already struggling. The cheapest way to shrink that error bar is to measure it: replay a controlled failure against a staging copy and watch what the clients actually send.
What the number rules in and out
The 8 TPS quota is a design constraint, and it eliminates whole integration shapes on sight.
- A web tier that calls the policy system once per page render is ruled out. A modest 200 requests per second of page traffic is 200 TPS against a system with 35 to give — roughly six times the entire budget. That consumer must read a replicated copy fed by change data capture, or a cache with a published staleness budget, and must not hold a synchronous dependency on the legacy system at all.
- An event-driven consumer reacting to policy changes is ruled in, because its load is proportional to change volume rather than to read volume, and change volume on a policy book is orders of magnitude lower than read volume.
The decision rule: if a consumer's steady-state demand exceeds roughly half the measured headroom, it is a data-replication problem rather than an API problem. Below that, a quota and a circuit breaker are enough.
When this is the wrong answer
If the legacy system is being decommissioned inside six months, do not build replication for it. Standing up a change-data-capture pipeline, its schema mapping and its reconciliation costs more than the throttling and the short period of degraded response times, and the pipeline becomes a second thing to decommission. Buy the time with quotas and queueing instead.
Common weak answers
- "Add capacity to the legacy system." On a fourteen-year-old platform the licence, the hardware model or the single-writer design usually forbids it, and where it is possible the lead time exceeds the integration deadline.
- "Cache everything." A cache in front of an unbounded key space with no staleness contract moves the failure from overload to stale answers that nobody can bound.
- "Use the headroom and monitor it." Ninety TPS of measured headroom is not ninety TPS of allocatable capacity, and an alert fires after the legacy system has already shed load.