concept

Keep-Lit Capacity Floor

also called Keep-the-Lights-On Floor, Legacy Demand Floor, Maintenance Floor

The share of engineering capacity a legacy system consumes no matter what else is planned, which decides whether a rebuild is funded at all - and which a change freeze defers rather than removes.

rebuild-vs-rearchitectcapacity-planningchange-freezestaffingbusiness-case

Nine engineers maintain a legacy order system that absorbs about 14 change requests a month. A two-year rebuild is approved, six of the nine move onto it, and the legacy system is put under change freeze so no effort is wasted on code destined for deletion.

The freeze does not reduce demand. It queues it. 14 a month over 24 months is about 336 deferred changes, and the three engineers left are carrying a flow sized for nine. The arithmetic only works if the legacy system's demand is elective, and it is not: a measurable share is regulatory, contractual or revenue-blocking. The keep-lit capacity floor is that irreducible share. It is not an estimate of how much maintenance the team would like to do; it is the demand that will be served whatever the plan says.

Why it matters

Rebuild decisions are argued on architecture and decided on capacity. A rebuild needing six engineers, from a team of nine whose floor is five, is not under-resourced by one person; it is unfunded, and it shows up as a rebuild running at a third of its planned rate while nobody can say why.

The floor also makes the honest comparison possible. Incremental re-architecture is done by the people already serving the floor, so it competes with maintenance request by request, while a rebuild needs the floor subtracted before any capacity exists at all. That difference, rather than any argument about code quality, usually decides which approach can be delivered.

Implementation patterns

  • Measure the floor from history. Take 12 months of change requests, incidents and compliance items, classify each as deferrable or not, and express the non-deferrable share in engineer-months. On a business-critical system the answer is frequently 40 to 60% of the current team.
  • Subtract the floor before sizing the rebuild team, and show both numbers to the sponsor together.
  • Price a freeze rather than assuming one. Change rate times freeze duration gives the deferred count; then estimate leakage, the share done anyway. 20% is a reasonable start where there is regulatory exposure.
  • Plan the unfreeze release as its own event, because the first release after a long freeze bundles months of unrelated changes into one deployment and lands next to the cutover.
  • Rotate rather than split permanently, and revisit the floor quarterly, because it falls as capabilities genuinely move.

Industry example

The honest public evidence is the pattern of outcomes rather than any company's staffing table. The documented scaling migrations that worked kept the old system fully supported throughout: Figma's databases team grew a running Postgres stack roughly 100x from 2020 by vertically partitioning and then sharding rather than stopping to build a replacement, and Notion resharded a live fleet in 2023 without changing the application's routing. Neither is a rebuild, and that is the point.

The counter-archetype is the long rebuild with a frozen predecessor, which a vendor in the position of an established ERP or developer-tools company faces on every platform transition. The organisations that survive it fund both; the ones that fail fund the successor from the predecessor's maintenance slack.

Failure scenarios

  • The invisible stall. The rebuild is staffed at six on paper and runs at two, because four are continuously pulled back. Velocity is blamed on the technology.
  • Leakage without context. The engineers who knew the legacy system moved to the rebuild, so the regulatory change that breaks the freeze is implemented by whoever is left, slowly and with defects.
  • The unfreeze release. Months of unrelated changes deploy together next to cutover, and the resulting incident is attributed to the migration.
  • Attrition on the keep-lit side, because maintaining a system everyone knows is being replaced is a poor job, so the floor ends up served by nobody.
  • The floor hidden in the business case, where the same nine people appear in both the maintenance line and the delivery line.

Trade-offs

Funding the floor properly makes the rebuild look more expensive, often by 50% or more, and some rebuilds will not survive that honesty. That is the point of measuring it: a rebuild affordable only when maintenance is assumed free was never affordable.

The alternative is to shrink the floor deliberately rather than deny it. Retiring unused features, pushing a class of requests to a later release, and accepting slower response on legacy changes are all legitimate, and each is a negotiation with a named business owner. A floor reduced by agreement is real; a floor reduced by a freeze memo is deferred.

When not to use it

The floor is not a veto. Used as one it becomes an argument that nothing can ever be replaced, which is false. It is an input: when the floor plus the rebuild team fits the headcount, the architectural argument should be had on its merits. A system with a change rate near zero and a stable interface is the narrow case where a frozen predecessor plus a dedicated team actually works.

Measure it when the system is business-critical, when the programme runs longer than a year, and when the same people are named in both the maintenance line and the delivery line. That last condition is the tell.

Interview question

Q: "A nine-person team maintains a legacy order system taking about 14 change requests a month. The CTO has approved a two-year rebuild with six of the nine on it and a change freeze. What do you check before agreeing?"

What a strong answer covers: measuring the non-deferrable share of the last 12 months of change requests rather than accepting the freeze as an assumption; the deferred-change arithmetic and the leakage rate, with the point that freezes queue demand; the unfreeze release as the riskiest deployment of the programme; continuity and attrition on the keep-lit side; and the comparison that follows, since incremental work competes with maintenance request by request while a rebuild needs the floor subtracted first. A strong answer proposes a negotiated reduction of the floor with named business owners instead of a freeze.

Quick check

Quiz: A legacy system takes 14 change requests a month and is frozen for a 24-month rebuild. What has the freeze decided? — That roughly 336 changes are deferred into one eventual release and that about a fifth will leak through anyway to be done by whoever is left, so the rebuild is funded from capacity that does not exist.

Flashcard: Why does the keep-lit floor decide between rebuild and incremental re-architecture? — Because incremental work is done by the people already serving the floor and competes request by request, while a rebuild requires the floor to be subtracted from headcount before any capacity exists.