practice

Manual Fallback Capacity

also called Degraded-Mode Throughput, Paper-Process Capacity

The transactions per hour a business can actually complete without its system - the measured number that says whether a recovery-time objective is survivable and how long the catch-up takes.

business-continuityrtodegraded-modedrillsbacklogoperations

A continuity plan states a recovery-time objective of 4 hours. Nobody asks what the business does during those 4 hours. Normal throughput is 900 orders an hour; staff with paper forms and a phone complete about 60. A 4-hour outage therefore leaves roughly 3,360 orders unserved and a backlog whose re-entry takes days, so the binding constraint is not the restore time.

Manual fallback capacity is that 60, paired with the person-hours to re-enter what it produced. Together they say whether an RTO is a plan or a hope.

Why it matters

Recovery targets are derived from the system's perspective and validated against an impact analysis that describes consequences qualitatively. The arithmetic connecting the two is the manual rate, and without it an RTO has no relationship to survival.

The asymmetry is severe. A business at 50% of normal throughput in fallback absorbs a day. One near zero cannot absorb twenty minutes, and for it continuity must be bought as redundancy rather than planned as a procedure. Both can hold identical documents with very different exposure.

Implementation patterns

  • Measure it in a drill, not a workshop. Run 60 minutes of degraded operation and count completed transactions. An unexercised fallback has a capacity of zero until proven otherwise, because forms are stale and the pricing rules live only in the application.
  • Record two numbers per process: transactions per hour in fallback, and re-entry person-hours per 100 transactions.
  • Derive the RTO from the backlog, not the restore. A plan that restores in 4 hours and needs 3 days of catch-up has missed its objective.
  • Design for re-entry: a bulk import path, an as-at timestamp so backdated records price correctly, and idempotency so a re-entered transaction cannot double-charge.
  • Keep the fallback's inputs outside the system — a printable price list, a customer extract, a stock snapshot — and hold a numeric degraded-mode trigger, since a plan with no trigger for "slow rather than down" has none at all.

Industry example

A grocery marketplace that syncs retailer inventory runs picking and substitution through its own application. During a 90-minute regional provider failure, store staff fall back to printed pick lists. Measured rate: about 35 baskets an hour per store against a normal 140, because a substitution the application resolves in a second needs a phone call.

That leaves roughly 160 unserved baskets per store, and re-entry takes about 3 person-hours per 100 baskets, since each substitution must be recorded against the right line for settlement. The outage lasted 90 minutes and the settlement team felt it for two days. The fix is not a shorter RTO: a bulk re-entry path with as-at pricing and an offline substitution sheet moves the rate to near 90 and cuts re-entry to under an hour per 100.

Failure scenarios

  • The fallback nobody has run — documented, audited, zero in practice.
  • Re-entry omitted from the plan, so the system is restored inside RTO and the business still misses its day.
  • No as-at pricing, so transactions re-entered later are priced at today's rules, which becomes a reconciliation and sometimes a regulatory problem.
  • Fallback depends on the failed system for a price lookup, so there is no degraded mode at all.
  • Capacity measured at a quiet hour and applied to a peak four times busier.
  • Duplicate transactions at cutover, because neither path was idempotent.

Trade-offs

Measuring this costs real disruption: a 60-minute drill on a live operation has a throughput cost and must be repeated at least annually, because the number decays as the product gains features paper cannot express.

The alternative is to buy capacity instead of measuring it — redundancy, a second provider, an offline client. That is right where the manual rate is structurally near zero and wrong where a human can do the work slowly, since redundancy costs far more than a rehearsed paper path.

When not to use it

When no manual equivalent can exist. For algorithmic decisions with no human fallback, invest in redundancy and graceful degradation.

When failure is tolerable. A batch reconciliation that can run tomorrow needs an RPO discussion, not a fallback rate.

When the drill would create the risk it studies. Rehearsing a manual path for a safety-critical process may be worse than the outage; walk it through on paper instead.

Interview question

Q: Your continuity plan has a 4-hour RTO signed off by the business. What would you measure to find out whether it is survivable, and what would you change in the architecture once you had the number?

What a strong answer covers: the manual rate and re-entry person-hours, obtainable only from a drill, with an unexercised fallback counted as zero; backlog arithmetic deriving the RTO from clearability by the next business day; the architectural consequences (bulk re-entry, as-at pricing, idempotency, offline reference data); and when to buy redundancy instead.

Quick check

Quiz: Normal throughput is 900 orders an hour and the measured manual rate is 60. What does a 4-hour outage really cost? About 3,360 unserved orders plus a backlog that may take days, so the RTO should come from backlog clearability rather than restore time.

Flashcard: Which two numbers make an RTO testable? — Transactions per hour in fallback and re-entry person-hours per 100 transactions, both from a timed drill; an unexercised fallback counts as zero.