concept

Compensating Manual Control

also called Load-Bearing Workaround, Downstream Correction, Shadow Correction Step

A human step outside a legacy system that corrects its output before anyone consumes it, which means the real specification is the code plus the correction - so a replacement validated against the code alone is validated against an incomplete system.

legacy-assessmentrequirementsshadow-processcutoverparallel-run

A replacement for a 16-year-old claims system passed every acceptance test and went live. Six weeks later finance rejected the reinsurance return it produced. Nothing in the test pack was wrong. The old system had always produced that return with one category misallocated, and for nine years an analyst had corrected it in a spreadsheet before forwarding it. The correction was a requirement, and it was in nobody's repository.

Sixteen years of defects nobody thought worth fixing do not disappear. They are absorbed downstream by a reconciliation spreadsheet, a monthly correction journal, a rule one person applies by hand above a threshold. Each is behaviour the business depends on, executed by a person, and invisible to every artefact an engineer would normally read. The new system is faithful to the old system's code; the business was never using that code on its own.

Why it matters

These controls are the most reliable cause of a post-cutover business failure no test could have caught, because the gap is in the specification rather than the implementation.

They also decide whether a parallel run is worth running. A comparison of the old system's raw output against the new system's raw output will agree and prove nothing, because both are wrong where the human was correcting. The comparison has to sit after the human step, which is an assessment decision rather than a testing one. They matter for funding too: 8 hours a month of senior finance time, plus one person who cannot take leave at quarter end, is often the strongest number in a case that otherwise offers only technical obsolescence.

Implementation patterns

  • Inventory outputs, not interfaces. Every recurring thing the system produces and who receives it. A core system of this age typically has 20 to 60, and a handful are annual.
  • Interview the recipient, not the owner. The owner believes the output is correct, which is why the correction lives downstream.
  • Require written confirmation per output from a named person that they change nothing. Every "actually, I do adjust the…" is a requirement with an owner, a frequency and an acceptance test.
  • Ask what the system last got wrong that nobody fixed. Operations and finance can answer; engineering cannot, because the defect was closed as working-as-designed in 2014.
  • Watch one period close. Two weeks on a system with 30 outputs covers this: a dozen conversations and one observed month-end, not a code audit.

Industry example

The condition is documented mostly through the failures it causes, since nobody publishes their own workarounds. The usable archetype is an enterprise application vendor's long-lived customer installs, of the kind an ERP vendor in SAP's position supports: decades of customer-specific corrections applied in spreadsheets outside the product, because the product's own change process was too slow or too expensive to absorb them. No upgrade path reads those spreadsheets. That is the condition, and it is not specific to any vendor.

Failure scenarios

  • Rejected statutory output weeks after cutover, with the link to the migration invisible because nothing errored.
  • A perfect parallel run. Compared at the wrong boundary it agrees throughout, is signed off, and the discrepancy appears at the first period close after the old system is off.
  • The annual control, which a 6-week parallel run cannot see.
  • The person retires. The control was never a process; when they leave mid-programme the output silently becomes wrong on both systems.
  • Over-correction. The new system fixes the defect, the human keeps correcting, and the output is now wrong in the opposite direction.

Trade-offs

Finding these controls costs assessment time the programme would rather spend on technical discovery, and the interviews create scope: once a control is documented, the business asks for it to be fixed.

The alternative is cheaper only until cutover. A control discovered after go-live costs a business incident, credibility with the function that received the wrong output, and remediation under pressure, against a pre-cutover cost of one conversation and one test. The real judgement is which controls to absorb: some are a defect to fix, some are policy that belongs with a human, and some should stop.

When not to use it

A full output-by-output walk is not warranted for a system with a handful of internal consumers and no external, financial or statutory output, and a short-lived internal tool with three users does not need it.

It is also the wrong focus where the outputs are machine-consumed end to end. There are no compensating controls to find and the risk lives in data quality instead. Spend the effort where a person sits between the output and its use; where nobody does, assess contracts and volumes.

Interview question

Q: "A replacement for a 16-year-old claims system passed every acceptance test, went live, and six weeks later produced a reinsurance return finance rejected. Nothing in the test pack was wrong. How would you have found that during the assessment?"

What a strong answer covers: that this is an assessment failure rather than a testing failure, because the missing behaviour was never in the code; the method of walking recurring outputs backwards from their consumers and requiring a named human per output; why the system's own owner is the wrong person to ask; the parallel-run change to compare after the human step; why run duration must cover the longest reporting cycle; and which controls should not be reimplemented at all.

Quick check

Quiz: A monthly report is corrected by hand before finance sees it. What does that do to your parallel run design? — The comparison must happen after the human step, on the output the business consumes; comparing raw system outputs agrees perfectly and proves nothing.

Flashcard: Why can a replacement pass every acceptance test and still produce a wrong business answer? — Because part of the legacy specification is executed by people downstream, so a replacement validated against the legacy code is validated against an incomplete system.