advanced 3 min answer

A finance-reporting platform runs a nightly batch job and a streaming job that compute the same daily revenue figure, and the batch number is the one the controller signs. You are asked to retire the batch path. Sequence it so every step is reversible and say where the point of no return is.

lambdareconciliationauditparallel runreprocessing
Show the full answer Hide the answer

The sequence

  1. Write down what the batch job actually computes. Timezone boundary for "day", refund and chargeback attribution, currency rounding, what it does with events that arrive after the boundary. This is the deliverable the streaming path has to match, and it is usually undocumented and partly accidental. Reversible: nothing has changed yet.
  2. Make the streaming path emit a signed daily close, not a running total: a final, versioned result published after the lateness allowance expires, with a completeness statement attached. A controller cannot sign a number that is still moving. Reversible: it is a new output nobody depends on.
  3. Parallel run and compare per account, not per total. A 0.02% difference on the total can hide offsetting per-account errors of 5% in both directions. Compare at the grain the business would dispute. Reversible: reporting still reads the batch number.
  4. Flip the report's source to the streaming close, with the batch job still running and reconciling silently for at least two close cycles of the kind that gets signed — if the controller signs quarterly, that means two quarter-ends. Reversible: flip the source back in one configuration change.
  5. Stop scheduling the batch job but keep it runnable, with its inputs retained. Reversible: run it on demand for any period still in retention.
  6. Delete the batch code only after the audit retention period, which is a compliance question and not an engineering one.

Where data can diverge, and how you would know

The boundary minute (an event at 23:59:59.7 in one timezone), late events beyond the watermark, restatements after a correction, and rounding applied at a different point in the calculation. Each of these diverges because the two paths make a different implicit choice about a case neither specification mentions, which is also why the batch job that has run in production for years is the only complete statement of the rule. Detection is a per-key diff with a tolerance and an owner, run every close, with a non-zero diff triggering a defined action. A reconciliation nobody acts on is a cost, not a control.

The point of no return

Not the source flip. It is the moment your input retention falls below the period you might be asked to re-derive. Until then any disagreement can be settled by reprocessing; after it, the streaming close is the only record and its correctness is an assertion. Tiered storage moves that line out cheaply and is the single highest-value purchase in this migration.

Rollback at each stage

Steps 1 to 3 need no rollback. Step 4 rolls back by configuration. Step 5 rolls back by running the job. Step 6 does not roll back, which is why it comes last and is gated on a retention policy rather than on confidence.

How long it really takes

One iteration per close cycle. With monthly closes and two signed cycles of parallel running, plus the definition work, this is two to three quarters, and most of it is waiting rather than building. Teams that promise six weeks are counting the code.

When keeping both paths is right

If the batch job is the audit evidence and costs one job-hour a night, retire the duplicated logic rather than the second run: generate both outputs from one definition, or keep the batch job purely as an independent check. The real cost of Lambda is two implementations of one business rule drifting apart, not two executions.

Common weak answers

  • "Run both for a month and compare totals." Totals agree while accounts do not.
  • "Cut over and keep the batch code in git." Code in git is not a runnable path; its inputs and its environment are the part that decays, and you discover that during the first dispute, when the job fails on a dependency that moved two releases ago.