TSB's 2018 core banking migration moved 1.9 million customers in a single weekend and failed publicly. What would you have required before approving that cutover?
Show the full answer Hide the answer
The case, as publicly reported
In April 2018 TSB migrated from a platform rented from Lloyds Banking Group to Proteo4UK, built by its parent Banco Sabadell. The migration moved around 1.9 million customers and roughly 5 billion records over one weekend.
On reopening, customers could not access accounts, some saw other customers' data, and disruption continued for months. The independent review commissioned by TSB's board (Slaughter and May) concluded the platform was not ready for the migration and that testing had not adequately covered the migration itself. Total cost was reported at around £330 million, the CEO resigned, and in 2022 the FCA and PRA jointly fined TSB £48.65 million.
What I would have required
1. Evidence from a full-scale rehearsal, not a scaled-down one. The failure mode was capacity and behaviour under real production volume. A rehearsal at 10% of data with 5% of the load tests the procedure, not the system. Require at least one dress rehearsal at full data volume with production-like concurrency, and treat its results as the go/no-go evidence.
2. A staged migration, or a defensible reason there cannot be one. Big-bang is occasionally unavoidable in banking — accounts are interlinked and dual-running two ledgers has its own correctness problems — but "unavoidable" must be argued rather than assumed. Migrating by customer cohort, starting with a small internal population, bounds the blast radius to something recoverable. Where the coupling genuinely forbids that, the compensating controls have to be much stronger.
3. A rollback plan that has been executed, not written. The question to ask is not "can we roll back" but "when did we last roll back, how long did it take, and what was the data position afterwards?" Rollback after customers have transacted on the new platform is a data-reconciliation problem, and if the answer is "we cannot roll back once we open", then the go/no-go decision is the only control and it needs to be correspondingly rigorous.
4. Named, quantitative go/no-go criteria agreed in advance. Written down before the weekend, with thresholds, and with a named person empowered to say no. Under time and cost pressure, criteria decided in the room are criteria that get met.
5. An honest readiness view, structurally protected. The board review pointed at optimism in reporting. Independent assurance reporting to the board rather than to the programme is the structural answer, because the programme cannot be the sole judge of its own readiness.
6. A capacity model validated against measurement. Not a projection — a measurement from the rehearsal.
The general principle
The cutover is not the risky part; the readiness assessment is. Every large migration failure looks like a technical failure on the night and turns out to be a decision failure weeks earlier — proceeding on a date rather than on evidence.
What a strong answer adds
Acknowledging the commercial pressure honestly: TSB was paying substantial fees to run on its former owner's platform, which created real urgency. The architect's job is not to pretend that pressure does not exist but to make the risk it is buying explicit and costed — "moving on this date rather than when criteria are met carries these specific risks, at this estimated cost" — so the decision is made with open eyes rather than by default.