Southwest cancelled about 16,700 flights in December 2022 when crew scheduling could not recover from a storm, after years of deferred modernisation. How do you make that argument before the failure rather than after?
Show the full answer Hide the answer
The case, as publicly reported
A severe winter storm caused widespread cancellations across US carriers. Most recovered within days; Southwest did not. Its crew scheduling system could not reconcile the volume of reassignments, and staff reportedly resorted to phoning crews individually. Around 16,700 flights were cancelled over roughly ten days, the company reported a pre-tax hit of roughly \(1.1 billion, and the DOT imposed a \)140 million penalty in December 2023.
Union representatives had publicly warned about the scheduling technology for years. The modernisation had been repeatedly deferred in favour of higher-return investments.
Why the argument usually fails beforehand
Because it is made in the wrong currency. "The crew scheduling system is old and needs replacing" competes against proposals with revenue attached, and loses every planning cycle — correctly, from the perspective of whoever is allocating.
The system had also been working. Deferred modernisation of something that functions daily looks like prudence right up until the tail event, and no individual deferral decision was obviously wrong.
How to make it land
1. Convert it into a quantified tail risk. Not "this is old" but: under a disruption of this magnitude, this system cannot re-solve the schedule; the consequence is N days of cancellations at £X per day, with a probability of roughly P per year. Expected annual loss is a number that sits in the same table as revenue projections. This is the single most important move.
2. Find the leading indicators and report them. Time to re-solve the schedule after a disruption, the size of disruption the system has actually handled, manual interventions per week, hours of overtime spent working around it. A metric that has been trending in the wrong direction for three years is far more persuasive than an opinion, and it converts a subjective concern into an observable one.
3. Test the tail, do not argue about it. A game day that replays a historical severe-weather day against the current system produces evidence. "It could not complete" is unanswerable in a way that a slide is not.
4. Name the concentration explicitly. This was not a general legacy problem; it was one system whose failure could stop the airline. A capability map with criticality heat-mapping against technical health puts that single square in front of the board.
5. Offer a proportionate option. "Replace the crew scheduling platform" is a multi-year programme that will not be funded. "Add capacity and a bulk re-solve path for the disruption scenario, in two quarters" might be, and it addresses the actual tail risk. Modernisation proposals fail partly because they are all-or-nothing.
6. Record the decision. If it is deferred again, an ADR or risk register entry naming the accepted risk and the trigger conditions puts the decision on record. That is not blame-shifting; it is what makes the risk visible at the next review instead of being re-argued from zero.
What a strong answer adds
Naming the general pattern — deferred maintenance concentrates risk into rare, correlated events. The system works every ordinary day, which is exactly why the investment case never clears the bar, and the loss when it arrives is not proportional to the savings that accumulated. Architects who can price tail risk get modernisation funded; those who appeal to elegance do not.