advanced 3 min answer

A SaaS vendor's documented recovery posture is point-in-time restore for any individual data store. Atlassian's April 2022 incident tested it at scale - sites for 775 customers were deleted at once, no customer lost more than five minutes of data, and the last sites came back on 18 April. Which number had the organisation actually been buying, and what would you change first?

atlassianrportodisaster-recoveryblast-radiusrestore-drills
Show the full answer Hide the answer

The situation they were in

On 5 April 2022, between 07:38 and 08:01 UTC, a maintenance script intended to remove a legacy app from customer sites was given cloud site identifiers instead of app identifiers, and the deletion API accepted both kinds without confirming which it had been handed. 883 sites belonging to 775 customers were permanently deleted in 23 minutes. Atlassian's post-incident review records that the first customers were restored on 8 April and the last on 18 April, that no customer lost more than five minutes of data, and that over 99.6% of customers were unaffected throughout.

What the two numbers say

Those two figures are the whole lesson. Five minutes of data loss is an excellent RPO. Thirteen days to restore is a catastrophic RTO. An organisation does not arrive at that pairing by neglect; it arrives there because RPO and RTO are bought with different money and only one of them is exercised by normal operation.

RPO is a property of the write path: replication, continuous backup, point-in-time snapshots. It runs on every transaction, so it is continuously tested and it works. RTO is a property of a procedure. It runs only when someone invokes it, which in practice means restoring one tenant for one unhappy customer, a handful of times a year. Atlassian's review is explicit that the restore process had been used to recover a small number of sites and that recovery at this scale was not well defined. The capability existed and had never been exercised at the cardinality of the worst plausible event.

The arithmetic shows why that gap is not recoverable under pressure. If restoring one site takes 45 minutes of largely manual work with verification, 883 sites run serially cost about 660 hours, which is 27 days of wall clock. Getting to 13 days means they were already parallelising hard. No amount of effort during the incident closes a two-orders-of-magnitude gap in a procedure's design.

What it cost them

Thirteen days of unavailability for 775 customers, a 7,500-word public review, and a programme of work to automate multi-site multi-product restoration that they committed to in the same document. The quieter cost is reputational: every enterprise buyer now asks the question, and for most vendors the honest answer is the same.

What I would change first

Not the script, and not a second reviewer. Measure the thing nobody measures: the largest number of tenants you have restored in a single run, in the last 90 days, under drill conditions. If that number is 3 and your worst plausible event is 800, your stated RTO is fiction, and you now have a defensible number to fund against.

Then make restore capacity a sized, tested property rather than a runbook: bulk restore driven by a tenant list, with per-tenant verification automated, and a drill every quarter that restores a sampled set into an isolated environment. Alongside it, the blast-radius control that costs least: a destructive bulk API that requires the caller to declare the object type and the expected count, and refuses when the count is off by more than a stated factor. A wrong-type deletion of 883 sites should have failed a precondition, not succeeded quickly.

When not to copy this

A product with 20 enterprise customers, or a single-tenant deployment, should not build multi-tenant bulk restore automation. There the right answer is a soft-delete window long enough to notice (7 to 30 days), a rehearsed manual runbook, and one drill a year. Bulk restore automation is justified by tenant count multiplied by the blast radius of your admin tooling, and if either number is small the automation costs more than the risk. The transferable part is not Atlassian's fix. It is the question: what is the largest recovery you have actually performed, and how does that compare with the largest one your tooling can cause?