advanced 2 min answer Multiple choice

On 5 April 2022 Atlassian ran a maintenance script to deactivate a legacy app. The script carried both a mark-for-deletion mode used in daily operations and a permanent-delete mode kept for compliance erasure. It ran in the wrong mode against site identifiers rather than app identifiers and destroyed 883 sites belonging to about 775 customers. Restoring them took up to two weeks. Which design change most directly prevents a repeat?

atlassianerasureblast radiussoft deleterestore granularity
Pick one
Show the full answer Hide the answer

The situation

Atlassian's published post-incident review describes a single script carrying two capabilities: the routine mark-for-deletion used day to day, and the permanent delete retained because compliance erasure requires data to actually go. The run used the wrong mode and the wrong class of identifier. Internal monitoring did not fire, because the deletions travelled the standard deletion workflow and looked exactly like intended work.

Why the blast radius was that large

Two properties multiplied. The destructive capability sat one flag away from the routine one, so the difference between a reversible daily operation and an irreversible one was a parameter. And the identifiers accepted were sites rather than apps, so a single argument list addressed whole customers instead of a component inside them. Recovery was slow not because backups were missing but because restoring individual customer sites into live multi-tenant environments is a granular operation the platform was not built to perform at scale, and the deleted records included the contact details needed to reach the affected customers.

Why separating the capability is the structural fix

It removes the class of mistake rather than one instance of it. A permanent-delete tool with its own binary, its own credential and its own approval cannot be reached by a routine maintenance run at all, no matter which flags an engineer types. It also concentrates the expensive controls, dry runs, scoped identifiers, rate limits, staged execution, in the one place where they are worth paying for, instead of taxing every ordinary script.

Why the other options fail

  • Two-person approval on every run. The script ran daily for routine work. A control applied to a routine operation is approved by reflex within a fortnight, and it does not distinguish the run that mattered from the hundred that did not.
  • A dry-run mode. Useful, and it fails here. The output would have been 883 identifiers, which is exactly the volume a human confirms without reading. Dry runs catch mistakes of intent, not mistakes of scope.
  • More frequent backups. This addresses recovery time, not the deletion, and the constraint was the granularity of restore into shared environments rather than the age of the backup. It also leaves the erasure obligation unmet, because a backup that can resurrect deleted data is precisely what a compliance delete must not leave behind.
  • Type-validate the identifiers. A real improvement and specific to this script. Write it, then notice that it protects one tool while every other script with a delete path remains untouched.

When this is the wrong lesson to draw

For a small estate with one operator and no multi-tenant restore problem, a second tool is ceremony and a soft delete with a 30-day recovery window does most of the work. The separation earns its cost when two things are both true: deletion has a genuinely irreversible mode because regulation requires one, and the identifier namespace lets one argument address a whole customer. Systems with only reversible deletion have a different problem, which is proving that erasure ever completed.