pattern

Deletion Safety Interval

also called Quarantine Window, Reclamation Hold

A mandatory quarantine period between detecting an unused resource and destroying it, which turns a cost cleanup from an irreversible change into a reversible one.

waste-eliminationsoft-deletereapersblast-radiusreversibility

A cleanup script deletes resources nobody has touched in 30 days. It runs in production for eight months and removes a genuine six figures of annual waste. In month nine it deletes the disaster-recovery environment, which is untouched by design, and the rebuild takes four days during which the company has no tested recovery path.

The script was correct about every resource it deleted. The defect is not the detection rule; it is that detection was wired directly to an irreversible action. A deletion safety interval breaks that wiring: the reaper moves a resource into a quarantine state, and only time plus silence promotes quarantine to destruction.

Why it matters

Waste elimination is the one cost activity with no reliability cost, which is why teams automate it first and least carefully. A reaper saving $8,000 a month is worth $96,000 a year and one bad deletion costing three engineer-days is about $5,000, so the arithmetic says keep going.

It changes completely when the resource is unrecoverable. Legally retained records, the only copy of a dataset, an audit log, a signing key: these have no rebuild cost because there is no rebuild. A control that is economically irrational for recoverable resources is mandatory for those, so the interval is applied by tier rather than uniformly.

Implementation patterns

  • Quarantine instead of delete. Detach the volume, deny the policy, stop the instance, rename the bucket. The resource stops doing its job and stops costing most of its money, and it still exists.
  • Hold for 14 to 30 days, long enough to cross a monthly or quarterly cycle. A quarterly DR exercise needs a longer interval than a developer sandbox.
  • Tier by recoverability, not by cost. Tier one is freely reclaimed (idle compute, orphaned addresses); tier two is quarantined; tier three is never touched by automation.
  • Rate-limit the reaper. With no cap it will eventually delete an estate at machine speed; a few resources per run makes a wrong rule a small incident.

Industry example

The pattern's value is clearest in the class of incident where a maintenance script runs against the wrong set of identifiers. A provider executes a routine deletion job, the identifier list is subtly wrong, and hundreds of customer environments disappear in minutes. Recovery then depends on backup granularity rather than on the script, because a per-customer restore from a shared backup is a far slower operation than a per-customer delete, and restoration in incidents of this shape has run into weeks against a deletion that took minutes.

A quarantine state changes that event's shape: the same wrong list leaves environments offline and recoverable. Soft delete is cheap precisely because it is not a backup: it is the original data with its access removed.

Failure scenarios

  • The reaper deletes capacity insurance. Standby replicas, spare zone capacity and DR environments look identical to waste on a utilisation chart, because being unused is their function.
  • The reaper deletes retained records, which are untouched by design and unrecoverable by definition.
  • Quarantine that keeps billing. Detaching a volume without snapshotting and releasing it saves nothing, so the reported saving never reaches the invoice.

Trade-offs

The interval delays the saving by its own length and keeps part of the cost during the hold: for a 30-day quarantine the first-year saving falls by roughly a month of the reclaimed spend, and quarantined resources still cost storage.

In exchange every deletion becomes reversible, which changes who will approve the automation. A reaper that can be undone gets permission to run against production; one that cannot stays confined to development, where the waste is not.

When not to use it

For resources that are trivially cheap to recreate and genuinely stateless, quarantine is pure overhead. An empty security group, an orphaned address, a stopped instance with no storage: delete them and move on. A quarantine pipeline is a week of engineering plus ongoing operation, so if the reclaimable estate is a few hundred dollars a month it costs more than everything it will protect.

Interview question

Q: A reaper has saved a genuine $100,000 a year for eight months and has just deleted the disaster-recovery environment. Leadership wants it switched off. Argue for keeping it, and say exactly what you would change.

What a strong answer covers: that the detection rule was right and the action was wrong; tiering by recoverability; a quarantine hold spanning the DR exercise cycle; rate limiting so the next wrong rule is small; an expiring exemption marker; and the honest cost, a delayed saving plus residual storage. A strong answer also says what it would not do: switch the reaper off, because the waste returns within two quarters and nobody restarts it.

Quick check

Quiz: Why is sorting reaper candidates by monthly cost the wrong priority order? Because cost measures the saving, not the risk; rank by recoverability instead, since the most expensive candidate is often the most recoverable and the cheapest may be the only copy of something.

Flashcard: What does a deletion safety interval change about a cleanup programme? — It converts an irreversible action into a reversible one, which is what allows the automation to run against production at all. The price is a delayed saving and residual storage during the hold.