practice

Restore Drill

A scheduled, timed exercise of restoring from backup into a clean environment — the only thing that converts a backup from a hope into a control.

drtestingverification

Backup job success is not evidence of recoverability. The recurring findings when an organisation first restores properly: the backup excluded a schema or a volume; the encryption key is stored in the system being restored; restore takes far longer than the stated RTO; the runbook is years stale; the person who knew the procedure has left; or the backup has been silently failing because the job reports success while writing nothing.

A drill fixes what a dashboard cannot. The properties that make it worth doing: into a clean environment, not over the original; timed, with the result compared to the stated RTO; verified, by checking data integrity rather than that the process completed; and performed by someone other than the author of the runbook, which is how documentation gaps surface.

Two operational habits worth adding. Alert on backup success rather than failure — a job that stopped running produces no failures at all, which is the silent case. And treat the drill as a compliance obligation with a date, not as something to do when there is spare capacity, because there is never spare capacity.

Quarterly is a reasonable cadence for critical systems; the important property is that it is scheduled and that its output is a measured number.