Reclamation Notice Window
also called Interruption Notice, Preemption Warning
The seconds between a provider announcing it is taking an instance back and actually taking it, which sets a hard upper bound on how much work a shutdown handler can finish and therefore how large a unit of work may be.
A training job checkpoints 30 GB of optimiser state. The team adds a shutdown handler that writes a checkpoint when the reclamation notice arrives, tests it on an idle machine, and ships it. In production the notice arrives, the handler starts a 30 GB upload, and the instance disappears four minutes before the upload would have finished. Every reclamation loses the full interval since the last scheduled checkpoint, and the handler contributes nothing except a false sense of safety.
The notice window is a budget, not a warning. Whatever cannot complete inside it does not happen. EC2 spot instances give two minutes, surfaced in instance metadata and as an event, with the provider recommending a poll every 5 seconds. Compute Engine spot VMs document up to 30 seconds of best-effort shutdown time. Azure spot signals 30 seconds. A design that works on one of those windows can be worthless on another, which makes the window a portability constraint as well as a cost one.
Why it matters
Spot capacity is 60-90% cheaper, and the discount is paid for with interruption tolerance. Whether a workload has that tolerance is usually decided by a single number: how much work the application must externalise when told to stop, divided by its own throughput. A handler that needs 90 seconds fits on two-minute notice and nowhere else, and one that needs five minutes is not a handler at all.
The number also sets the size of a unit of work. Batch stages, map tasks, encode segments and training intervals should be small enough that losing one is cheap, because the notice window cannot be negotiated and the loss is bounded only by granularity.
Implementation patterns
- Size the work unit to the window: a task that completes in under the notice period needs no handler, because it can simply be allowed to finish.
- Checkpoint continuously, not on notice. Periodic checkpoints make the handler a flush of a small delta rather than a bulk upload.
- Give the handler a deadline and a fallback. Drain for at most half the window, then force the stop so shutdown hooks run at all.
- Requeue before you checkpoint. Returning in-flight messages to the queue is milliseconds and recovers the work elsewhere; a checkpoint only helps the instance that is leaving.
- Act on earlier signals where they exist. Rebalance recommendations and preemption-risk signals arrive before the formal notice and are the only way to get more than the documented window.
- Measure the handler in production, as a histogram of completion time against the window, and alert when the 95th percentile crosses half of it.
Industry example
Distributed training makes the constraint unusually sharp. In collective-communication frameworks of the kind NVIDIA's NCCL implements, every rank participates in each all-reduce, so losing one worker stalls the whole job rather than degrading it. The notice window is then not a per-instance concern but a job-level one: either the job can rebuild its communicator and continue from the last checkpoint, or 400 machines wait for one. Teams running training on interruptible capacity therefore spend their effort on checkpoint frequency and fast restart, not on what a handler can achieve in two minutes.
Failure scenarios
- A handler that cannot finish, discovered only in aggregate as unexplained lost work.
- Correlated reclamation. Many instances in one pool are taken within minutes, so the handlers all contend for the same storage endpoint and each one gets a fraction of the bandwidth it tested with.
- Disruption budgets blocking eviction. On Kubernetes the handler's eviction calls are refused while a budget is at its limit, the window expires, and the pod is killed ungracefully.
Trade-offs
Frequent checkpointing costs throughput and storage writes, and buys a bounded loss per interruption. The honest framing for a business conversation: the spot discount is paid for in engineering work whose size is set by the notice window, and for workloads whose state is large that work can exceed the saving.
When not to use it
Designing around the window is wasted effort for workloads that should not be on interruptible capacity at all: databases, stateful singletons, anything holding a user-facing request, and any job whose rework cost after an interruption exceeds the discount. Whether the workload belongs on spot in the first place is a separate question about interruption tolerance; this concept only sizes the mechanics once the answer is yes.
Interview question
Q: Your batch platform moves from one cloud's spot instances to another's and the same code starts losing work. Nothing in the application changed. What do you look at first, and what would you change structurally?
What a strong answer covers: the notice window shrinking from two minutes to about 30 seconds; measuring handler completion time against the new window; converting the handler from a bulk flush to a delta flush backed by periodic checkpoints; shrinking the work unit below the window; and requeueing in-flight work before attempting any state write.
Quick check
Quiz: How long does a spot reclamation handler actually have? Two minutes on EC2 spot and up to 30 seconds on Compute Engine and Azure spot, so a portable handler must fit the smallest window.
Flashcard: Why is a shutdown handler the wrong place to write a large checkpoint? — The notice window is 30 seconds to 2 minutes, so anything larger than a small delta does not finish and the work is lost anyway.