A 300-node interruptible fleet runs stateless batch workers from a 4 GB container image. Reclamation replaces roughly 400 nodes a day and the discount is quoted at 70%. Estimate what the churn costs and at what interruption rate the discount disappears.
Show the full answer Hide the answer
The assumptions, stated
Each replacement pays four fixed costs: an image pull, boot and warm-up time billed but unproductive, re-executed work from the task that was killed, and scheduler churn. Assume a 90-second boot, a 60-second image pull and a 30-second warm-up, a mean task of 20 minutes, and a reclamation notice measured in tens of seconds.
The arithmetic, step by step
Image traffic. 400 pulls a day at 4 GB is 1.6 TB a day. If pulls cross a managed address-translation gateway or come from a registry outside the network, that traffic is metered per gigabyte processed on top of any transfer charge, so at a per-gigabyte processing rate in the low cents this is on the order of $50 to $90 a day, or $1500 to $2700 a month. Check your own region's rate rather than trusting that range.
Unproductive instance time. 400 × 3 minutes = 1200 node-minutes = 20 node-hours a day. The fleet supplies 300 × 24 = 7200 node-hours a day, so this is 0.3%.
Re-executed work. A task killed at a uniformly random point loses half a task on average: 10 minutes × 400 = 67 node-hours a day, or 0.9% of the fleet.
Total: about 1.2% of fleet-hours plus roughly $2000 a month of image traffic. Against a 70% discount that is noise, and spot is the correct call.
Where it stops being noise
Hold the interruption count and change one assumption: the task runs four hours instead of twenty minutes, with no checkpointing.
- Lost work per interruption: 2 hours. 400 × 2 = 800 node-hours a day = 11% of the fleet.
- Triple the reclamation rate during a regional capacity squeeze: 33%. You now pay 30% of the on-demand price for 67% useful work, so the effective cost per useful hour is about 45% of on-demand and the 70% discount has become roughly 55%, on work that also takes half again as long in wall-clock time.
- Push further and the fleet fails to finish at all: when expected lost work per unit exceeds the work completed between interruptions, progress goes to zero regardless of price, because each restart loses more than the previous attempt gained.
Which assumption dominates the error
The ratio of task duration to mean time between interruptions. Image size and boot time move the answer by fractions of a percent; the duration ratio moves it from 1% to 30%. Everything else is detail.
The decision rule, and the cheap levers
Keep the unit of work at or below roughly a tenth of the mean time to reclamation, or checkpoint at that interval. With 400 reclamations a day across 300 nodes, mean time to reclamation per node is about 18 hours, so units of under two hours are safe and a four-hour uncheckpointed task is not.
Two levers cost almost nothing. Shrink the image: 4 GB to 400 MB cuts the traffic tenfold. Put a pull-through cache inside the network, which removes the metered gateway path entirely.
When this is the wrong answer
Do the engineering arithmetic before the infrastructure arithmetic. Interruption handling, checkpointing and a diversified fleet is realistically two engineer-weeks, call it $8000. If the fleet costs $400 a month, the 70% discount saves $280 a month and pays that back in 29 months, so choose on-demand and leave it alone. At 300 nodes the saving is five figures a month and the work pays back in days.