A CI platform launches a fresh virtual machine for every job. 40000 jobs a day each do about 25 seconds of real work, yet the compute line is roughly four times the total job time. Nothing is throttled and the instance family is correct. Where is the money going?
Show the full answer Hide the answer
What is being tested
Whether you read a bill in the units the provider bills in, rather than the units your work happens in. Everyone sizes instances. Almost nobody checks the granularity of the meter against the duration of the task, and for short tasks the granularity is the whole story.
The mechanism
EC2 has billed Linux instances per second since October 2017, with a 60-second minimum per instance. Compute Engine charges a one-minute minimum and then per second. Both meters start when the instance starts, not when your job starts, so every second of firmware boot, image pull, agent registration and shutdown is billed at the same rate as useful work.
Put the numbers together. A 25-second job on an instance that takes about 45 seconds to boot and register bills at the 60-second minimum at least, and in practice closer to 70–105 seconds of instance time. That is three to four billed seconds for every second of work, which is exactly the ratio on the invoice. At roughly $0.10 per instance-hour for a 2-vCPU general-purpose instance at 2026 US list prices, 40000 jobs a day at 105 billed seconds each is about 1170 instance-hours a day, near $117 a day, against about $28 if only the work were billed.
Why the other options fail
- Rounding to the next hour. This was the pre-2017 model and the instinct survives it. Hourly rounding would make the bill roughly 140 times the work time, not four times, so the arithmetic rules it out.
- Burstable credits. A real failure, but it lengthens job duration, which the question says is 25 seconds of real work measured by the job itself. Credit exhaustion shows up as slow jobs, not as a gap between job time and billed time.
- Queueing. Queued work means jobs waiting for instances, which costs latency and nothing else. Idle instances waiting for work would cost money, which is the pooled model the fix recommends and the opposite of launch-per-job.
What to change
- Reuse workers with an idle timeout rather than launching per job. One pooled runner handling 60 short jobs an hour pays for one instance-hour, not 60 minimums.
- Batch related short jobs into one so the unit of work exceeds the unit of billing.
- Cut startup time, since it is billed: pre-baked images, warm caches, no package installation at boot.
- Decision rule: when the median task is shorter than roughly twice the minimum billing increment plus startup time, the scheduler is the cost lever and instance size is noise.
When this is the wrong answer
When each job must run on a clean machine because it executes untrusted code, reuse is a security decision rather than a cost decision, and the right move is to make boot fast and accept the minimum. The whole effect also disappears once tasks run for minutes: at a 10-minute job the increment is under 1% of the bill and attention belongs on instance family and purchase model instead.