A rightsizing proposal moves 180 steady services from 4-vCPU general-purpose instances to 2-vCPU burstable instances because average CPU is 9%, projecting a 55% saving. Several of those services run a two-hour nightly batch. Review the proposal - what would you change, what would you keep, and what would you leave alone?
Show the full answer Hide the answer
What is actually required
The proposal's evidence is sound and its instrument is wrong. 9% average CPU on 4 vCPU is 0.36 vCPU of sustained demand, and paying for 4 is waste worth removing. The question is whether burstable instances are the shape that removes it.
What I would change
The ceiling, not the average, decides. A burstable instance earns CPU credits at a rate that defines its baseline - the utilisation at which earning matches spending. A t3.large sits at 30%, a t3.nano at 5%, and credits accumulate up to a 24-hour supply. The batch services need about 3.4 vCPU for two hours. A 2-vCPU instance cannot deliver 3.4 vCPU at any credit balance, because credits buy time above baseline and not vCPUs that do not exist. The nightly window stretches from two hours to three and a half, every night, and no dashboard flags it as a cost decision.
Then the billing behaviour. T3 instances launch in unlimited mode by default, so sustained use above baseline is allowed and charged at about $0.05 per vCPU-hour of surplus at 2026 US list prices. The services that fit within baseline will be fine. The ones that do not will quietly bill surplus rather than throttle, and the saving shows up on the instance line while the surplus charge appears somewhere nobody mapped to this project. In standard mode the same services throttle to baseline instead and p99 triples. Both outcomes are produced by the same proposal; only the symptom differs.
What I would keep
Sizing on measured demand instead of on the original request, and the 180-service sweep. Split the fleet by duty cycle: services whose 24-hour average sits comfortably under the baseline of the target size move to burstable; spiky services move to a smaller fixed-performance size or a newer generation, which is frequently a 15–20% price-performance gain with no credit mechanics at all.
What I would leave alone
Any service whose low utilisation is deliberate: the instance sized to absorb a failover, the one holding headroom for a known seasonal peak. That capacity looks identical to waste on a utilisation dashboard and is the cheapest reliability in the estate. It needs an owner and a comment, not a smaller instance.
The rule I would argue for and when it is wrong
Not as "burstable is bad". As a rule the team can apply without me: choose burstable when the rolling 24-hour average CPU stays under the target size's baseline and bursts are short; the choice flips the moment either the average exceeds baseline or a burst needs more vCPUs than the size has. Then ask for the surplus-credit line and the p99 to be on the same slide as the projected saving at the next review, because a rightsizing programme that reports only the instance line cannot see its own cost.