concept

Billing Meter Granularity

also called Billing Increment, Meter Minimum

The unit and the minimum each provider meter charges in, which makes a cost model built in units of work wrong by the ratio between the two whenever the work is smaller than the meter.

cost-modellingbilling-incrementsminimum-durationrequest-chargesestimation

A continuous-integration platform launches a fresh virtual machine for every job. The jobs do 25 seconds of real work each, the compute line is four times the total job time, and the per-second rate on the price page is exactly what was modelled.

The model was built in seconds of work. The provider bills in instance-seconds with a minimum per instance, and the meter starts when the machine starts rather than when the job does. Every meter has two properties and the price page advertises only one: the rate, and the increment plus minimum that the rate is applied to. The same gap appears in storage (a minimum storage duration in a colder class, a per-object metadata floor), in object access (a charge per request regardless of size), in addresses (per address-hour whether traffic flows or not) and in managed services (a minimum capacity unit).

Why it matters

A rate error of 10% is a rounding difference; a granularity error is a multiplier. A 25-second task on a meter with a 60-second minimum and a 45-second boot bills at three to four times the work performed. A 4 KB object in a class with a 128 KB metadata floor bills at over thirty times its size. A lifecycle rule that moves an object into a class with a 90-day minimum and deletes it on day 40 pays for 90 days plus the transition request.

Implementation patterns

  • Write each meter as a triple: rate, increment, minimum. A model line with only a rate is unfinished.
  • Compare the meter's unit against the work's unit first. If the work is smaller than the increment, the ratio between them is your error.
  • Batch work up to the increment. Pack short jobs onto a warm pool rather than a machine each; combine small objects; hold a connection rather than reconnecting.
  • Check the floor before assuming a discount applies. Features with a minimum size silently do nothing below it: Anthropic's documentation in 2026 gives a minimum cacheable prompt prefix of 512 to 4096 tokens depending on model, and shorter prefixes are processed without caching and without an error.

Industry example

Content-addressed storage shows the problem in one design. Hugging Face's Hub documentation describes Xet deduplication as splitting files into content-defined chunks of about 64 KB, with boundaries chosen by a rolling hash over the file's own bytes. Deduplication wants small chunks, and small chunks are unforgiving: a 200 GB repository at 64 KB per chunk is roughly 3.1 million chunks, and 3 million stored objects costs more in requests and metadata than in bytes.

The Xet team's February 2025 post "From Chunks to Blocks" states the trade-off directly, that smaller chunks maximise deduplication while adding infrastructure overhead. The deduplication unit and the billing unit are allowed to be different sizes, and separating them is the design.

Failure scenarios

  • The short-task invoice. Compute three to four times the modelled figure, with correct rates and high utilisation, and nothing visibly wrong on any dashboard.
  • The archive that cost money. Tiering 400 million small objects raises the bill, because transition requests plus per-object floors exceed the per-gigabyte saving.
  • The early deletion charge. Retention deletes inside the minimum storage duration, so each deletion bills the remainder of the window.

Trade-offs

Choose Gains Pays
Batch up to the increment Meter and work align Latency, and a pool that costs money when idle
Keep fine-grained work Simplicity and isolation A billed-to-performed ratio of 3x or worse

When not to use it

Do not instrument granularity for meters that are a rounding error in your estate. If object-storage requests cost $12 a month, do not model them. The analysis earns its keep where the work is smaller than the meter and the volume is high - short jobs, small objects, brief connections, tiny messages. For long-running tasks on large objects the rate genuinely is the whole story.

Interview question

Q: A team's CI bill is four times the total runtime of its jobs, every rate checks out against the price list, and utilisation is high. Find the error, then tell me what you would change and what the change costs them.

What a strong answer covers: rate versus increment; that the meter includes boot, image pull and shutdown; the arithmetic showing a 25-second job billed at 60 to 105 seconds; a warm pool rather than a cheaper instance; and the cost of that fix, which is pool idle time plus the isolation a fresh machine was providing. The best answers add the second-order effect: once jobs share a machine, one job's resource use becomes another's flakiness.

Quick check

Quiz: Your model prices 400 million 4 KB objects a month at the per-gigabyte rate for a cold class. Which two meters are missing? Per-request charges for transitions and reads, and the per-object metadata floor; for 4 KB objects the floor alone can exceed the byte charge by more than an order of magnitude.

Flashcard: What two properties does every billing meter have, and which does the price page lead with? — The rate, and the increment plus minimum it applies to. The page leads with the rate, and the increment is what makes a model wrong by a multiplier rather than a percentage.