pattern

Provisioning Quota

also called Creation-Time Limit, Control-Plane Spend Cap

A hard limit on how much of a resource an account may create, enforced by the control plane at request time, which is the only cost control that acts before billing data exists.

cost-governancequotaspreventive-controlbilling-lagrunaway

At 02:14 on a Saturday a storage-triggered function begins writing its output back into the prefix that triggers it. Each write creates an invocation and each invocation creates a write. By Monday the account has logged hundreds of millions of invocations against a monthly forecast in the hundreds of dollars.

Every detective control behaved correctly and arrived late. The budget alert fired once enough spend had been metered, consolidated and published to make the threshold true, which is hours to days after the money was committed. Detective cost controls cannot be fast, because they consume billing data, and billing data is a downstream artefact of spending. A provisioning quota sits upstream: the control plane evaluates it when the request to create something is made, from state it already holds.

Why it matters

A runaway doubles spend on its own timescale, which may be minutes. A daily budget evaluation gets at best one chance to see that, and a monthly alarm at 100% of forecast fires after the month's money has gone. Cost telemetry lag sets a floor on how fast any detective control can react, and no alert tuning moves that floor.

Quotas also change who has to be awake. A limit that denies the 10001st concurrent execution needs no on-call engineer and no judgement about whether the spike is legitimate. It converts an unbounded financial exposure into a bounded capacity failure, which engineering teams already know how to handle.

Implementation patterns

  • Quota the dimension that multiplies: concurrency, instance count, accelerator count, resource creation rate. Spend is the wrong dimension, because spend is what you learn about late.
  • Set each quota above the rehearsed peak from a load test or a known seasonal peak, not above the average.
  • Make raising it a ten-minute self-service action with an audit record. A quota behind a ticket queue becomes the thing that blocks incident response, and is then removed permanently.
  • Quota per blast-radius boundary - account, project or cluster. An organisation-wide ceiling is consumed by whichever workload runs away first.
  • Separate classes of spend, so incident response, production serving and experimentation do not share a ceiling. The spend you most want to deny is the one you least want denied during an outage.

Industry example

Hosted inference platforms are built on this pattern rather than on budget alerts. Capacity is scarce, expensive and shared, so access in production is governed by per-organisation request rates, concurrency limits and compute ceilings enforced at admission, with raised limits available through a defined process. A customer whose integration enters a retry loop is rate limited within milliseconds, and exposure is bounded by the limit rather than by how long somebody takes to notice an invoice.

The same architecture underlies every cloud provider's default service limits. What is unusual is applying it deliberately inside your own estate rather than inheriting the defaults.

Failure scenarios

  • The quota that denies the fix. An incident needs capacity, the limit is reached, the raise path is a ticket, and the control caused the outage it was meant to prevent.
  • One shared ceiling. A batch experiment consumes the organisation's whole concurrency allowance and production serving is denied.

Trade-offs

A quota trades availability for a bounded bill. It will sometimes deny legitimate work - a delayed batch, a shed request - and in exchange the worst case stops being unbounded.

The second cost is operational: quotas are state that must be owned and reviewed, and one with no owner degrades into a ceiling somebody raises to an enormous number to silence an error message.

When not to use it

For spend that cannot run away, a quota adds friction and prevents nothing. A fixed reserved fleet, a flat-rate licence or a committed contract is already bounded, and the useful control there is the purchasing decision.

A quota is also wrong for gradual overspend: spend creeping up 5% a month is an allocation problem, and a quota loose enough not to interfere never triggers. Use quotas against unbounded and fast, and allocation plus review against gradual and structural.

Interview question

Q: A recursive function has produced a six-figure surprise invoice over a weekend. Leadership asks for better budget alerts. Explain why that is the wrong request and specify what you would put in place instead.

What a strong answer covers: that billing lag makes every alert detective; that the control belongs in the control plane at creation or admission; the dimensions to limit and why spend is not one of them; per-account scoping; separate classes so incident response is not denied; a self-service raise path with an audit trail; and a fix for the specific defect, a trigger pattern that can feed itself.

Quick check

Quiz: Why is "spend per hour" a poor dimension for a hard quota? Because spend is derived from billing data, which lags the activity producing it, so a spend-denominated limit inherits the latency of the alert it was meant to replace; limit the countable thing that multiplies instead.

Flashcard: Which cost control can stop a runaway, and why can no alert do it? — A creation-time or admission-time quota in the control plane, evaluated from state it already holds. Alerts consume billing data, which lags spend by hours to days, so by construction they report rather than prevent.