Non-Functional Budget
also called Quality Attribute Budget, Latency Budget Allocation
Splitting a system-level quality target into per-component allowances with named owners, so a change can be shown to have spent someone's allowance rather than merely argued about.
A checkout page has a p99 target of 300 ms. Eighteen months later it sits at 620 ms. Nobody shipped a slow change: recommendations added 40 ms, a fraud check 55 ms, a loyalty lookup 30 ms. Every one was approved by someone who checked that their own change was small. The target was nobody's to defend, so it was spent by consensus.
A non-functional budget is the accounting that makes this impossible. The system target is divided into named allowances with named owners, the allowances sum to less than the target, and a change that exceeds its allowance needs someone else to give theirs up. Availability decomposes multiplicatively down a chain, error budgets across services, cost per request. ISO/IEC 25010 — nine characteristics in its 2023 second edition — names which attributes exist; the budget turns one into something a build can fail on.
Why it matters
Quality attributes are consumed in every pull request and defended once, in a design review, before any of those pull requests exist. That asymmetry is the whole problem.
A budget makes the allowance the unit of review. The question stops being "is 40 ms a lot?" — to which the honest answer is always no — and becomes "recommendations have 35 ms and you want 75, so what gets cheaper?". That has answers: cache it, move it off the critical path, or accept a lower target with the product owner's signature. A budget that will not close also tells you the structure is wrong before you build it, because if the minimum plausible allowances exceed the target no optimisation gets there.
Implementation patterns
- Reserve headroom. Allocate about 70% of the target and hold the rest centrally. A budget allocated to 100% is over budget on its first surprise.
- Budget the tail. Allowances in p99, because means hide the component with a bimodal distribution.
- Account for fixed costs first: client network and TLS, edge, gateway, serialisation, response path. On a 300 ms mobile target these commonly take 80 to 120 ms before application code runs.
- Make the allowance a test: a CI assertion against a fixture, plus an alert when a component's p99 exceeds its allowance for an hour. Without one, the ledger is a spreadsheet.
- One owner per allowance, because an unowned allowance is spent first.
- Decompose availability multiplicatively. Four synchronous dependencies at 99.95% cap you near 99.8%, so that budget is really a limit on hop count.
Industry example
Etsy's Measure Anything, Measure Everything (Ian Malpass, Code as Craft, 15 February 2011) introduced StatsD, whose design is the precondition for any budget being enforceable. Metrics go out as fire-and-forget UDP packets, so the application never waits for the collector and does not fail when the collector is down; the daemon aggregates and flushes on a 10-second interval, and a sample rate lets a hot path emit cheaply.
The load-bearing property is the cost of instrumentation, not the storage. If adding a timer is one line and carries no risk, engineers instrument everything, including paths nobody planned to measure — and a budget can only be defended per component if numbers exist for the components nobody nominated.
Failure scenarios
- Allowances in averages. Every component is inside its mean allowance and p99 is double the target.
- No owner for the client. Every server budget is met and mobile users see three times the target, because 180 ms of radio latency was never in the ledger.
- Retries uncounted. A component inside its allowance per attempt spends triple on the 1% that retry twice — precisely the tail the target is about.
Trade-offs
What it buys: a defensible refusal. An engineer declines a change on evidence, and a product owner sees the price of a feature in the currency of the target.
What it pays: allocation is political, roughly a day for a first pass and a recurring argument each quarter. It can ossify structure — a team owning 35 ms optimises inside its allowance rather than proposing the hop be deleted, because that is somebody else's budget. The bill arrives at the first hard target change: tightening 300 ms to 150 ms is a short negotiation with a budget and a six-week investigation without one.
When not to use it
Do not budget an attribute nobody has measured, or one with no agreed target — an allowance derived from a guess imports the guess into every component. Below roughly three components on a path it is bureaucracy. Skip it entirely for attributes that do not decompose arithmetically: usability and testability have no per-component split, and their equivalent is a quality attribute scenario.
Interview question
Q: You have a 200 ms p99 target for a mobile API call that fans out to three services, one of which calls a partner. Allocate the budget and tell me what you would refuse.
What a strong answer covers: fixed costs first — radio, TLS, edge, gateway — which realistically leaves 60 to 100 ms of server time. Parallel fan-out costs the slowest branch rather than the sum, and the partner call is the branch you do not control, so it is refused inside this budget and replaced by a cached value with a freshness target. Then about 30% headroom, and the assertion and alert that make each allowance real.
Quick check
Quiz: A checkout p99 drifts from 300 ms to 620 ms over eighteen months with no single change responsible. Which control was missing? A per-component latency budget with owners and an automated assertion — the target existed but no component had an allowance to breach.
Flashcard: How much of a latency target should be allocated, and in which statistic? Roughly 70%, in p99 terms, with the rest held centrally as headroom.