Guarantee Surface
also called Promise Surface, Service Guarantee Footprint
The set of customer-facing promises a product has made, each of which obliges the architecture to detect its own breach and pay for it - so this surface, not the feature count, drives a whole class of build cost.
A product manager changes one sentence in an app: "delivered in 30 minutes" becomes "delivered in 30 minutes or it is free". No feature has been added. What has been added is a liability the system must now detect, price, pay and defend, and that work is larger than the feature the sentence describes.
The guarantee surface is the count and shape of such promises. It grows one sentence at a time, usually in marketing reviews, and is rarely represented anywhere in the architecture.
Why it matters
A promise stated as a percentile buys a measurement obligation; the same promise stated as a guarantee buys a payout pipeline. The difference is invisible in a product brief and decisive in the build.
Cost in this class scales with the number of promises, not the number of features, because each promise needs its own detection rule, remedy, accounting treatment and abuse limit. And each is close to permanent, since withdrawing a published guarantee is a public act, so the surface ratchets upward unless somebody owns it. Organisations that never name it end up with six promises, three of which nobody can prove they are meeting.
Implementation patterns
- Keep a register of live promises with, for each: the trigger definition, the clock, the exclusions, the remedy, the owner and the monthly cost. One page, reviewed quarterly.
- Define the clock before anything else. Order placed, payment captured or picker assigned are three different metrics, differing by the minutes that decide breach.
- Detect breach without the customer claiming it. A remedy that requires a claim is a promise in marketing and a complaint queue in operations.
- Accrue the liability. Breach detection must write to a ledger that reconciles with the payment processor and the accounting system, or the cost appears a month late as an unexplained variance.
- Cap and flag every remedy. A per-line price ceiling for substitutions, a per-customer monthly limit for credits, and a kill switch per promise. A guarantee is an incentive to game the trigger, so assume some traffic will optimise against it.
- Report per cohort, not in aggregate, because one slow city hides comfortably inside a global percentile, and ship behind a rule you can narrow, starting where the promise is cheapest to keep.
Industry example
Grocery-delivery economics in the mould of Instacart make the arithmetic concrete. "Substituted automatically and never charged more" means the basket is re-priced at pick time and the difference is absorbed. On explicit assumptions — a 25-line basket, 12% of lines substituted, the substitute averaging 8% more on a £4 item — that is roughly £1 per order, about 1% of a £100 basket and a large share of a thin grocery margin. These are illustrative figures rather than measurements, which is the point: the promise is affordable only inside a measured range of substitution rate and price delta.
Failure scenarios
- A guarantee shipped as a copy change, with no breach detection, discovered when customers start claiming it by email.
- Two promises that conflict — a delivery window and a substitution guarantee — where meeting one breaches the other and nothing says which wins.
- Abuse at the trigger, where a small number of accounts learn to produce breaches, or a percentile nobody can evidence when a large customer asks for the history.
- Silent degradation in one cohort, invisible in the aggregate, until a review finds a city that has been breaching for months.
Trade-offs
Guarantees convert uncertainty into a competitive claim and generally do increase conversion; they pay in permanent build and permanent liability. Percentile statements are cheap to make and weak to sell.
| Choose | Gains | Pays |
|---|---|---|
| Guarantee | Sharper proposition and clearer accountability | Detection and payout pipeline plus an accrued liability and fraud surface |
| Percentile | Cheap to state and easy to revise | Measurement and reporting obligation with less commercial force |
When not to use it
Do not add a guarantee while the operation is still changing weekly. A guarantee freezes a process you have not stabilised, and early markets breach it for reasons nobody controls yet. State a percentile, publish it honestly, and convert later.
And do not treat the surface as an architecture problem when it is a supply problem. If 20% of lines are unavailable because stock data is hours stale, a substitution promise automates apologies; fix the inventory feed. The honest answer in that case is that no design helps and the constraint is operational.
Interview question
Q: Marketing wants to change "usually within two hours" to "within two hours or your delivery fee is refunded". You have four weeks. What do you build, what do you refuse, and what do you need measured before you agree?
What a strong answer covers: the clock definition and exclusions as the first artefact · automatic breach detection rather than a claims process · the accrual and reconciliation path, where the real cost sits · per-customer caps and a kill switch · the current breach rate by cohort, without which the liability estimate is meaningless · and a narrower promise by region or slot that can be widened once measured.
Quick check
Quiz: What does "30 minutes or it is free" require that "95% within 30 minutes" does not? Breach detection on every order, an automatic remedy, an accounting accrual, abuse limits and a dispute path - a payout pipeline rather than a reporting obligation.
Flashcard: What drives the cost of this class of work? — The number of live promises, not the number of features, and each one is effectively permanent because withdrawing a guarantee is a public act.