Outcome Measurement
Determining whether an architectural change achieved what it was meant to — the step most often skipped, which is why the same mistakes recur.
Definition
Outcome measurement closes the loop: the change was made to achieve something, and this is how we know whether it did. Without it, architecture is a sequence of confident assertions with no feedback.
Why it is skipped
The work is finished, the team has moved on, and measuring risks discovering that the expensive thing did not help. That risk is precisely the reason to do it — the alternative is an organisation that repeats decisions without learning from them.
What to do before the work starts
State the expected outcome and its measure, in the design document. "This migration should reduce p99 checkout latency from 1.2 s to under 500 ms and reduce infrastructure cost per order by 20%. We will measure at 30 and 90 days."
Written afterwards, the outcome is chosen to fit whatever happened. Written beforehand, it is a prediction, and predictions are how anyone learns anything.
Also state what would indicate failure, and what you would do about it. A change with no failure criterion cannot be evaluated.
What to measure
- The intended outcome, in the terms originally stated.
- The counterfactual, where possible. Did the metric move because of this, or because traffic changed, or because something else shipped that week? A holdout or a staged rollout gives a comparison; without one, attribution is a story.
- Unintended effects. Cost, latency elsewhere, incident rate, developer experience. Improvements in one dimension frequently move cost into another.
- Whether the assumptions held. The capacity model, the traffic forecast, the cost per unit.
Making it happen
- A review scheduled at design time, in the calendar, with a named owner.
- A short written record of what was predicted and what occurred — a paragraph, not a report.
- Honest reporting of the disappointing ones. An organisation that only records successes learns nothing, and everyone knows it.
- Feeding results into the next decision. "Last time we assumed X and observed Y" is the most valuable input to any estimate.
Failure scenarios
- No measure defined, so success is asserted.
- Measured immediately, before the effect could appear or before the novelty wore off.
- No counterfactual, so an unrelated change gets the credit.
- Only successes recorded, which destroys the credibility of the whole practice.
- The result never influencing anything.
Interview question
"How would you determine whether a migration you completed six months ago was worth doing?"