metric

Change Failure Rate

The proportion of deployments to production that cause a degradation requiring remediation.

doraqualitydelivery

The DORA metric that balances the speed ones, and the reason the four must be read together: lead time and deployment frequency can always be improved by taking more risk, and change failure rate is what makes that visible.

Its definition needs local precision before it can be tracked, because "failure" is ambiguous. Anything requiring a rollback, a hotfix, or a forward fix outside the normal cadence is the usual line. A rollback triggered automatically by a canary that caught a defect is a judgement call — it represents the system working, though it still represents a change that should not have got that far.

The counter-intuitive empirical finding, consistent across years of the research, is that teams deploying most frequently have lower change failure rates, not higher. The mechanism is batch size: small changes are easier to review, to reason about and to diagnose, and they fail less often per change. This is the single most useful piece of evidence to bring to a governance function that believes fewer, larger, more heavily reviewed releases are safer.

The number to be suspicious of is a very low rate combined with a low deployment frequency, which usually indicates unrecorded failures rather than reliability.