DORA Metrics
Four measures of software delivery — deployment frequency, lead time, change failure rate, time to restore — that resist gaming better than most.
Definition
- Deployment frequency — how often you release to production.
- Lead time for change — commit to running in production.
- Change failure rate — proportion of deployments causing a degradation.
- Time to restore service — how long recovery takes.
The first two measure throughput; the last two measure stability.
Why the set is well chosen
The central finding of the research behind them is that throughput and stability move together rather than trading off. Teams that deploy frequently with short lead times also have lower failure rates and recover faster. That contradicts the intuition that speed costs safety, and the mechanism is batch size: small changes are easier to review, verify, understand and revert.
The set also resists gaming reasonably well because the four constrain each other. Deploying more often to improve frequency raises the failure rate if quality drops. Deploying rarely to protect the failure rate destroys frequency and lead time.
How to use them properly
- Trend, not benchmark. Comparing your team to published elite figures generates anxiety and no information. Compare to yourselves three months ago.
- Per team and per service, not organisation-wide. An average across a hundred teams describes none of them.
- As a diagnostic, not a target. The question is "what is preventing us from deploying more often", and the answer is usually specific: a manual approval, a slow test suite, a schema migration process, a shared environment.
- Alongside outcome metrics. Delivering the wrong thing quickly is not success, and DORA measures delivery capability rather than value.
Where each number usually points
| Metric poor | Usual cause |
|---|---|
| Deployment frequency | Manual gates, long-lived branches, batched releases |
| Lead time | Slow pipeline, review latency, environment contention |
| Change failure rate | Weak testing, no canary, no automated rollback |
| Time to restore | No rollback, poor observability, unclear ownership |
Failure scenarios
- Made a target, so teams split deployments to inflate frequency.
- Averaged across the organisation, hiding both the excellent and the struggling.
- Measured without acting, so it becomes a reporting exercise.
- Treated as the only measure of engineering, ignoring quality, cost and outcome.
Interview question
"Your deployment frequency is monthly and your change failure rate is 30%. Which do you attack first and why?"