metric

DORA Metrics

Four measures of software delivery — deployment frequency, lead time, change failure rate, time to restore — that resist gaming better than most.

doradeliverythroughputstabilitymeasurement

Definition

  • Deployment frequency — how often you release to production.
  • Lead time for change — commit to running in production.
  • Change failure rate — proportion of deployments causing a degradation.
  • Time to restore service — how long recovery takes.

The first two measure throughput; the last two measure stability.

Why the set is well chosen

The central finding of the research behind them is that throughput and stability move together rather than trading off. Teams that deploy frequently with short lead times also have lower failure rates and recover faster. That contradicts the intuition that speed costs safety, and the mechanism is batch size: small changes are easier to review, verify, understand and revert.

The set also resists gaming reasonably well because the four constrain each other. Deploying more often to improve frequency raises the failure rate if quality drops. Deploying rarely to protect the failure rate destroys frequency and lead time.

How to use them properly

  • Trend, not benchmark. Comparing your team to published elite figures generates anxiety and no information. Compare to yourselves three months ago.
  • Per team and per service, not organisation-wide. An average across a hundred teams describes none of them.
  • As a diagnostic, not a target. The question is "what is preventing us from deploying more often", and the answer is usually specific: a manual approval, a slow test suite, a schema migration process, a shared environment.
  • Alongside outcome metrics. Delivering the wrong thing quickly is not success, and DORA measures delivery capability rather than value.

Where each number usually points

Metric poor Usual cause
Deployment frequency Manual gates, long-lived branches, batched releases
Lead time Slow pipeline, review latency, environment contention
Change failure rate Weak testing, no canary, no automated rollback
Time to restore No rollback, poor observability, unclear ownership

Failure scenarios

  • Made a target, so teams split deployments to inflate frequency.
  • Averaged across the organisation, hiding both the excellent and the struggling.
  • Measured without acting, so it becomes a reporting exercise.
  • Treated as the only measure of engineering, ignoring quality, cost and outcome.

Interview question

"Your deployment frequency is monthly and your change failure rate is 30%. Which do you attack first and why?"