intermediate 2 min answer

An engineering organisation reports velocity, deployment frequency and story points. Leadership cannot tell whether the investment is working. What is missing?

chargebeeoutcomesmetricsdeliveryvalue
Show the full answer Hide the answer

What is missing

Every one of those is an output measure. They describe how much work was done, not whether it achieved anything. A team can double its deployment frequency while shipping features nobody uses, and the metrics will look excellent.

What to measure instead

  • The business outcome each investment was meant to change, stated before the work starts. Reduced churn, faster onboarding, increased conversion, lower cost per transaction. The discipline is naming the number and the expected direction in advance, because retrospective justification always finds something that improved.
  • Adoption of what was shipped, which is the cheapest and most neglected measure. A feature used by a handful of accounts is a signal about the roadmap process, not about engineering.
  • Time from decision to customer impact, which is the flow metric that actually matters and which deployment frequency only partially proxies.

Why output metrics still have a place

They are diagnostic rather than evaluative. Deployment frequency, lead time, change failure rate and recovery time are excellent indicators of delivery system health, and a team that cannot deploy frequently cannot respond to what outcome measurement tells them.

The error is presenting them to leadership as evidence of value, which invites the reasonable response that none of it appears in the business results.

The structural requirement

A hypothesis per significant investment: what we believe, what we will measure, what result would cause us to stop. Without the last clause, nothing is ever stopped, and the portfolio accumulates work that is not delivering.

For a subscription business the natural outcomes are well defined — activation, expansion, retention, cost to serve — which makes this easier than in many domains and therefore less excusable to omit.

The cultural risk

Outcome measurement used punitively produces gaming and defensive target-setting. Teams pick outcomes they can guarantee, which are by definition the ones least worth pursuing. The framing has to be that a disproved hypothesis is a successful experiment, and that requires leadership to visibly treat it that way at least once.