An engineering director tells you that team A has a change failure rate of 2% and team B's is 18%, and asks you to find out what team A is doing right so the practice can be spread. Walk me through how you would answer.
Show the full answer Hide the answer
What the interviewer is testing
Whether you audit a metric before acting on it, and whether you can tell a stakeholder their comparison is invalid without sounding evasive. The trap is that both numbers are probably correct and the comparison is still meaningless.
The clarifying questions that change the answer
- What counts as a failure? An incident ticket, a rollback, any unplanned remediation, or a gate that reverted a canary automatically?
- Who records it? A human filing a ticket, or the pipeline recording its own reverts?
- What is the denominator? Deploys, commits, or user-visible releases?
- How often does each team deploy, and how much is in each change?
Those four answers usually resolve the whole puzzle before any practice is examined.
The arc of a strong answer
Here is the shape it nearly always takes. Team B canaries every change and auto-reverts on a gate, and every auto-revert is recorded as a failed change. Team A deploys straight to production and records a failure only when a human files an incident. Team B is measuring attempts; team A is measuring complaints.
Put numbers on it. Team B at 40 deploys a day with 18% is about 7 auto-reverts a day, each stopped at roughly 1% of traffic within a couple of minutes, so user-visible impact is seconds of degraded experience for a small cohort. Team A at 2 deploys a week with 2% is about one failure every five weeks, found by customers, at full traffic, lasting as long as it takes someone to notice. Team A's measured rate is nine times better and its users are worse off.
The general mechanism: better rollout automation raises measured change failure rate while lowering harm, because automation converts invisible failures into recorded ones. Any metric whose numerator is defined locally will reward the team that measures itself least.
So the answer to the director is not "team B is fine". It is: normalise the definition first, then compare the thing they actually care about, which is minutes of error-budget burn per week, a quantity both teams measure the same way because it is derived from user-facing signals. Keep change failure rate as a within-team trend line and stop comparing it across teams with different rollout automation.
Common weak answers
- "Interview team A and document their practices." Produces a confident write-up of practices that are not causing the number, and the recommendation will be to deploy less often.
- "Change failure rate is a bad metric." It is a useful metric with a locally defined numerator. The senior move is to fix the definition, not to dismiss the measure.
- "Average the two and set a target." A target on a metric with inconsistent definitions is a target on reporting behaviour.
What a strong answer adds
The organisational second-order effect, stated out loud: if this comparison is published, team B will turn off its auto-revert, because the cheapest way to improve the number is to stop recording reverts. Naming that before the dashboard ships is the difference between an analyst and an engineering leader.
The follow-on is a definition the whole organisation can use: a failed change is one that required unplanned remediation and reached users, which counts team A's incidents and excludes team B's automated catches without pretending the catches were free. Each auto-revert still costs an engineer roughly 30 minutes of rework, so the catches belong on a separate line rather than in the same number.
Decision rule: compare a team's change failure rate only against its own history, unless both teams use the same numerator definition and the same rollout automation. For anything cross-team, use error-budget burn, which is derived from user-facing signals and cannot be improved by recording less.