concept

Toil

Manual, repetitive, automatable operational work that scales with service growth and produces no lasting improvement.

The defining characteristics: manual, repetitive, automatable, tactical, devoid of enduring value, and growing at least linearly with the size of the service. Work that is merely tedious but non-recurring is not toil; work that is manual but requires genuine judgement is not toil either.

Why it matters as a measured quantity rather than a complaint: toil scales with growth, so a team that spends 30% of its time on it will spend more as the service succeeds, until there is no capacity left for the engineering that would reduce it. That is the trap, and it closes quietly.

The conventional ceiling is 50% of a team's time, with the rest protected for engineering work. The number matters less than measuring it at all — most teams have no idea what theirs is, and the estimate they give is consistently lower than the measurement.

Sources to look for: manual deployments and rollbacks, access requests, certificate renewals, capacity adjustments, routine restarts, manual data corrections, repeated incident responses to the same cause, and reports produced by hand.

The organisational point: if reducing toil is not funded, on-call load grows, engineers leave, and the knowledge leaves with them. Reliability culture is largely the willingness to spend engineering time on work that produces no feature.