Provisioning a database on this platform takes about 4 engineer-hours of real work. The median request takes 6 working days end to end. The platform team of 5 is fully occupied. Why is the lead time roughly twelve times the work time and which change reduces it most?
Show the full answer Hide the answer
The mechanism
6 working days is about 48 working hours. 4 of those are work. The other 44 are waiting, and waiting has two separate causes.
Queueing. A fully occupied team is a queue, and wait time does not grow linearly with utilisation - it grows roughly in proportion to u/(1-u). At 90% utilisation an item waits about 9 times its own service time; at 95%, about 19 times. This is why a team that is "slightly too busy" produces lead times that look inexplicable: the last 5% of utilisation does more damage than the previous 50%.
Handoffs. Each round trip between requester and platform is bounded by how often the other side looks. A team that triages once a day adds up to a day per exchange, so three clarifying questions cost three days regardless of how fast anyone types. Lead time is set by the number of handoffs and the responder's batching interval, not by the difficulty of the work.
Why removing the human wins
Automating the common case attacks both causes at once. The request no longer enters a queue, so the u/(1-u) term disappears for that class of request, and there is nothing to hand off, so the batching intervals disappear too. Lead time collapses from days to the duration of the provisioning itself, typically minutes.
The precondition is that safety stops being a review and becomes a boundary: a fixed catalogue of sizes, enforced naming and tagging, encryption and backup on by default, a cost ceiling per request, and network placement chosen by the platform rather than the requester. Inside those bounds there is nothing for a human to decide, which is what makes approval removable rather than merely unpopular.
Why the other options fail
- Hire two engineers. This does work, and it works through utilisation: 5 engineers at 95% going to 7 at 80% cuts the queueing multiplier from roughly 19 to roughly 4. But the handoff days remain, demand grows with the estate, and you have bought a permanent two-engineer cost to make a queue shorter rather than removing it.
- A two-day SLA. An SLA changes reporting, not mechanics. Teams that are measured on breaches hold capacity in reserve to protect the number, which lowers throughput and lengthens the queue for everything not covered.
- One provisioning day a week. Batching raises average wait. A request arriving the day after provisioning day waits six days by construction, and the variance a requester plans around gets worse, not better.
- A more detailed form. Worth doing, and bounded by arithmetic: it attacks handoffs only, and even eliminating every clarification leaves the queue. The work was 4 hours of 48.
When this is the wrong answer
If the request arrives twice a month, automation is the expensive option. Building a safe self-service path for databases is commonly four to eight engineer-weeks, and at 24 requests a year it will not repay that in lead-time terms. Fix the form, publish the catalogue, and keep the human. The threshold is request frequency, not team size: self-service pays when the same decision is made often enough that making it once in code is cheaper than making it each time in a ticket.