advanced 3 min answer

A platform group of 12 serves 45 product teams through one intake queue of about 70 requests a week, with a median lead time of 6 working days. Leadership has approved restructuring into capability-owning teams with self-service interfaces. Plan the change so the queue is never dropped and you can stop at any stage.

platform teamticket opsself-servicereorganisationsequencing
Show the full answer Hide the answer

The sequence

  1. Classify two weeks of the queue by request class and count it. Expect the top three classes to be somewhere around two thirds of the volume. Costs nothing, changes nothing, and it is the only input that makes the rest of the plan arguable rather than ideological.
  2. Fix the intake contract before touching the organisation. One channel, a request template per class, and published response times. This makes the queue measurable and kills the side channels that hide demand. Fully reversible.
  3. Put two of the twelve on a weekly duty rota that absorbs all intake. The other ten can then hold a block of uninterrupted build time. Write the rota's scope down, including what it refuses, or it becomes a permanent dumping ground. Reversible week to week.
  4. Automate the largest class behind a self-service interface while the ticket path keeps working. Run both in parallel and publish the share of that class served without a ticket.
  5. Retire the ticket path for one class only when it is above roughly 80% self-served, with a named owner for the remainder. Per class, not per programme.
  6. Re-form the teams last. Team boundaries should follow the interfaces that now exist. Drawing the org chart first produces capability teams that own a name and a queue.

Where it diverges, and how you would know

Shadow demand is the failure you cannot see in the queue, because it consists of requests that stopped arriving. Teams quietly build their own provisioning scripts, and the queue graph falls while the estate fragments. Detect it by asking a sample of teams what they stopped asking for, and by looking for duplicate tooling in team repositories - that search finds more than any dashboard.

The second divergence is the duty rota becoming the new permanent job for the same two people, which shows up as those two never appearing in build work for a quarter. Rotate it and cap it.

The point of no return

Per class, it is retiring the ticket path, because the people who ran it have moved on to other work and the tacit knowledge decays within weeks. Everything before that is reversible by reopening the channel. There is no single point of no return for the whole programme, and designing it that way is the main reason it can be stopped.

The rollback at each stage

Stages 1 and 2 need no rollback. Stage 3 rolls back by ending the rota. Stage 4 rolls back by routing the self-service path's failures to the still-live ticket queue - which is also how you discover its gaps. Stage 5 rolls back only by reopening an intake for that class, which costs credibility, so the 80% threshold exists to make it unnecessary. Stage 6 is the expensive one to undo, which is why it is last.

How long it really takes

Per class: two weeks to measure, four to eight weeks to build and parallel-run, four weeks to retire. Three classes is therefore two to three quarters, with the first class slowest because the self-service pattern is being invented. The queue does not shrink in the first two months - it becomes visible, which usually looks like it got worse.

When this plan is the wrong answer

If the queue is 70 requests a week of genuinely varied one-off work, there is no top class to automate and this plan has nothing to bite on. The fix there is scope reduction: stop offering some of it. And with four platform engineers rather than twelve, the duty rota consumes half the team, so the honest move is to publish a narrower catalogue of what the platform does at all.