Process Ownership Gap
also called Unowned Workflow, Orphaned Business Process
The condition in which every step of a multi-service process has an owning team but the process itself has none, so a stall produces no error, no deadline breach and no page.
Order fulfilment runs across six services connected by events. Each team owns its step, each step works, and every service meets its availability target. Last quarter about 0.4% of orders ended in a state nobody noticed for days.
No service failed. What is missing is not a safeguard but a component: nobody owns the statement "this order reached a terminal state". Because no one owns it, nothing defines the terminal state, nothing has a deadline, and no rota is responsible. Nothing is ever late, because lateness is a property of a process and the process does not exist as a thing.
This is the characteristic failure of choreography adopted by default. Choreography distributes the happy path well and does not distribute responsibility for the whole.
Why it matters
The failure is invisible to every standard signal. Error rates are normal, because nothing errored. Queue depths are normal, because the message that would have continued the process was never sent. Dead-letter queues are empty, because the absence of a message is not a message. All six dashboards are green.
The cost scales with volume and is paid by people. At 20,000 orders a day, 0.4% is 80 stuck orders a day, a full-time support role plus the customer contacts it generates. At 200 orders a day it is a weekly report, and the report may genuinely be the right response.
A process with no owner accumulates no deadlines, no state machine and no runbook, so every stuck-order investigation starts from nothing.
Implementation patterns
- Name the process and give it a terminal state in code, not in a diagram. If it exists only in people's heads there is nothing to alert against, and defining it is the first deliverable.
- Give every step a deadline. This is the detection mechanism: a step that has not produced its completion event inside its deadline raises an alert naming the instance and the step.
- Introduce a process manager that owns sequencing and timeouts and nothing else - no pricing rules, no inventory logic - and keep each service's internal work choreographed.
- Run it in shadow for two weeks first: consume the events that already flow, assert the expected sequence, alert on violations, drive nothing. You learn the real distribution of stuck states before the coordinator becomes a release dependency for six teams.
- Assign the page to a team that owns a business outcome. If the only candidate is the platform team, the finding is that the process has no business owner.
Industry example
Published multi-service workflow write-ups from 2018 onward converge on the same remedy, which is why durable-execution and workflow engines exist as a product category: the gap they sell into is not message delivery, which brokers already solve, but process-level state and timeouts, which nothing in a pure event mesh provides. The archetype in the other direction is a fulfilment platform where answering "where is this order" means an engineer querying six services by hand - a procedure that works, costs an hour each time, and is never written down because no team owns the question.
Failure scenarios
- A step's completion event is never emitted after a deploy changes a condition, and the process stops with every service healthy.
- A terminal state nobody defined, so instances parked in a valid intermediate state forever look identical to instances in progress.
- Compensation that nobody owns, so a half-completed process is neither finished nor undone.
- Support becomes the monitoring system, so detection latency is however long a customer waits before complaining.
- The process manager given authority on day one with guessed deadlines, generating enough false alerts to be switched off.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Pure choreography | Teams ship independently; no coordinator to operate | No process-level state, deadlines or owner |
| Process manager for sequencing only | Deadlines, a terminal state and a page that fires | A new stateful service that is down when the process is down |
The process manager becomes a dependency of six teams' releases and sits in the critical path, and its deadlines will be wrong at first. That is the argument for the shadow period, not against the component.
When not to use it
Where the process is a fire-and-forget notification rather than a business transaction, there is nothing to own. If no instance has a required outcome, a stall is not a defect and a coordinator is cost with no benefit.
It also does not apply where the whole process sits inside one team and one database, where the honest answer is a single transaction or a state column. At low volume the economics favour a report: a handful of stuck instances a week handled by a weekly review is cheaper than a new stateful service.
Interview question
Q: A process spans six services with events and no coordinator, and a small percentage of instances silently stall. I do not want to hear "add tracing". What structural change do you make, and who do you page?
What a strong answer covers: that the missing thing is a component rather than a safeguard - a terminal state, per-step deadlines and an owner; that the deadline is the detection mechanism while tracing only shortens an investigation someone already started; the shadow-first rollout; the volume arithmetic deciding whether this warrants a service or a report; and that the page must land on a team owning a business outcome.
Quick check
Quiz: Why does a dead-letter queue not catch a stalled choreographed process? Answer: A dead-letter queue catches messages that failed, and this process stopped because an expected message was never sent - the absence of a message is not a message.
Flashcard: What three attributes does an unowned multi-service process lack? - A terminal state defined in code, a deadline per step, and one team that owns the whole process and gets paged when it stalls.