An interviewer says - we run one orchestrator with about 4,000 tasks a night, all written by the platform team. Leadership wants to open it to 12 product teams as self-service so they stop queueing behind us. Where do you take this?
Show the full answer Hide the answer
What the interviewer is testing
Whether you treat an orchestrator as a multi-tenant platform rather than as a tool. The failure modes of self-service scheduling are shared-resource failure modes, and a candidate who talks only about DAG authoring conventions has missed the whole question.
The clarifying questions that change the answer
- Does task code run inside the scheduler's process or in an isolated executor? If it runs in-process, self-service means 12 teams sharing one Python environment, and the first dependency conflict is a platform outage.
- Does the scheduler parse every pipeline definition on a loop? This is the one people miss. A shared parse loop means one team's slow top-level import delays scheduling for everybody, and the symptom is every task queued while workers sit near idle.
- Who is paged at 03:00 for a failed task? If the answer is still the platform team, this is not self-service, it is the same work with more meetings.
- What is the freshness commitment of the slowest downstream consumer? That sets whether per-team queues can be allowed to grow or must be capped.
A strong answer's arc
- Name the shared resources: the scheduling loop, the metadata database, the worker pool, and the concurrency limits on shared sinks such as the warehouse. Every tenancy control maps to one of these.
- Isolate execution. One container image per team, so dependencies are a team's own problem.
- Quota each shared resource. A per-team concurrency pool, a default task timeout, a cap on tasks per pipeline, and a limit on definition-parse time enforced by the platform.
- Ship a paved path rather than a blank orchestrator. A scaffold with the team's pool, ownership tag, alert routing and a tested template. Ownership metadata is mandatory, because alert routing is what makes the on-call boundary real.
- Publish the platform's own commitment: what uptime the control plane has, what happens on a scheduler restart, and which classes of task are unsafe to retry.
Common weak answers
- "Add workers." The characteristic self-service failure is a scheduling-loop or metadata database bottleneck, where workers are idle. Adding workers costs money and changes nothing.
- "Give each team its own orchestrator." Twelve control planes and twelve upgrade paths, and cross-team dependencies become unexpressible, which removes the reason the orchestrator exists. It is a defensible answer only when teams genuinely share no data.
- "Use scheduled jobs in Kubernetes." That deletes the dependency graph, which is the product.
What a strong answer adds
The cost curve goes the wrong way first. Self-service moves the platform team from writing pipelines to reviewing, supporting and operating other people's, and for the first two quarters expect roughly a quarter of platform capacity on support before the queue actually shortens. Saying that out loud, with a plan to measure it, is the senior layer.
When this is the wrong answer
With 3 teams and 300 tasks a night, building a tenancy model costs more than the queue it removes. A shared repository with a code-review gate and a naming convention is the right answer until either the review queue or the blast radius of a bad pipeline becomes the visible pain.