Pinterest has described replacing Pinball - its own open-sourced orchestrator - with Spinner, an internal platform built on a branch of Apache Airflow, moving more than 3000 workflows and 45000 tasks with translation tooling, workflow tiers and an in-house Kubernetes executor (Airflow Summit 2021). What forced the change, what did they take on, and where would copying it be a mistake?
Show the full answer Hide the answer
The situation they were in
Pinball was built when nothing suitable existed, and it worked for years. Pinterest's account of the move names scalability and performance limits as execution demand grew - but the structural cost is the one every in-house orchestrator eventually presents: every operator, every source integration, every UI affordance and every scheduler fix is yours alone to write. The ecosystem tool gets a connector contributed by someone else; the in-house tool gets a ticket.
What they chose
Not "adopt Airflow" but adopt Airflow and keep control of it: an internal platform on a branch of the project, a translation layer from the old definitions, mass-migration and verification tooling, tiering and prioritisation so workflows of different importance do not share one queue, and their own Kubernetes executor so each task gets its own pod and its own resources.
Why it fit their constraints
With 3,000 workflows and 45,000 tasks written by many teams, the migration cost is user migration, not scheduler migration. That reframing explains every choice: translation tooling exists because asking hundreds of engineers to rewrite definitions does not finish; schedule-semantics alignment exists because the old system expressed cadence differently and a silent shift duplicates or skips a night's run; per-task resource provisioning exists because a task that was fine inside a shared worker fails alone in a pod.
What it cost them
- A fork is a standing tax. Every upstream release is a rebase decision, and the alternative is drifting onto a version that no longer receives fixes.
- One control plane for everyone makes the scheduler a tier-1 service with its own on-call, its own metadata database as a shared bottleneck, and tiering as a permanent governance conversation about whose workflow is important.
- Migration verification is a product. Proving 3,000 workflows do the same thing after the move needs tooling nobody wants to own afterwards.
When copying this is the wrong move
- A team with 40 nightly jobs. Managed Airflow, or the warehouse's own scheduler, until there are several teams and thousands of tasks. The platform Pinterest built answers a problem created by scale and team count.
- Forking. The branch is defensible when you have engineers dedicated to the orchestrator and requirements upstream does not serve. Fork at 40 jobs and you have bought the in-house-tool problem back, with extra steps.
- Taking the Kubernetes executor as the default. Per-task pods buy isolation and pay start-up latency - seconds to tens of seconds - which is wrong for thousands of short tasks and right for heterogeneous heavy ones.
The general lesson
The decision was not which scheduler but who pays for change. In-house means you pay for every feature; upstream means you pay in migration and in the constraints of someone else's model; a fork means you pay a smaller amount forever. Pick knowingly, and write down which one you chose.