Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
5 to work through
-
intermediate
An interviewer says - we run one orchestrator with about 4,000 tasks a night, all written by the platform team. Leadership wants to open it to 12 product teams as self-service so they stop queueing behind us. Where do you take this?
3 min answer -
intermediate
Every task in the data platform sits in a queued state. Worker CPU is around 6%, the metadata database is healthy, the task logs show nothing, and no task has failed. Someone proposes adding workers. What do you look at first, and what is probably happening?
2 min answer -
intermediate
What should be evaluated when choosing a workflow scheduler for a data platform?
2 min answer -
intermediate Multiple choice
Why is cron insufficient for a pipeline of interdependent data jobs?
2 min answer -
advanced
Pinterest has described replacing Pinball - its own open-sourced orchestrator - with Spinner, an internal platform built on a branch of Apache Airflow, moving more than 3000 workflows and 45000 tasks with translation tooling, workflow tiers and an in-house Kubernetes executor (Airflow Summit 2021). What forced the change, what did they take on, and where would copying it be a mistake?
2 min answer
3 terms in this topic
Airflow: A Scheduler Born From Pipeline Sprawl
Airbnb built Airflow because cron cannot express dependencies, and a data platform's failures are mostly dependency failures.
conceptScheduler Control Plane
The orchestrator's own state and availability, which becomes a critical dependency for every pipeline it runs.
toolWorkflow Schedulers
Tools that run interdependent tasks in the right order with retries and visibility — where the value is the explicit dependency graph, not the schedule.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.