Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
3 to work through
-
advanced
A data platform's batch jobs have complex interdependencies and a failure early in the night cascades. What orchestration properties matter?
2 min answer -
advanced
On the last Sunday in October a UK data platform's 01:30 daily job ran twice and duplicated rows in a revenue table. On the last Sunday in March the same job did not run at all. The scheduler is configured in local time. What failed, and what is the structural fix?
3 min answer -
advanced
The batch orchestrator's scheduler process dies at 02:40 with about 30 tasks already running on workers. It is restarted at 03:00. The nightly load includes a MERGE into a revenue table and a push to a payments vendor. What happens between 02:40 and 03:15?
2 min answer
6 terms in this topic
Backfill
Reprocessing historical periods through a pipeline after fixing a defect or adding a field, at a scale the pipeline was not sized for.
conceptBatch Orchestration
Running interdependent data jobs in the right order with retries, backfills and failure handling, expressed as a dependency graph rather than a schedule.
conceptDAG Dependency
The declared edge stating that one task must not start until another has succeeded, and the difference between that and merely running later.
practiceLogical Date Partitioning
Making every batch task a pure function of a logical date, so that any run can be repeated safely and a backfill is a range of independent executions.
conceptWall-Clock Schedule Hazard
The class of batch faults caused by triggering work on a civil clock that repeats, skips or shifts, producing duplicate or missing runs that no job r…
conceptZombie Task
A task whose worker has stopped reporting to the scheduler but whose process may still be running and writing, so the orchestrator retries work that …
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.