Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
6 to work through
-
beginner
A team runs 40 SQL scripts nightly, numbered 01 through 40 and executed in order by a shell script. Someone proposes adopting a transformation framework. A senior analyst pushes back - the scripts already run in the right order, so what does a framework actually add? What is the honest answer, and when is the analyst right?
3 min answer -
intermediate
A platform adopts a transformation framework and analysts build hundreds of models. What governance is needed to prevent a tangle?
2 min answer -
intermediate
Review this transformation project. 300 models are all views with no materialisations and chains up to 11 deep. Every model has not-null and unique tests on every column - about 2400 tests. CI runs a full build plus the whole test suite on every pull request - 70 minutes. Each developer has a personal schema. What would you remove, what would you change and what would you leave alone?
3 min answer -
advanced
A transformation project has grown to 600 models with chains twelve deep. A change at the base has an unknowable blast radius. What do you do?
2 min answer -
advanced
A transformation project has grown to hundreds of models with a build taking hours. What structural problems produce that, and what fixes them?
2 min answer -
advanced
Your transformation project has 2,400 models. A full build takes six hours and one change rebuilds half the graph. How did this happen and what do you do?
2 min answer
3 terms in this topic
Model Graph Pruning
Periodically removing transformation models with no downstream consumers - the cheapest available reduction in a build's cost and duration, and the o…
conceptSQL Transformation Model
A named, versioned, tested SELECT statement that declares its own dependencies, turning transformation logic into reviewable software.
toolTransformation Framework
A tool that turns warehouse transformations into version-controlled, tested, dependency-aware code rather than a collection of scheduled scripts.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.