Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
5 to work through
-
intermediate
A platform team proposes one ingestion standard: every source publishes to Kafka and the warehouse loads only from Kafka. That includes 40 SaaS connectors that pull once a day. The argument is one path, one set of tools, and replayability everywhere. Review it — what would you remove, what would you keep and what would you leave alone?
3 min answer -
intermediate
An analytics team extracts daily using a `modified_at` watermark. Reconciliation shows the warehouse has 3% more customers than the source. Why?
2 min answer -
advanced Multiple choice
A connector extracts a SaaS object by paging with offset and a limit of 500 under a modified-at filter. Every nightly run loads exactly 10000 rows and finishes green. Reconciliation shows 61000 matching records at source. The API returns HTTP 200 with an empty page at offset 10000 and support confirms an undocumented deep-paging cap. Which change actually closes the gap?
3 min answer -
advanced
A data-ingestion platform runs hundreds of connectors, each with different APIs, rate limits, failure modes and schemas. How should isolation, retries, scheduling, checkpointing, backfills and schema evolution be designed?
2 min answer -
advanced
A platform ingests high-volume telemetry from many sources. Which ingestion pattern properties matter most?
2 min answer
3 terms in this topic
Ingestion Pattern
The mechanism by which data leaves a source system, which determines freshness, source load and how faithfully change is captured.
patternMerge Upsert
Applying a batch of changes to a target by matching on a key and inserting, updating or deleting per row, which is expensive and frequently avoidable.
conceptSilent Coercion
An ingestion pipeline converting a value it did not expect rather than rejecting it - the worst available failure mode, because the load succeeds and…
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.