Data Platform Architecture
also called Analytical Platform Architecture, Data Stack Architecture
The arrangement of ingestion, storage, transformation, serving and governance that decides whether analytical data arrives on time, means the same thing to everyone and costs a predictable amount.
A company has a warehouse, a scheduler, a transformation project and 42 dashboards, and it still cannot answer "what was revenue last month" without two people arguing. Nothing is broken. Every job is green. The problem is that the estate was assembled tool by tool, and no layer promises anything to the layer above it.
A data platform architecture is the set of contracts between layers, not the set of tools. Two companies can run identical software and have completely different platforms: in one, a consumer can state when a table will be ready, who owns it, what its grain is and what happens when it is late; in the other, all four answers are "ask the data team".
The tooling decisions that dominate procurement conversations - which warehouse, which orchestrator, which ingestion vendor - are the least consequential part. The consequential parts are where data enters, what is guaranteed at each hop, who owns each dataset, and what the platform refuses to do.
Why it matters
Three costs land on the organisation when the contracts are absent, and none of them look like a technical failure.
Freshness becomes unknowable. Without a published commitment, every consumer builds a private belief about when data is current, and those beliefs diverge. The first symptom is a report run at 08:00 that was correct on Tuesday and wrong on Wednesday because a feed slipped by an hour.
Definitions fork. Three teams each compute "active customer" from the same tables and get three numbers, so the organisation spends its analytical capacity reconciling rather than deciding.
Cost grows faster than usage. Consumption pricing means an unowned dashboard refreshing every 5 minutes runs roughly 100,000 times a year and costs money forever. Without ownership metadata there is nobody to ask whether it is still needed, and a first audit commonly finds that roughly a fifth to two fifths of scheduled refreshes serve nobody.
Implementation patterns
- Publish a contract per dataset: grain, owner, freshness commitment, schema-change policy and a deprecation path. This is the artefact that repays the most per hour spent on it, and it is usually a table in the catalogue rather than a document.
- Layer by guarantee, not by name. Raw is a faithful replayable copy with nothing corrected; the modelled layer is where meaning is applied; the serving layer is shaped for consumption. A layer that makes no distinct promise is a copy with a different prefix.
- One entry path per source class. Event streams through a log, operational databases through change capture, vendor systems through scheduled extracts. Three well-operated paths beat one universal path that every source has to be bent into.
- Make transformation code, not scripts: dependency graph derived from the code, tests in the graph, environment substitution so changes can be tested against production shapes.
- Attribute cost from the start. Tag compute by team and dataset on day one, because retrofitting attribution to a live estate is months of work and the conversation about spend arrives long before the tags do.
- Expose one semantic definition for the handful of metrics that appear in board packs, and make every downstream tool read it rather than reimplement it.
Industry example
Notion's published account is a useful shape for the whole discipline. Its product runs on Postgres sharded into 480 logical shards from 2021, and analytics originally loaded into a managed warehouse through off-the-shelf connectors. By 2024 it had published a data lake built on change data capture into Hudi tables on object storage, because the change stream is dominated by updates rather than inserts and the connector-and-warehouse path could not keep up at an acceptable cost.
The architectural lesson is not the tool list. It is that the ingestion path was chosen from a measured property of the data - the ratio of updates to inserts - rather than from a category preference. That is what distinguishes an architecture from an accumulation.
Failure scenarios
- The platform becomes a queue. Every new dataset requires the central team, so lead time grows with adoption and the team is measured on tickets rather than capability.
- Silent staleness. A feed stops and downstream tables retain yesterday's data, so dashboards show plausible numbers with no error anywhere. Detection comes from a human noticing.
- The bypass estate. Analysts query an operational read replica directly because the platform is slow to serve them, and one afternoon a federated query degrades the production database.
- Definition drift after a reorganisation, where the team that owned a metric no longer exists and nobody notices for two quarters.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Central team owns everything | Consistency, one definition, controlled cost | Lead time grows with demand; the team is the bottleneck |
| Domains own their datasets | Throughput, domain knowledge where it belongs | Needs contracts, a catalogue and real platform investment first |
| Governed self-service | Speed within guard rails | Ongoing cost of maintaining the paved path and the templates |
Decentralising without contracts produces fragmentation rather than autonomy, which is the most common expensive mistake in this area.
When not to use it
A company of 40 people with one product database does not need a data platform; it needs a read replica, a scheduled export and a BI tool. Layering, contracts and a catalogue are overhead paid against future coordination cost, and if there are two consumers who sit near each other, that cost is near zero. The honest triggers for building one are the third team asking for the same data, the first reconciliation argument in a leadership meeting, and the first month where analytical spend is a line someone questions.
Interview question
Q: You are the first data platform engineer at a 200-person company with three analysts, a Postgres product database and a growing pile of spreadsheets. Leadership expects a warehouse in the first quarter. What do you build, and what do you deliberately not build?
What a strong answer covers: picking the two or three decisions that are expensive to reverse (where data lands, how it is keyed, who owns what) and leaving the rest cheap; delivering one genuinely useful dataset with a published contract rather than a full layered estate; declining to build a catalogue, a semantic layer and a governance process before there is anything to govern; and naming the measurement that proves it worked, which is lead time for the next dataset rather than the number of tables loaded.
Quick check
Quiz: Two companies run the same warehouse, orchestrator and transformation tool, and one has an architecture while the other has an accumulation. What is the observable difference? - Whether a consumer can state a dataset's grain, owner, freshness commitment and late-data behaviour without asking a person.
Flashcard: Your data estate has every modern tool and still cannot answer "what was revenue last month" without an argument. What is missing? - Contracts between layers. Tools decide capability; published grain, ownership and freshness commitments decide whether anyone can rely on the output.