Feature Store

Architecture Views

21 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

A feature store for a consumer delivery marketplace — the platform behind the ETA you see when you order dinner, the fraud score on the card, and the order the stores appear in. Read it in seven acts: the boundary first, then the people and what they get to do, then the structure, the data, the runtime, the operations, and finally why it is safe. One rule runs through all of it: a feature is defined once and materialised twice, and the online store is a rebuildable projection rather than a system of record.

Context and scope

What sits inside the boundary, who reads it, and the systems the platform reads but never writes to.

People and journeys

Who the platform is for, and what each of them actually gets to do with it — including the three machines in the cast.
03 Producers and consumers Data scientist 60 across 9 teams Goal — Find the feature that already exists before I build a fourth copy of it — and when I do have to build one, ship it today rather than next sprint. Core journeys Define and ship a feature ~40 / month Assemble a training set Search the catalogue ML engineer 45 models in production Goal — Read the vector my model needs in under 25 ms, and know which values are stale before my accuracy tells me. Core journeys Respond to stale features Onboard a model Read the skew report Accountability Feature group owner 120 groups Goal — Know who depends on my feature before I change it, and hear about a freshness breach from the platform rather than from a model owner. Core journeys Approve a breaking change Answer a freshness alert Data governance 2 people, 1.1M entities Goal — Be certain a restricted column has not quietly become an unclassified feature that somebody logs at 100% sampling. Core journeys Classify a feature group Run an erasure request Platform SRE on call 24×7 Goal — Rebuild the online store before any model notices it was gone, and keep a backfill from taking serving with it. Core journeys Rebuild the online store Drain a backfill Machines in the cast Online model service 320k reads/s peak Goal — Get the vector, or a clear reason why not — never a silent zero dressed up as a real value. Core journeys Read a feature vector Materialisation scheduler 120 groups Goal — Land every batch group before the models wake up, and say so loudly when I cannot. Core journeys Run the 05:00 batch wave Stream processor checkpointed Goal — Stay inside five seconds of the event, or declare myself behind rather than look fresh. Core journeys Aggregate a 30-minute window Who the Feature Store is For, and What They Get To Do Person or role Journey / task External / third party Security / platform Three of the eight actors are machines. Two of them are expected to fail in specific, designed-for ways. v 1.0 · owner Data Platform Architecture · date 2026-09 Actors and Journeys Eight actors, three of them machines, each with a goal in their own voice. HTML page SVG draw.io

Structure

The parts, the layers, and the interfaces the rest of the estate sees.

Data

What is stored, how long it survives, and which stores can be thrown away.

Runtime

What happens when a model reads a vector, when a window closes, and when a training set is assembled.

Operations

How a definition reaches production, what is watched, and how the loop closes.

Assurance

Why it is safe to read, and what is assumed to fail.

Architecture One-Pager

The problem, the shape of the answer, the decisions that carry it, and what a prototype should prove.

A feature exists only as a compiled definition with two materialisations of one computation, and the online store is a rebuildable projection of the offline store — never an independently written system of record.

When someone opens a delivery app at eight in the evening and sees "arriving in 32 minutes", a model produced that number from a few hundred signals: how busy this restaurant has been in the last fifteen minutes, how long this courier's last three pickups took, what rain does to this neighbourhood on a Friday. Nine teams own models that need those signals, and the hard part is not computing them. The hard part is that a signal computed one way in a training notebook and another way in a request path produces a model that scores 0.91 offline and performs like 0.74 in production — and the bug is nearly impossible to find afterwards, because both numbers are individually plausible and the training data no longer exists in the form that produced them. The second problem is organisational: all nine teams need "orders at this store in the last 30 minutes", each builds it differently, and none of them can find the others' version.

Hold every feature as a declarative definition in reviewed source control, and compile it into two execution plans — an offline plan and an online plan — from one artefact, refusing to publish anything that cannot express both. Materialise streaming features from one computation writing to both stores, and batch features into the offline store with the online store materialised down from it. Write two timestamps onto every offline value, event time and ingestion time, so a training set can be assembled as of any instant by joining the value that was readable then. Treat the online store as a derived, latest-value projection with weak durability and absolute correctness, rebuildable from the offline store and a seven-day stream log. Serve a vector inside a 25 ms client budget with per-feature status, so absence is always an explicit value with a reason code rather than a silent zero. Then prove the whole claim continuously: log a sample of served vectors, replay them through the offline path, and treat disagreement as a defect with a version stamp pointing at its cause.

What it is, and what it is not

One definition compiled into two materialisation plansTwo pipelines that are supposed to agree
An online store that is derived and disposableA fast database that is the system of record for features
Absence returned as a value with a reason codeA zero substituted where a value was missing
A training set pinned to snapshots, versions and a lookbackA query over whatever the feature tables hold today
Skew measured continuously as a controlA skew dashboard nobody opens
A catalogue with owners, adoption and lineageA wiki page listing what exists

The decisions that are the architecture

01One definition, compiled into two plans, or it does not publish

The compiler emits an offline plan and an online plan from a single artefact and fails the build if either cannot be expressed. That single gate is the parity guarantee, mechanised rather than reviewed.

ADR-01

02On-demand transformations run in the serving path, from the same artefact

The transformations most likely to skew are exactly the ones that cannot be precomputed. They are compiled from the shared definition and sandboxed inside a 2 ms budget, which buys parity where it is most at risk and accepts consumer code in the serving failure domain.

ADR-02

03Streaming features dual-write; batch features materialise down

Freshness where it is required, provenance where it is affordable. Two write topologies rather than one, with the version stamp and nightly replay as the detector for the partial-failure risk dual-write introduces.

ADR-03

04Skew is a continuous control, not a report

One per cent of served vectors are logged and replayed nightly through the offline path. That number is the only ongoing evidence that the two materialisations still agree, and the platform is exactly as trustworthy as it.

ADR-04

05Two timestamps on every offline value, or the feature is refused

Event time says when the fact became true; ingestion time says when the platform could first have known it. Without both, a point-in-time join is guesswork, and the platform refuses the feature rather than approximating it.

ADR-05

06The online store is derived, latest-value and disposable

The largest and hottest store in the platform has the weakest durability requirement and the strictest correctness requirement. It is rebuildable in 90 minutes, which makes online materialisation a cost lever and regional recovery a rebuild rather than a failover.

ADR-08

07Absence is a value; criticality belongs to the consumer

A missing feature returns a reason code and the declared default that training also used, and only a consumer-declared critical group fails the call. A fraud model may fail closed where a ranking model would rather rank slightly worse.

ADR-09

08The feature version is the unit of versioning and of pinning

A transformation, window or dtype change creates a new immutable version and leaves pinned consumers where they are. Training sets name the exact versions they were built from, which is what makes a regenerated set bit-identical.

ADR-11

09Definitions arrive only through reviewed source control

A feature that can be created by an API call is a feature with no review, no version history and no duplicate check. Registration is a merge, and near-duplicates are blocked there rather than discovered a year later.

ADR-12

10Backfill is isolated, budgeted and pre-emptible

Rewriting history is normal in this platform and must never be able to take serving with it. The cost is estimated before the run and the pool yields to live materialisation.

ADR-14

Why this should still be right in ten years

Table formats, stream processors and key-value stores will all be replaced inside the life of this platform. These are the properties that should outlast them.

Compilation outlives the engines

The contract is that one artefact produces both materialisations. Whether the offline engine is Spark, a warehouse, or something that does not exist yet, and whether the online store is a key-value store or a format nobody has shipped, the guarantee is unchanged and every existing definition survives the swap.

Two timestamps are a property of the world, not of a product

The gap between when a fact became true and when a system could know it is created by physics and network partitions, not by a vendor. Any platform that wants reproducible training data will need both columns, whatever it is built on.

Derived state stays cheap to be wrong about

As long as the online store is a projection, its technology is an implementation detail and its loss is an inconvenience. The decision that will age worst in most feature stores — betting correctness on a fast store — is the one this design refuses to make.

Explicit absence survives every model generation

Reason codes and declared defaults are a contract between platform and consumer that does not care whether the consumer is a gradient-boosted tree, a transformer, or whatever replaces them.

Skew replay is a measurement, not a tool

Replaying served inputs through the batch path to check agreement is a technique, not a product. It will still be the cheapest available proof of parity in ten years.

The organisational claim is the durable one

The platform's real product is that nine teams share one definition of a signal. That value grows with the number of teams and models, and it is the last thing a replacement platform would be allowed to give up.

Non-functional targets

Every figure below is a stated assumption for this design, chosen to be defensible and arguable rather than measured. A reviewer who changes one can follow it to the decision that depends on it.

QualityTargetHow it is metView
Serving availability ≥ 99.95% monthly for the online path Three AZs, stateless serving pods, a derived store that can lose a partition without losing truth 15
Online read latency p50 3 ms, p99 15 ms, p99.9 40 ms server-side; p99 25 ms client-observed Locally cached compiled plans, one batch read across groups, short-TTL cache for hot keys 12
Streaming freshness event-to-readable p99 ≤ 5 s; windows ≤ 30 min readable p99 ≤ 30 s after close Checkpointed stream processor dual-writing both stores from one computation 13
Batch freshness hourly groups within 20 min of the hour; daily groups by 06:00 local p95 Dependency-aware scheduling with partition-level retry, separate from the backfill pool 13
Read throughput 320,000 vector reads/second peak, 3× burst for 10 minutes On-demand key-value capacity, per-consumer quotas, request coalescing on hot keys 08
Offline join performance 500M spine rows × 200 features p95 ≤ 45 min, ceiling 90 min Snapshot-isolated as-of reads with partition pruning on entity type, group and event date 14
Point-in-time correctness zero violations tolerated; regenerated sets bit-identical Event and ingestion timestamps on every offline value; snapshot, version and lookback pinned in the manifest 11
Training–serving skew ≤ 0.1% of sampled requests per feature beyond declared tolerance 1% serving log replayed nightly through the offline path, attributed by definition version 18
Staleness detection group marked stale within 60 s of exceeding 2× its freshness SLO; owner alerted ≤ 5 min Freshness SLO per group evaluated continuously, marker surfaced in serving responses 17
Recovery registry RPO 0 / RTO 1 h; offline RPO 15 min / RTO 4 h; full online rebuild ≤ 90 min; serving RTO 15 min cross-region Small authoritative registry, re-derivable offline store, disposable online store rebuilt rather than replicated 10
Time to production merged definition readable online ≤ 30 min for streaming, ≤ 1 cadence for batch Compile-and-publish pipeline with the parity gate as the only blocking check 16
Cost ≤ $0.40 per million feature-vector reads; zero-read online features reported at 30 days Online materialisation flag per feature, cost attributed to producing and consuming teams per group 17

Scope

In scope

  • Declarative feature definitions, a registry, and a compiler emitting an offline and an online plan from one artefact
  • Batch and streaming materialisation, with contract checks that quarantine rather than publish a bad batch
  • An offline store with event and ingestion timestamps, and point-in-time training set assembly with a reproducible manifest
  • A low-latency online serving path with per-feature status, declared defaults and per-consumer quotas
  • On-demand request-time transformations compiled from the shared definition artefact
  • Freshness SLOs, staleness marking, skew replay and drift attribution
  • A catalogue with owners, lineage, adoption evidence and duplicate detection
  • Feature-group authorisation, sensitivity classification and entity-level erasure

Explicitly out of scope

  • Model training, model serving and the model registry — the platform serves them and does not host them
  • Experiment tracking, hyperparameter search and evaluation methodology
  • Labelling and annotation
  • The source systems that emit events, and their schemas
  • Business intelligence reporting over the same underlying data
  • Feature selection — which features a model should use is the data scientist's decision

What a four-week prototype should prove

The prototype's job is to falsify the central claim — that one definition can be compiled into two materialisations that provably agree — on two features and one model, not to build a platform.

  1. Compile one streaming feature and one batch feature from a single definition artefact into both an offline and an online plan, and show the build failing when only one plan can be expressed
  2. Write both stores, then deliberately corrupt the online value for one entity and show nightly skew replay finding it and naming the definition version
  3. Assemble a training set over a 10M-row spine and prove point-in-time correctness by showing that a value whose ingestion timestamp postdates its spine row is excluded
  4. Regenerate the same training set from its manifest a week later and show it bit-identical after new data has landed
  5. Delete the online store for one feature group and rebuild it from the offline store inside ten minutes with serving up throughout
  6. Serve a vector with one group deliberately unavailable and show the reason code, the declared default, and the same default applied in the training path
  • A source event schema changes mid-run: the batch is quarantined, the last good values keep serving, and the owner is alerted rather than the values being published
  • The stream processor falls three minutes behind: the group is marked stale within 60 s and the reading model sees the age, not a fresh-looking number
  • One feature group is made unreachable: the vector returns with a reason code and the declared default, and a consumer that declared the group critical fails its call instead
  • A definition is edited to change a window: a new version is published, the pinned consumer stays on the old one, and both versions materialise side by side
  • A thirteen-month backfill runs during the morning batch wave: it is pre-empted, restarts at partition grain, and no scheduled group misses its landing time

Open risks, carried rather than hidden

RiskIf it landsResponse
The parity gate is slower than writing a feature by hand, so teams route around it Features get computed in model services again, the central claim quietly stops being true, and the platform becomes a catalogue of what some teams did Treat time-to-production as a first-class NFR at 30 minutes and measure it per team; keep the parity gate the only blocking check in the pipeline; make the catalogue and backfill estimate good enough that using the platform is the fast path rather than the compliant one
The online store rebuild path is documented but never exercised The 90-minute rebuild and the 15-minute cross-region RTO are both fiction, and the decision to treat the online store as disposable becomes the platform's largest single point of failure Rebuild one production feature group on a schedule as a routine operation rather than a drill, and publish the measured time alongside the SLO so a regression is visible
On-demand transformations become a general-purpose compute surface in the serving path Consumer-authored code grows past its 2 ms budget, and a bad transformation takes the serving tier down for every model, not just its own Sandbox with a hard CPU and wall-clock ceiling enforced per call, cap the transformation surface to pure functions over request payload and fetched values, and hold on-demand features to phase 2 so the MVP proves parity before it accepts the risk
Group-granularity authorisation is too coarse for restricted features Teams split groups to model access control, the group stops being a materialisation unit, and storage layout follows the permissions model Keep sensitivity classification at the group and enforce it at registration; where a feature genuinely needs narrower access, accept the new group as the answer and price it, rather than adding per-feature ACLs to the read path
Nightly skew replay is too slow for a regression that appears at peak A bad definition version serves wrong values for up to a day before the control that exists to catch it reports anything Keep the per-consumer health view as the live signal for freshness and absence, raise the sampling rate per model on demand to 100%, and treat replay as the proof of parity rather than the alarm for it

Feature Store — Architecture Decision Record

Why every component and every technology on these 21 views is what it is, and what each choice costs.

Seventeen decisions make up this architecture. Everything else across the twenty-one views is either a consequence of one of them or a detail that could be decided differently next quarter without anybody having to redraw the set. Each record carries more than the classic context / decision / consequences triple: the forcing question, how the decision is realised on AWS, the alternatives including the ones that are right for a different organisation, the conditions that would flip the choice, why it should still hold as scale and technology change, and the transferable lesson.

Status of this document. This is a design, not a report on a running system. Every rate, latency, threshold, tolerance and retention figure is a stated assumption chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a consumer delivery marketplace: 40M monthly active users, 1.1M partner stores, 180k couriers, 2.2M orders a day peaking at 4,500 a minute, 45 models in production owned by 9 teams, and 1,400 features in 120 groups across 6 entity types. The named AWS services are the ones this exercise commits to; the requirement in ask.md stays vendor-neutral so it survives a change of cloud. A reviewer who disagrees with a number can follow it to the record that depends on it.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

Parity — one definition, two materialisations 4

The decisions that make the value a model trains on and the value it is served the same number.

ADR-01One definition artefact, compiled into both plans, or it does not publish ADR-02On-demand transformations run in the serving path, compiled from the shared artefact ADR-03Streaming features dual-write from one computation; batch features materialise down from offline ADR-04Skew is proven continuously by replaying sampled served vectors through the offline path

Correctness in time 3

Timestamps, as-of joins and late data — how the platform avoids training on the future.

ADR-05Every offline value carries event time and ingestion time; a feature without a reliable ingestion time is refused ADR-06Point-in-time correctness lives in the table format's snapshots; lookback and tolerance live in the join engine ADR-07Late-arriving events are accepted and change future training sets, but never restate a generated one

Serving inside a request budget 3

What the online path holds, what it does when a value is missing, and how it stays inside 25 ms.

ADR-08The online store holds latest value only, and is a derived, rebuildable projection ADR-09Absence is a value with a reason code; criticality and absence policy belong to the consumer ADR-10One shared online keyspace now; per-consumer projections only for the heaviest fan-out

Reuse, versioning and accountability 3

The decisions that decide whether nine teams share one definition or build nine.

ADR-11The feature version is the unit of versioning, of pinning and of training-set provenance ADR-12Definitions arrive only through reviewed source control, and near-duplicates are blocked at registration ADR-13Every group declares an owner and a freshness SLO, and staleness is marked rather than hidden

Durability, cost and recovery 2

What can be thrown away, what must not be, and what a backfill is allowed to disturb.

ADR-14Backfill runs in an isolated, budgeted, pre-emptible pool and writes online only where a consumer reads online ADR-15Region loss is met by rebuilding the online store in a warm second region, not by replicating serving state

Access and data protection 2

Who may read a feature, on what basis, and what happens to the evidence trail.

ADR-16Sensitivity classification is enforced at registration, and the serving log inherits the highest class it contains ADR-17Workload identity for every caller, authorisation at feature-group granularity, no shared keys

Technology by capability

Amazon Web Services was chosen for this exercise deliberately rather than because the topic demands it. A feature store is not an AWS-shaped problem — the requirement in ask.md stays vendor-neutral throughout, and every decision above would hold on another cloud. AWS was picked because it is the least-used major cloud across this practice's recent use cases and had never carried a data-platform one, and because the offline/online split maps onto its services without a proprietary lakehouse in the middle. The table below is what this exercise commits to, what each choice would be replaced by elsewhere, and which decision record it serves.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Offline store S3 with Apache Iceberg tables and the Glue Data Catalog AWS + open format Delta Lake on ADLS Gen2, or BigQuery with time travel Snapshot isolation lets a 45-minute training join read a stable view while materialisation commits, and the snapshot id becomes a pinnable input to the manifest ADR-06
Online store DynamoDB, one item per entity and feature group, on-demand capacity AWS Cosmos DB, Bigtable, or self-hosted Cassandra / Redis Single-digit-millisecond reads at 320k/s with no capacity planning for the stated 3× burst, and cheap enough to treat as disposable ADR-08
Stream log Kinesis Data Streams with 7-day retention AWS MSK / Kafka, Event Hubs, or Pub/Sub Retention is the replay window that every streaming recovery path in view 21 depends on; 7 days is the assumption the RPO rests on ADR-07
Stream processing Managed Service for Apache Flink, checkpointed AWS + open engine Self-hosted Flink, Spark Structured Streaming, or Dataflow Windowed aggregates from 1 minute to 24 hours with exactly-once checkpointing, and a job graph the compiler can emit ADR-03
Batch materialisation EMR Serverless, two pools AWS + open engine Databricks jobs, Glue, or Dataproc Separate pools give backfill its own capacity and pre-emption without a second cluster to keep warm ADR-14
Registry Aurora PostgreSQL Serverless v2 AWS Azure SQL, Cloud SQL, or self-hosted PostgreSQL Small, strongly consistent and the only store whose loss is unrecoverable; relational because the model in view 11 is relational ADR-11
Serving tier gRPC services on EKS behind an internal NLB, three AZs AWS + Kubernetes AKS, GKE, or ECS Fargate Stateless pods scaled independently of materialisation, with IRSA giving each caller a workload identity rather than a key ADR-17
Hot-key cache ElastiCache with a short TTL, values immutable within their window only AWS Momento, Redis Enterprise, or an in-process cache Coalesces reads on a small number of very popular stores without becoming a second, undetectable staleness source ADR-10
Serving log Kinesis Data Firehose to S3, partitioned by model and hour AWS Event Hubs Capture, or a direct async write to object storage Buffered, cheap, and classification-aware by prefix, which is what lets the log inherit the vector's maximum class ADR-16
Offline access control Lake Formation column and row grants AWS Unity Catalog, Purview, or Ranger Scoped credentials for a human's offline read, so purpose and ACL are enforced before the data leaves the table ADR-17
Definition pipeline GitHub with required review and GitHub Actions External GitLab CI, Azure DevOps, or CodeBuild Registration is a merge, so review, duplicate detection and the parity gate all happen where the author is already working ADR-12
Observability Amazon Managed Prometheus and Managed Grafana, with CloudWatch for platform metrics AWS + open ecosystem Azure Monitor, Cloud Monitoring, or self-hosted Prometheus Freshness and lag are SLO series rather than logs, and they need to be queryable by feature group at 120-group cardinality ADR-13
Encryption and keys KMS with customer-managed keys for restricted groups AWS Key Vault, Cloud KMS, or Vault A separate key per classification is the boundary that makes a classification more than a tag ADR-16
Cross-region recovery S3 cross-region replication plus a DynamoDB global table replica AWS Paired-region storage replication with a rebuild job The offline replica is what the rebuild reads; the online replica shortens the rebuild rather than guaranteeing correctness ADR-15

The decisions, and the alternatives that lost

Parity — one definition, two materialisationsThe decisions that make the value a model trains on and the value it is served the same number.

ADR-01

One definition artefact, compiled into both plans, or it does not publish

Accepted

What stops the value a model trains on and the value it is served from being two different numbers computed by two different pieces of code?

Context
Training and serving have opposite shapes. Training wants full history over many entities, read once, in a columnar format; serving wants one entity's latest value, read three hundred thousand times a second. Almost every organisation therefore ends up with two implementations of the same feature — a GROUP BY in a notebook and a counter incremented by a service — and no mechanism that notices when they diverge. The resulting defect is uniquely expensive: the model scores well offline, performs badly in production, both numbers are individually plausible, and the training data no longer exists in the form that produced them. Reviewing the two implementations against each other does not work, because the divergence appears later, in a change to one of them.
Decision
A feature exists only as a declarative definition artefact in reviewed source control. A compiler reads that artefact and emits two execution plans — an offline plan and an online plan — and the build fails if either cannot be expressed. Nothing reaches the registry, and therefore nothing is materialised, without both plans. Materialisation engines execute a compiled plan and never read a definition directly.
How it is realised on AWS
The definition lives in a Git repository; a GitHub Actions job runs the compiler, which validates the declaration against the registered entity types and emits a Flink job graph for the online plan and a Spark plan for the offline plan. Both are hashed into a version record written to Aurora PostgreSQL. EMR Serverless and Managed Flink are handed the compiled plan and its version id; neither has credentials to read the registry's definition tables.
Options weighed
  • ChosenOne artefact compiled into two plans, publish gated on both: Makes divergence impossible to introduce rather than possible to detect
  • RejectedTwo hand-written pipelines with a reconciliation report: Cheapest to start and the defect it is supposed to catch is exactly the one that appears between reports
  • RejectedServe from the offline store directly, one implementation only: Perfect parity and cannot meet a 15 ms p99 on a single-entity read
  • Right elsewhereA registry that catalogues pipelines teams already wrote: Right for an organisation retrofitting governance onto existing features; it documents parity rather than providing it
Consequences
What it buys
  • Training–serving skew becomes a detectable defect with a named cause rather than an unexplained accuracy gap
  • A feature's meaning is reviewable, because it exists in exactly one place and arrives through a pull request
  • Engine choices become replaceable: a new stream processor is a new code generator, not a re-specification of 1,400 features
What it costs
  • The platform becomes a gate on every new signal; a data scientist cannot compute a feature in a notebook and ship it
  • Transformations the declaration cannot express are refused, which will be experienced as the platform being less capable than a notebook
  • The compiler is now the most load-bearing component in the system, and a bug in it is a bug in every feature
Choose differently when
If one team owned every model and every feature, and features were rewritten rather than reused, the compiler would be pure overhead and two hand-written pipelines with a reconciliation job would be cheaper and faster. The decision is justified by the nine teams, not by the technology.
Why it holds up over time
The contract — one artefact, two materialisations, publish gated on both — is independent of every engine that implements it. Spark, Flink, a warehouse, or something that does not exist yet can each satisfy it, and every existing definition survives the substitution. This is the one decision in the set that should never need revisiting.
LessonWhen two systems are supposed to agree, do not build a detector for disagreement; build a generator that cannot produce it. A gate at compile time is worth more than any amount of monitoring downstream of it.
Shown on views02 16 13
ADR-02

On-demand transformations run in the serving path, compiled from the shared artefact

Accepted

Where does a feature live that cannot be precomputed — the distance between this courier and this restaurant right now, the time since this user's last order?

Context
Some transformations depend on values only present in the request. They cannot be materialised, so they are computed somewhere at request time. The tempting answer is that this is the model service's business, and that is exactly where the worst skew historically comes from: a haversine distance implemented in Python in the training notebook and in Go in the serving path, with a different earth radius and a different rounding rule. These transformations are simultaneously the smallest part of a feature set and the largest contributor to skew, because they are the only ones where two implementations genuinely exist.
Decision
On-demand features are first-class platform features. They are declared in the same artefact, compiled by the same compiler, and executed by a sandboxed runtime inside the serving path with a hard budget of 2 ms p99. The identical compiled transformation is replayed by the training-set service when assembling history, so the training and serving values come from the same code.
How it is realised on AWS
The compiler emits the transformation as a restricted expression over a fixed input set — request payload fields plus already-fetched feature values — with no I/O and no unbounded loops. The serving API evaluates it in-process with a wall-clock ceiling enforced per call and a circuit breaker per feature. EMR Spark evaluates the same expression as a UDF during point-in-time assembly, reading the request payload from the spine.
Options weighed
  • ChosenPlatform-owned runtime, shared compiled artefact: Buys parity exactly where skew is most likely, at the cost of consumer code in the serving failure domain
  • RejectedLeave request-time transformations to the model service: Keeps the platform simple and leaves the largest source of skew entirely unaddressed
  • RejectedPrecompute over a discretised grid and look up: Works for distance and nothing else; a grid fine enough to be accurate is a table too large to materialise
  • DeferredA sidecar the model service calls, owned by the consumer: Defensible isolation model; revisit if a transformation genuinely needs more than 2 ms
Consequences
What it buys
  • The transformations most likely to skew are the ones where parity is strongest
  • A model service stops carrying feature logic, so its deploy cadence stops being a feature release mechanism
  • The point-in-time training set includes request-time features rather than approximating them
What it costs
  • Consumer-authored code runs in the serving path's latency budget and failure domain
  • The expression language is deliberately restricted, so some transformations will be refused as on-demand features
  • The 2 ms budget is a real constraint that will push some logic back to the model service anyway
Choose differently when
If the on-demand surface were more than a small fraction of the feature set, or if the transformations needed I/O, the sandbox would become a general-purpose compute platform inside the serving path and the sidecar option would be right. Held to pure functions over a fixed input set, it stays a compiler feature rather than a runtime.
Why it holds up over time
The principle — the transformation most at risk of skew is the one that must share its implementation — survives any change in how the expression is evaluated. The specific restriction to pure functions may loosen as isolation technology improves, without changing the guarantee.
LessonPut the platform's strongest guarantee where the risk actually is, not where it is cheapest to provide. The features that are hardest to own are usually the ones worth owning.
Shown on views12 13
ADR-03

Streaming features dual-write from one computation; batch features materialise down from offline

Accepted

Is the online store written beside the offline store by the same job, or derived from the offline store after the fact?

Context
Two topologies are available and each fails differently. Dual-write — one computation, two sinks — gives the freshest online value and one code path, but has two sinks that can partially fail, leaving the paths silently divergent with no single place that noticed. Materialise-down — write offline, then project into the online store — makes the offline store unambiguously the source of truth and the online store trivially rebuildable, at the cost of added latency. The mistake is choosing one answer for the whole platform: a five-second freshness SLO on a kitchen-load feature cannot pay for a round trip through an Iceberg commit, and a nightly store-attributes feature has no reason not to.
Decision
Both, chosen per modality rather than per platform. Streaming features dual-write from the same Flink computation to the offline and online stores. Batch features are written to the offline store only, and the online store is materialised down from it as a separate step. Every written value carries its definition version, and nightly skew replay is the detector for the divergence that dual-write makes possible.
How it is realised on AWS
The Flink job writes an Iceberg append and a DynamoDB put from the same operator, with the offline write first so a crash between them loses the online value rather than the record. EMR Serverless writes batch groups to Iceberg; a following job reads the committed snapshot and writes the latest value per entity into DynamoDB. The materialise-down job is idempotent on snapshot id, so a retry is free.
Options weighed
  • ChosenDual-write for streaming, materialise-down for batch: Freshness where it is required, provenance where it is affordable
  • RejectedMaterialise-down for everything: One topology, unambiguous truth, and cannot meet a 5 s streaming freshness SLO through a table commit
  • RejectedDual-write for everything: Uniform and fresh, and makes every batch group's online values as hard to reason about as the streaming ones
  • RejectedOnline store as a cache populated by read-through from offline: Simple and turns every cold read into a table scan at p99
Consequences
What it buys
  • Each modality gets the topology its freshness SLO can actually pay for
  • Batch features — the majority — have a single unambiguous source of truth and a free rebuild
  • The materialise-down job's idempotency makes the online store's rebuild path the same code as its normal path
What it costs
  • Two write topologies to operate, reason about and document
  • Streaming features carry a real partial-failure risk, detected after the fact rather than prevented
  • The offline-write-first ordering means a streaming value can be in history before it is servable, which looks like staleness
Choose differently when
If the streaming freshness SLO relaxed to a minute, materialise-down for everything would be the right answer and this record would collapse into one topology. Conversely, if a table format offered second-scale commits with snapshot isolation at this write rate, the same simplification would follow.
Why it holds up over time
The reasoning — pick the write topology from the freshness requirement, not from a preference for uniformity — outlives both topologies. What will change is where the boundary sits, as commit latencies fall.
LessonResist making one choice for the whole platform when the workloads have different requirements. Two topologies with a stated rule for which applies is simpler in practice than one topology that someone has to work around.
Shown on views13 09 07
ADR-04

Skew is proven continuously by replaying sampled served vectors through the offline path

Accepted

How does the platform know, today, that its two materialisations still agree?

Context
Every guarantee in ADR-01 and ADR-03 is a claim about code that was correct when it was written. Engines get upgraded, a Flink window's allowed lateness gets tuned, a batch job's partition predicate gets edited, and none of those changes announce that they moved a value. A platform whose central promise is parity and which has no continuous measurement of parity is asking to be trusted rather than earning it. The measurement has to come from production, because the failure is in the difference between production and the training path, not inside either.
Decision
One per cent of served feature vectors are logged with their request id, entity keys and timestamps, configurable to 100% per model. Nightly, those logged vectors are replayed through the offline path as an as-of read at the request's own timestamp, and each feature's served value is compared with the offline value. Disagreement beyond the feature's declared tolerance is a defect attributed to a definition version, and the aggregate is the platform's published skew number.
How it is realised on AWS
The serving API writes sampled vectors through Kinesis Data Firehose to S3, partitioned by model and hour. An EMR Serverless job reads the partition, joins each row against the Iceberg offline store as of the logged timestamp, and writes a skew_measurement row per feature and consumer model to Aurora. The version stamp on both the served value and the offline value is what turns a disagreement into a cause.
Options weighed
  • ChosenSampled serving log replayed nightly through the offline path: The only proof that comes from production; cheap at 1% and exact at 100%
  • RejectedUnit tests comparing the two plans on synthetic inputs: Necessary and insufficient — it tests the compiler, not the running materialisations
  • RejectedCompare aggregate distributions of served and training values: Catches gross drift and misses a per-entity divergence, which is the failure that matters
  • RejectedSynchronous dual-read on every request, offline and online: Exact and impossible inside a 15 ms p99
Consequences
What it buys
  • The platform's central claim carries a number, published per feature and per consumer
  • A regression is dated and attributed to a definition version rather than argued about
  • The serving log doubles as training data with exactly the inputs production used
What it costs
  • Nightly cadence dates a regression to a day, not an hour, so this is a proof and not an alarm
  • The serving log is a copy of feature values and inherits their sensitivity classification
  • A 1% sample will not find a divergence confined to rare entities, which is the case sampling is worst at
Choose differently when
If serving traffic were small enough to log entirely and replay hourly, replay would become the alarm as well as the proof and the per-consumer health view would matter less. At 320,000 reads a second, 1% is already 3.2 million vectors a second sampled down to a manageable dataset.
Why it holds up over time
Replaying production inputs through the batch path to test agreement is a technique, not a product, and it will still be the cheapest available proof of parity a decade from now. The sampling rate and the cadence will change with the cost of compute.
LessonA guarantee without a continuous measurement is a hope. Decide early what number would falsify your central claim, and then pay to produce it every night.
Shown on views18 17 09

Correctness in timeTimestamps, as-of joins and late data — how the platform avoids training on the future.

ADR-05

Every offline value carries event time and ingestion time; a feature without a reliable ingestion time is refused

Accepted

When a training row says 19:42, which version of "orders in the last 15 minutes" should it see?

Context
The answer that feels obvious — the value whose event window ended at 19:42 — is wrong, and wrong in the direction that inflates offline accuracy. At 19:42 in production, the platform did not yet know the 19:42 value: the event was still in flight, the window had not closed, the batch had not landed. A training set built on event time alone therefore trains the model on information the serving path will never have, which is label leakage wearing the costume of a join condition. It is invisible in evaluation, because the evaluation set has the same leak.
Decision
Every value written to the offline store carries two timestamps: event time, when the fact became true in the world, and ingestion time, when the platform could first have read it. Point-in-time joins select on ingestion time. A feature whose source cannot supply a reliable ingestion time is refused registration for offline materialisation rather than approximated.
How it is realised on AWS
Iceberg tables carry event_ts and ingest_ts columns, with event date as a partition key for pruning and ingest_ts written by the materialising engine at commit time rather than derived from the payload. The as-of join predicate is ingest_ts <= spine_ts. The compiler refuses a definition whose source declaration does not establish how ingestion time is obtained.
Options weighed
  • ChosenBoth timestamps, join on ingestion time, refuse features that cannot supply it: The only join that reproduces what production could actually read
  • RejectedEvent time only: Simplest schema and systematically leaks future information into training
  • RejectedEvent time plus a fixed per-feature ingestion lag estimate: Cheap approximation whose error is largest exactly when the pipeline is late, which is when it matters
  • Right elsewhereProcessing time only: Right for a platform whose features are all stream-computed with negligible lag; it cannot express a late-arriving fact
Consequences
What it buys
  • A training row sees what production would have seen, which is the definition of point-in-time correctness
  • Late arrival becomes representable rather than a source of silent error
  • Leakage becomes a detectable condition — ingestion lag exceeding the spine's own horizon — rather than an unexplained accuracy drop
What it costs
  • Storage and write cost for a second timestamp on 2.6 billion values a day
  • Some features are refused, which will be read as the platform being obstructive
  • Joining on ingestion time rather than event time defeats naive event-date partition pruning and makes the join more expensive
Choose differently when
If every source emitted into a single log with negligible and bounded lag, event time and ingestion time would be within milliseconds and one column would do. A marketplace with nightly warehouse dimensions and a third-party weather feed is nowhere near that, and never will be.
Why it holds up over time
The gap between when a fact becomes true and when a system can know it is created by physics and partitions, not by a vendor. Any platform that wants reproducible training data will need both columns, whatever it is built on.
LessonModel the difference between what was true and what was knowable. Systems that collapse the two produce training data that flatters them.
Shown on views11 14 09
ADR-06

Point-in-time correctness lives in the table format's snapshots; lookback and tolerance live in the join engine

Accepted

Which layer is responsible for producing an as-of value — the store or the query?

Context
Two places can own it. A table format with snapshot isolation and time travel makes an as-of read native, cheap to get right and hard to bypass, but it answers "as of a commit" rather than "as of this row's timestamp", and it has no concept of a maximum lookback or a per-feature tolerance. A join engine above a plainer store can express all of that, and is portable across stores, but re-implements correctness on every query path and can be bypassed by anybody with read access to the bucket.
Decision
Split it. The table format owns snapshot isolation and the physical as-of read, so that a training job reads a stable view while materialisation writes, and so a snapshot id is a pinnable input. The join engine owns the semantics the format has no opinion on: the per-row as-of predicate, the declared maximum lookback, reason-coded nulls, and per-feature tolerance. Direct reads of the offline store are permitted for batch scoring but are not point-in-time joins and are not described as such.
How it is realised on AWS
Iceberg on S3 with the Glue Data Catalog provides snapshot isolation and time travel; the training-set service pins a snapshot id per table and records it in the manifest. EMR Spark executes the as-of join with ingest_ts <= spine_ts and a lookback floor of spine_ts minus the feature's declared window, emitting a reason code where no value falls inside it. Lake Formation grants govern the direct read path.
Options weighed
  • ChosenSnapshots in the format, semantics in the engine: Each layer owns what it can actually enforce
  • RejectedEverything in the join engine over plain Parquet: Portable, and gives up snapshot isolation, so a long training job races a materialisation commit
  • RejectedEverything in the format, via as-of-commit reads only: Cheapest and cannot express a per-row timestamp, a lookback or a tolerance
  • DeferredA dedicated point-in-time serving engine in front of the store: Worth revisiting if join cost rather than join correctness becomes the constraint
Consequences
What it buys
  • A training job and a materialisation run can proceed concurrently without a lock or a freeze window
  • The snapshot id becomes a pinnable input, which is what makes a regenerated training set bit-identical
  • Lookback and tolerance are declared per feature and enforced in one place
What it costs
  • Two layers to understand when debugging a wrong training value
  • The direct read path exists and is not point-in-time, which is a foot-gun that has to be documented and named
  • Iceberg metadata growth at 2.6 billion writes a day needs its own compaction and expiry operation
Choose differently when
If the store offered per-row as-of semantics and a lookback predicate natively, the engine's role would shrink to reason codes and this would become one layer. If snapshot isolation were unavailable, the whole design would need a freeze window during training-set generation, which at 45 models would be unworkable.
Why it holds up over time
The split follows a durable rule: put a guarantee in the layer that can enforce it against every access path, and put policy in the layer that can express it. Table formats will keep absorbing semantics upward, which moves the line without changing the rule.
LessonWhen two layers could own a guarantee, ask which one can be bypassed. Physical isolation belongs where it cannot be, and policy belongs where it can be read.
Shown on views14 10 11
ADR-07

Late-arriving events are accepted and change future training sets, but never restate a generated one

Accepted

A courier's position from 19:38 arrives at 19:51, after values from 19:45 are already written. What happens to it, and to the training sets built in between?

Context
Late arrival is the normal case, not the exception: mobile networks partition, the weather feed batches, the warehouse lands at six in the morning. Three responses are available and two are wrong. Dropping the event loses real information and biases the history towards well-connected couriers. Rewriting history in place makes a training set silently different from the one that was evaluated, which destroys the reproducibility everything else in this design is built on. The third response — accept it, timestamp it honestly, and let it affect only what has not yet been produced — is the only one consistent with ADR-05.
Decision
A late event is accepted and written with its true event time and its true ingestion time. Because point-in-time joins select on ingestion time, it becomes visible to any training set generated after it landed and invisible to any generated before. Already-generated training sets are never restated. A generated set's manifest records the snapshot ids that produced it, so its exact content remains derivable.
How it is realised on AWS
Iceberg appends carry the late value into the event-date partition it belongs to, with ingest_ts set at commit. Flink is configured with bounded allowed lateness for window computation and routes later arrivals to the offline store only, so a closed window is not silently recomputed in the online store. No update or delete is issued against a historical value.
Options weighed
  • ChosenAccept, timestamp honestly, affect only future sets: Preserves both the information and the reproducibility
  • RejectedDrop events past the window's allowed lateness: Simplest and biases history against exactly the entities whose connectivity is worst
  • RejectedRecompute affected windows and restate history in place: Most accurate view of the past and makes every generated training set unreproducible
  • DeferredVersioned windows with an explicit restatement event: Right for a platform with regulatory restatement duties; heavier machinery than a marketplace needs
Consequences
What it buys
  • No information is discarded, and no already-published dataset changes underneath a model
  • The mechanism is one rule rather than a special path: ingestion time does all the work
  • A model retrained on the same spine a week later legitimately improves as history fills in
What it costs
  • Two models trained a week apart on the same spine can differ, and that will be read as a bug
  • A closed window in the online store can be less accurate than the same window offline
  • Offline and online values for the same window can differ for legitimate reasons, which complicates skew interpretation
Choose differently when
If the platform owed regulatory restatement — a corrected figure must supersede a published one — versioned windows with explicit restatement events would be required and the no-restatement rule would be unacceptable. A marketplace's training data carries no such duty.
Why it holds up over time
The rule follows from ADR-05 rather than from any product, and it will hold as long as reproducibility is valued above retrospective accuracy. That trade is what makes model development possible at all.
LessonImmutability of what you have already published is usually worth more than accuracy of what you published. Let new information change the next answer, not the last one.
Shown on views14 11 21

Serving inside a request budgetWhat the online path holds, what it does when a value is missing, and how it stays inside 25 ms.

ADR-08

The online store holds latest value only, and is a derived, rebuildable projection

Accepted

Is the online store a database the platform must not lose, or a cache it can throw away?

Context
The online store is the most expensive and most read component in the platform: 45 million entities, 320,000 reads a second, inside a 15 ms p99. Treating it as a system of record is the default, and it is the decision that ages worst. It forces strong durability guarantees onto the hottest path, makes cross-region recovery a state-replication problem, makes every schema change a migration, and — most damagingly — makes the store's contents the answer to "what is this feature's value", which is a question the offline store and the definition should answer together.
Decision
The online store holds latest value plus event timestamp plus definition version, per entity and feature group, and nothing else. It is explicitly derived: fully rebuildable from the offline store and the retained stream log, with a rebuild path exercised on a schedule. Its RPO is best-effort and its correctness requirement is absolute.
How it is realised on AWS
DynamoDB with entity_type:entity_id as the partition key and one item per feature group, holding a value map, event timestamp, version id and a per-feature TTL. On-demand capacity absorbs the stated 3× burst. Rebuild is the materialise-down job of ADR-03 run over all groups, idempotent on snapshot id: all groups in 90 minutes, one group in ten.
Options weighed
  • ChosenDerived latest-value projection, rebuildable: Inverts the requirements on the hottest store: weak durability, absolute correctness
  • RejectedOnline store as a durable system of record with history: Familiar, and makes the most expensive store also the one that must never be lost
  • RejectedIn-memory store with periodic snapshots: Fastest reads and a cold-start problem measured in tens of minutes at 45M entities
  • RejectedRead-through cache in front of the offline store: Minimal storage and turns a cold read into a table scan at exactly the wrong percentile
Consequences
What it buys
  • Regional recovery is a rebuild rather than a state failover, which is why a 15-minute RTO is credible
  • Online materialisation becomes a cost lever: a feature nobody reads online need not be stored online
  • A partition loss is an availability event, never a data-loss event
What it costs
  • The 90-minute rebuild is a hard dependency, and one that rots quietly unless it is run in anger
  • Holding latest value only means the online path cannot answer any historical question, by design
  • During a rebuild the platform is serving from a partially populated store, which must be visible to consumers as staleness
Choose differently when
If the offline store could not reconstruct the online state — features computed from a stream with retention shorter than the rebuild window, or from a source that cannot be replayed — the online store would be the only copy and would have to be treated as one. The seven-day stream log is what buys this decision.
Why it holds up over time
As long as the online store is a projection, its technology is an implementation detail and its loss is an inconvenience. That property is what makes the rest of the platform able to change stores without changing guarantees.
LessonDecide which of your stores you are allowed to lose, and then make sure the expensive one is on that list. Durability spent on derived state is durability taken from somewhere it was needed.
Shown on views10 07 15
ADR-09

Absence is a value with a reason code; criticality and absence policy belong to the consumer

Accepted

A feature group is unreachable at 19:40. Does the request fail, get a default, or get the last known value?

Context
All three are correct answers for different models. A payment-risk model would rather refuse to score than score without its velocity features. A store-ranking model would far rather rank slightly worse than show an empty screen. An ETA model can use a kitchen-load value that is four minutes old and cannot use one that is an hour old. The failure mode to design against is none of these three: it is the silent zero, where a missing value is substituted with a plausible number and the model's output changes with nothing in the system recording that anything happened. That is the single most expensive bug this platform can ship, because it is indistinguishable from the model being wrong.
Decision
Every feature response carries a per-feature status and, where a value is absent, a reason code and the feature's declared default — the same default the training path applies. Whether an absent feature degrades the model or fails the request is declared by the consumer per feature group as a criticality level, recorded as a consumer pin. The platform never substitutes a value without saying so.
How it is realised on AWS
The serving API's gRPC response is a value map plus a parallel status map with codes for OK, STALE, MISSING, QUARANTINED and DENIED, each carrying the value's age. Criticality and absence policy live in the consumer_pin table in Aurora and are resolved into the compiled plan the serving tier caches. A critical group's absence returns an error status for the whole call rather than a partial vector.
Options weighed
  • ChosenExplicit status per feature, criticality declared per consumer: The only option where the model owner chooses their own failure mode
  • RejectedAlways fail the request on any missing feature: Safest for correctness and turns one stale group into a platform-wide outage
  • RejectedAlways substitute the declared default silently: Highest availability and produces wrong answers nobody can attribute
  • RejectedAbsence policy declared on the feature rather than the consumer: Simpler to model and wrong, because the same feature is critical to one model and incidental to another
Consequences
What it buys
  • A model owner chooses their own failure mode and can be held to it
  • Degradation is observable end to end, so an accuracy change has a cause in the logs
  • Applying the same default in training and serving removes a whole class of skew
What it costs
  • The response contract is wider, and every consumer has to handle a status map rather than a value map
  • Per-consumer policy has to be resolved into the serving path without the serving path becoming a policy engine
  • A consumer that declares nothing gets the platform default, which may not be the right failure mode for it
Choose differently when
If every consumer shared one criticality profile, the policy would belong on the feature and this would be a simpler design. With a fraud model and a ranking model reading the same group, it does not.
Why it holds up over time
The contract — absence is a value, and the consumer chooses what it means — is independent of model architecture and of transport. It will outlive both the gRPC schema and the model generation reading it.
LessonNever let a system invent a value silently. An explicit null with a reason is worth more than a plausible number, because only one of the two can be debugged.
Shown on views12 21 05
ADR-10

One shared online keyspace now; per-consumer projections only for the heaviest fan-out

Accepted

Does every model read the same multi-tenant online store, or does each get the exact vector it needs under one key?

Context
A shared keyspace maximises reuse, minimises storage and means a new consumer costs nothing. It also means a model reading five feature groups performs five key lookups, and tail latency is the maximum of five rather than one. Per-consumer projections — one item holding precisely the vector one model wants — turn that into a single read and make p99.9 predictable, at the cost of write amplification proportional to the number of consumers, and a store whose physical shape follows the model roster, so every new model is a schema change.
Decision
The MVP serves every consumer from one shared keyspace, batching reads across groups in a single call. Per-consumer projections are deferred to phase 3 and introduced only for consumers whose measured fan-out and tail latency justify them, as an additional materialisation of the same definitions rather than a replacement for the shared store.
How it is realised on AWS
One DynamoDB table keyed on entity_type:entity_id with one item per feature group, read with BatchGetItem across the groups a request needs. Hot keys are handled by request coalescing and a short-TTL ElastiCache entry for values immutable within their window. A projection, when introduced, is a second table written by the same materialise-down job under a different plan.
Options weighed
  • ChosenShared keyspace with batched multi-group reads: Reuse and simplicity first; buy predictability only where measurement says it is needed
  • RejectedPer-consumer projections from the start: Best tail latency and makes the store's shape a function of the model roster on day one
  • RejectedOne item per entity holding every feature: One read per entity and a write amplification that makes any single feature update rewrite everything
  • Right elsewhereProjections as the only store, no shared keyspace: Right for a platform with few, very high-traffic models; it gives up ad-hoc consumption entirely
Consequences
What it buys
  • A new consuming model costs no storage and no new materialisation
  • One definition, one physical representation, so there is one place a value can be wrong
  • Projections remain available as an optimisation with evidence behind it
What it costs
  • Tail latency for a wide-fan-out consumer is the maximum of its group reads, not the mean
  • Hot keys concentrate across all consumers rather than per consumer, so one popular store affects everyone
  • The phase-3 projection work is real and is being deferred rather than avoided
Choose differently when
If one model's fan-out reached ten groups, or if the p99.9 budget tightened below 40 ms, projections would move from phase 3 into the MVP. The decision is a measurement bet, and the measurement is the per-consumer read latency in view 17.
Why it holds up over time
The principle — start with the shared representation and buy a specialised one only with evidence — holds across store technologies. What changes is the fan-out at which the trade flips, which falls as multi-key reads get cheaper.
LessonDefer the optimisation that makes your storage layout depend on your consumer list. That coupling is easy to add later and very hard to remove.
Shown on views12 08 17

Reuse, versioning and accountabilityThe decisions that decide whether nine teams share one definition or build nine.

ADR-11

The feature version is the unit of versioning, of pinning and of training-set provenance

Accepted

When a definition changes, what exactly changed — the feature, the group it sits in, or the repository commit?

Context
Three granularities are available. Per-commit versioning is simplest and reproducible, and couples unrelated changes: one team's edit forces another team's re-validation, and a training set can only say "built at commit abc123", which is true and useless. Per-group versioning matches the materialisation unit and is too coarse — changing one feature's window reversions forty others. Per-feature versioning gives the smallest blast radius and the most honest consumer pin, at the cost of a group whose features are at different versions and therefore whose storage layout is heterogeneous.
Decision
The feature version is the unit. A change to a transformation, window or data type creates a new immutable version; the previous version stays published and consumers pinned to it stay there until they move. Every written value carries its version id, and a training set's manifest records the exact set of versions it was built from.
How it is realised on AWS
feature_version rows in Aurora hold the transform hash, window specification, lifecycle state and publication time. Both value tables carry version_id. The training-set manifest records the version id per feature alongside the Iceberg snapshot id, which together make a regeneration bit-identical. The registry refuses to retire a version with a live production pin.
Options weighed
  • ChosenPer-feature immutable versions with consumer pins: Smallest blast radius, honest provenance, heterogeneous group layout as the price
  • RejectedPer-commit versioning of the definition repository: Trivial to implement and couples every team's changes to every other team's validation
  • RejectedPer-feature-group versioning: Matches materialisation and forces forty unrelated features to reversion for one window change
  • RejectedMutable definitions with an audit log of edits: Simplest for authors and makes every historical training set unreproducible
Consequences
What it buys
  • A change affects only its own consumers, so a breaking edit is a conversation with three teams rather than nine
  • A training set can name exactly what produced it, which is what ADR-01's guarantee needs to be checkable
  • The version stamp on every value is what makes the skew replay in ADR-04 able to attribute a cause
What it costs
  • A feature group can hold features at several versions, so its physical layout is not uniform
  • Multiple live versions of a feature multiply materialisation variants and therefore storage and compute
  • Consumers can stay pinned to an old version indefinitely, which needs a deprecation process with teeth
Choose differently when
If backfill were free, versions could be collapsed by re-deriving history whenever a definition changed, and pinning would be unnecessary. Backfill costs six hours per feature per thirteen months, which is precisely why the old version survives.
Why it holds up over time
Immutable versions with explicit consumer pins is the same pattern that governs published APIs and library releases. It has outlasted several generations of tooling and will outlast this one.
LessonVersion at the granularity at which people actually consume. Coarser versioning makes unrelated teams each other's problem; finer versioning makes storage heterogeneous, which is the cheaper of the two.
Shown on views11 16 18
ADR-12

Definitions arrive only through reviewed source control, and near-duplicates are blocked at registration

Accepted

How does a feature get created, and what stops the fourth copy of "orders in the last 30 minutes"?

Context
The platform's organisational value is that nine teams share one definition of a signal. That value is destroyed quietly: a team cannot find the existing feature, or finds it and cannot tell whether its window matches, so they write their own. Six months later there are four features with the same name-ish, three owners, and no way to know which the ETA model is actually reading. A catalogue alone does not prevent this, because searching is optional and shipping is not. Nor does a creation API, which produces features with no review, no history and no duplicate check.
Decision
A feature is created by merging a definition into the repository, never by an API call. At registration the platform compares the definition against every existing one on entity, source and transformation equivalence, and a near-duplicate must be adopted, differentiated, or overridden with a recorded reason. Every group carries a named owning team resolved from the organisation directory, and an unowned group is not materialised.
How it is realised on AWS
GitHub with required review; a GitHub Actions job runs the compiler, the duplicate detector and the contract lint, and publishes to the registry on merge. The detector normalises the compiled offline plan and compares transformation hashes plus source and entity, surfacing candidates in the pull request. Ownership resolves against a nightly snapshot of the org directory, and a group whose owner no longer exists is flagged rather than silently orphaned.
Options weighed
  • ChosenSource control only, duplicates blocked at registration: Review and duplicate detection at the one moment the author is paying attention
  • RejectedCreation API with a catalogue for discovery: Fastest authoring and no review, no history, no duplicate check
  • RejectedCatalogue plus periodic duplicate reports: Finds duplicates after both are in production, when neither can be removed
  • Right elsewhereCentral team authors every feature: Right for a small platform with few features; at 40 new features a month it is a queue, not a service
Consequences
What it buys
  • Reuse is enforced at the only moment it is cheap — before the second implementation exists
  • Every feature has a review, a history and an author, so its meaning can be reconstructed
  • Ownership is a precondition for materialisation, which is what makes a freshness alert reach a human
What it costs
  • Authoring is slower than writing a query, and that gap is the platform's main adoption risk
  • Transformation equivalence is undecidable in general, so the detector will produce false positives and be overridden
  • Organisation directory drift will orphan groups after every reorganisation
Choose differently when
If features were disposable and rarely reused — a research platform where every model brings its own signals — review and duplicate detection would be friction with no return, and a creation API would be right. The nine teams are what make it pay.
Why it holds up over time
Definition-as-code with review has outlasted every generation of data tooling for the same reason it outlasted hand-configured servers: the artefact is diffable, attributable and replayable. Nothing about that depends on the current platform.
LessonPut the guardrail at the moment of creation, not in a report. A duplicate found at review is a conversation; a duplicate found in production is a migration nobody will fund.
Shown on views16 03 04
ADR-13

Every group declares an owner and a freshness SLO, and staleness is marked rather than hidden

Accepted

When a feature group stops being fresh, who finds out and what does a reader see?

Context
Staleness is this platform's characteristic failure, and it is dangerous precisely because it is not an error. The read succeeds, the value is well-formed and plausible, and the only thing wrong with it is that it describes a world four hours old. A model consuming it produces confidently wrong output, and the first signal anybody gets is a business metric moving. Detecting staleness requires knowing what fresh means for that group, which means somebody has to declare it — and acting on staleness requires a human who is accountable for that group, which means somebody has to own it.
Decision
Every feature group declares a freshness SLO and a named owning team at registration. The platform evaluates the SLO continuously; a group whose newest value exceeds twice its SLO is marked stale within sixty seconds, the marker is returned to every reader alongside the value's age, and the owner is alerted within five minutes. An unowned group is not materialised.
How it is realised on AWS
Freshness SLO is a field on feature_group in Aurora, evaluated by a controller against the newest event timestamp per group and published as an Amazon Managed Prometheus series. The stale marker is resolved into the serving path's cached plan and returned in the per-feature status of ADR-09. Alert routing resolves the owning team from the org directory snapshot.
Options weighed
  • ChosenDeclared SLO per group, staleness marked in the response, owner alerted: Makes an invisible failure visible to both the reader and the accountable human
  • RejectedPlatform-wide freshness threshold: One number for a five-second stream aggregate and a nightly store attribute is no number at all
  • RejectedAlert the owner but serve the value unmarked: The human finds out and the model does not, which is the wrong way round
  • RejectedRefuse to serve a stale value: Correct in principle and converts a degradation into an outage; ADR-09 gives the consumer that choice instead
Consequences
What it buys
  • The platform's most dangerous failure becomes observable to the code reading it, not just to a dashboard
  • Alerts reach a team that can act, because ownership is a precondition rather than an aspiration
  • Per-group SLOs let a nightly feature and a five-second feature coexist without one of them being permanently in breach
What it costs
  • Every group needs a declared SLO at registration, which is one more thing an author must think about
  • A group can be legitimately stale — a store that stopped trading — so the marker needs a TTL story to avoid permanent noise
  • Org directory drift means the alert can route to a team that no longer exists
Choose differently when
If every feature were stream-computed with the same freshness expectation, a single platform threshold would be sufficient and the per-group declaration would be ceremony. A mix of 1-minute windows and nightly warehouse dimensions makes it necessary.
Why it holds up over time
Declared SLOs with per-reader visibility is the same pattern that made service reliability measurable. It transfers to data with no modification and is unaffected by the technology underneath.
LessonA failure that does not look like a failure needs a declared expectation before it can be detected at all. Ask for the number at creation time, when the author still knows it.
Shown on views17 21 05

Durability, cost and recoveryWhat can be thrown away, what must not be, and what a backfill is allowed to disturb.

ADR-14

Backfill runs in an isolated, budgeted, pre-emptible pool and writes online only where a consumer reads online

Accepted

A new feature needs thirteen months of history. What is that allowed to disturb, and what is it allowed to cost?

Context
Backfill is not an exceptional operation in a feature store; it is how every new feature and every changed definition becomes usable. It is also the largest single compute and write event the platform ever performs, and it is naturally unbounded: the author wants all of history, at full grain, now. Run on the same capacity as scheduled materialisation, it delays the morning batch wave; run without a cost estimate, it produces an invoice nobody approved; and written naively to the online store, it rewrites 45 million entities for a feature no model reads online yet.
Decision
Backfill runs in a resource pool separate from scheduled materialisation, pre-emptible by live jobs, over a declared and bounded window. Its cost is estimated and shown before execution and is gated against a budget at review time. It writes to the online store only for features a consumer actually reads online; offline-only features are backfilled offline only.
How it is realised on AWS
A second EMR Serverless application with its own capacity limits, submitted by the backfill runner with a declared partition range. The estimate is derived from partition counts and historical DPU-per-partition and surfaced in the pull request by the same job that runs the compiler. The online write is conditional on the feature's online flag and on a non-zero consumer pin count.
Options weighed
  • ChosenIsolated, budgeted, pre-emptible pool; online write only where read: Treats a routine operation as a first-class one with its own capacity and its own price
  • RejectedBackfill on the scheduled materialisation capacity: No new infrastructure and one backfill delays every group's morning landing
  • RejectedBackfill only during a declared maintenance window: Protects serving and turns a 30-minute time-to-production into a next-week promise
  • RejectedBackfill on demand at read time, lazily: No upfront cost and makes the first training set over a new feature unboundedly slow
Consequences
What it buys
  • A backfill can never take a dinner-peak read or a morning batch wave with it
  • Cost is a review comment rather than an invoice, which changes the conversation before the spend
  • Skipping the online write for offline-only features removes the largest avoidable write in the platform
What it costs
  • A second compute pool to size, monitor and pay for even when idle
  • Pre-emption means a long backfill can be interrupted and must be restartable at partition grain
  • The cost estimate is a model and will be wrong, most likely on the features with the most skewed partitions
Choose differently when
If backfills were rare — a stable feature set with few definition changes — a maintenance window on shared capacity would be the right answer and the second pool pure cost. At forty new features a month, backfill is a daily operation.
Why it holds up over time
Isolating a bursty, low-priority, high-volume workload from a latency-critical one is a structural decision that survives any change of engine. What will change is the price of the isolated capacity, which is falling.
LessonThe operation you treat as exceptional is the one that will take production down. If it happens weekly, give it its own capacity, its own budget and its own restart semantics.
Shown on views13 15 04
ADR-15

Region loss is met by rebuilding the online store in a warm second region, not by replicating serving state

Accepted

The serving region is gone. Where does the next feature vector come from, and how long does it take?

Context
The obvious answer is to replicate the online store and fail over, and for a system of record that would be right. Here it is expensive for the wrong reason: the online store is the largest and hottest component, so replicating it continuously doubles the platform's dominant cost to protect data that ADR-08 already established is derivable. The alternative extreme — a cold second region — cannot meet any credible RTO, because bringing up a serving tier and populating 45 million entities from scratch is measured in hours.
Decision
The second region runs warm at ten per cent of serving capacity, with the offline store replicated cross-region and the online store present as a replica that is not relied upon for correctness. On regional loss, the serving tier scales out and the online store is rebuilt or completed from the replicated offline state, targeting a fifteen-minute RTO with an explicit latency-SLO breach during the scale-out.
How it is realised on AWS
S3 cross-region replication carries the Iceberg data and metadata; a DynamoDB global table gives the second region a warm replica that shortens the rebuild rather than guaranteeing it. EKS node groups in the second region run at ten per cent with a scaling policy triggered by the failover event, and the materialise-down job of ADR-03 runs against the replicated snapshot to complete any gaps.
Options weighed
  • ChosenWarm second region, online store rebuilt from replicated offline state: Buys a credible RTO without doubling the most expensive tier
  • RejectedActive-active with fully replicated online state: Best RTO and roughly doubles serving and storage cost for state that is derivable
  • RejectedCold standby, infrastructure as code only: Cheapest and cannot populate 45M entities inside any RTO worth stating
  • Right elsewhereAccept single-region availability and degrade models to their defaults: Right where models are advisory; an ETA and a fraud score on the checkout path are not
Consequences
What it buys
  • A 15-minute RTO on the read path without paying continuously for a second full-size online store
  • The recovery path is the platform's normal materialise-down job, so it is exercised by ordinary operation
  • The decision follows directly from ADR-08, so there is one reason the online store is cheap to lose rather than two
What it costs
  • The first minutes after failover are served by a tier that is still scaling, so the latency SLO is breached before it is met
  • Warm-at-ten-per-cent is a permanent cost for capacity that is almost never used
  • The RTO depends on a rebuild whose measured duration must be tracked, not assumed
Choose differently when
If regulation required continuous multi-region availability with no degradation, active-active would be mandatory and the cost argument irrelevant. If the models were advisory rather than on the checkout path, cold standby would be defensible.
Why it holds up over time
Recovering derived state by recomputing it rather than replicating it gets cheaper every year as compute costs fall relative to storage and egress. The decision's economics improve with time rather than eroding.
LessonBefore paying to replicate state, ask whether you could recompute it inside the same recovery objective. For derived data the answer is usually yes, and usually cheaper.
Shown on views15 10 21

Access and data protectionWho may read a feature, on what basis, and what happens to the evidence trail.

ADR-16

Sensitivity classification is enforced at registration, and the serving log inherits the highest class it contains

Accepted

What stops a restricted column becoming an ordinary feature that somebody logs at full sampling?

Context
A feature store is a machine for moving data from restricted places into convenient ones. A transformation over a table containing a card fingerprint or a home address produces a number, and a number does not look sensitive. Once that number is in a general-purpose group, it is readable by any consumer authorised for the group, materialised into a store with weaker controls, and copied into the serving log — which, at the default sampling rate, becomes an unclassified partial replica of restricted data. Checking this at read time is too late, because by then every copy already exists.
Decision
Every feature group carries a sensitivity classification declared at registration. The compiler refuses a definition that derives from a source column of higher classification than its target group. Features derived from personal data additionally declare a lawful basis and purpose, and a consumer must be authorised for the purpose and not merely for the group. The serving log inherits the classification of the most restricted group in the vector it sampled.
How it is realised on AWS
The compiler resolves source column classifications from the Glue Data Catalog and Lake Formation tags and fails the build on escalation. Restricted groups are encrypted with a customer-managed KMS key; Firehose writes the serving log to a prefix whose classification is set from the vector's maximum class, with its own KMS key and its own grants. Entity erasure removes online values within 24 hours and offline values within 30 days.
Options weighed
  • ChosenClassification enforced at registration; log inherits maximum class: Prevents the copy rather than governing it afterwards
  • RejectedClassification checked at read time: Cheap to implement and every copy already exists by the time it fires
  • RejectedManual review of definitions by data governance: Correct in principle; two people cannot review forty definitions a month against 1.1M entities
  • Right elsewhereForbid features derived from restricted columns entirely: Right in a regulated setting where the model may not see the data at all; here it removes fraud detection
Consequences
What it buys
  • Classification escalation becomes impossible to introduce rather than possible to discover
  • The serving log stops being the platform's quiet privacy hole
  • Erasure has a defined path through both stores and is reflected in future point-in-time joins
What it costs
  • Source classification must be accurate in the catalogue, so the platform inherits someone else's metadata quality
  • Restricted groups need their own keys and their own grants, which multiplies operational surface
  • Offline erasure at 30 days means a join run on day 29 can still return an erased entity's values
Choose differently when
If every source were uniformly classified, the compiler check would be ceremony. If the platform had to prove erasure within 24 hours across both stores, offline erasure would need to become a rewrite-on-demand operation rather than a scheduled compaction.
Why it holds up over time
Enforcing a data-protection property at the point of creation rather than the point of access is the pattern that survives every change of regulation, because it removes the copies that the regulation is about.
LessonGovernance applied after the copy exists is documentation. Applied at the moment the copy would be created, it is a control.
Shown on views19 11 20
ADR-17

Workload identity for every caller, authorisation at feature-group granularity, no shared keys

Accepted

How does the platform know which model is asking, and what is it allowed to see?

Context
A serving path taking 320,000 requests a second from four model services is the natural home of a shared secret: one API key per consumer, checked cheaply. That key is then in a config map, in a CI variable, in somebody's laptop history, and rotating it means coordinating four deployments. Worse, a shared key authenticates a deployment rather than a workload, so a compromised sidecar in an unrelated service can read every feature the model could. Meanwhile a human reading the offline store for training has an entirely different problem: they are authorised for a purpose, not just for a resource.
Decision
Every service-to-service caller authenticates as a workload identity, with no shared API keys anywhere in the platform. Authorisation is evaluated per feature group. A human's offline read is authorised on group ACL plus declared purpose, and the credential issued to them is scoped to the columns that survive both checks. A denied group is refused explicitly and by name, never silently dropped from the returned vector.
How it is realised on AWS
IRSA binds each model service's Kubernetes service account to an IAM role; the serving API resolves the role to an authorised group set held in the registry and caches it with the compiled plan. Human access is corporate OIDC to the registry API, which checks group ACL and purpose before asking Lake Formation for a scoped credential against the Iceberg tables. Grants expire and must be renewed.
Options weighed
  • ChosenWorkload identity per caller, group-granular authorisation, explicit denial: No secret to rotate, and a narrower vector is never silent
  • RejectedPer-consumer API keys checked at the serving edge: Fastest to build and puts a long-lived shared secret in four deployment pipelines
  • RejectedAuthorisation per feature rather than per group: Finer control and pushes access control into the read path and the data model at once
  • RejectedNetwork-level authorisation only, trusted VPC: Cheapest and makes every workload in the VPC equivalent to every other
Consequences
What it buys
  • There is no long-lived credential to leak or rotate on the serving path
  • A model that loses authorisation for a group learns about it explicitly, so an accuracy change has a cause
  • Human and machine access are governed by different checks, which matches the fact that they are different risks
What it costs
  • Group granularity means authorising a consumer for a group authorises it for every feature in that group
  • Splitting a group for access reasons makes the permissions model influence the materialisation unit
  • Resolving purpose for human reads adds a control-plane dependency to a path that would otherwise be pure data
Choose differently when
If features carried genuinely per-feature access requirements, per-feature authorisation would be necessary and the read path would have to absorb the cost. Where that arises, the answer here is a new group, priced rather than hidden.
Why it holds up over time
Short-lived, platform-issued workload identity has replaced shared secrets everywhere it has been available, and the direction of travel is one-way. The mechanism will change; the absence of a shared key will not.
LessonAuthenticate the workload, not the deployment, and refuse out loud. A silently narrower answer is worse than an error, because only the error gets investigated.
Shown on views20 19 08

Every package used, in one table

The terms below are used precisely in this package. Several are used loosely in the wider feature-store literature, and the difference matters when reading the decision records.

PackageWhat it isWhat it does hereConsidered instead
Feature A named, typed signal about one entity, defined declaratively and materialised by the platform. The unit of definition, versioning and consumer pinning. A column in a training table, which is a feature's output rather than the feature.
Feature group A set of features sharing an entity key, a source and a refresh cadence. The unit of materialisation, authorisation, freshness SLO and ownership. A namespace, which carries none of those four responsibilities.
Offline store The full history of every materialised value, carrying both event and ingestion timestamps. The source of truth for batch features and the origin of every rebuild. A data warehouse, which holds the sources rather than the materialised features.
Online store Latest value per entity and group, plus event timestamp and definition version. A derived, disposable projection read inside a user-facing request. A cache, which implies a read-through path this design deliberately does not have.
Point-in-time correctness Every value in a training row is the one whose ingestion timestamp was at or before that row's timestamp. The property that makes a training set reflect what production could actually have read. An as-of-event-time join, which selects values the platform did not yet know and inflates offline accuracy.
Event time and ingestion time When the fact became true, and when the platform could first have read it. The two columns on every offline value; joins select on ingestion time. Processing time, which cannot represent a late-arriving fact.
Training–serving skew A difference between the value a model was trained on and the value it is served, beyond the feature's declared tolerance. The defect the whole architecture exists to make detectable and attributable. Data drift, which is the world changing rather than the platform disagreeing with itself.
Materialise down Writing the online store from a committed offline snapshot rather than from the original computation. The write path for batch features, and the rebuild path for every feature. Dual-write, which is the write path for streaming features only.
On-demand feature A transformation over request payload and fetched values, computed at read time from the shared definition artefact. The parity boundary: the features most likely to skew, and the ones where parity is strongest. A derived field in the model service, which is where this skew traditionally comes from.
Spine The entity-and-timestamp rows a consumer supplies to request a training set. The left side of every point-in-time join, and one of the three pinned inputs in the manifest. A label set, which is what the spine usually carries but is not what the platform needs from it.
Criticality A consumer's declaration of whether an absent feature degrades its model or fails its request. What decides between a partial vector and an error, per consumer and per group. A feature's importance, which is a modelling property and not a platform contract.
Backfill Materialising a new or changed feature over historical data in a bounded, budgeted window. A routine daily operation with its own capacity, not an exceptional one. A migration, which happens once; a backfill happens with every new feature.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.