# Architecture One-Pager

*Feature Store · Solution Architecture v1.0 · Amazon Web Services · Data Platform Architecture · 2026-09 · 21 views · 14 architecture decision records*

**A feature exists only as a compiled definition with two materialisations of one computation, and the online store is a rebuildable projection of the offline store — never an independently written system of record.**

When someone opens a delivery app at eight in the evening and sees "arriving in 32 minutes", a model produced that number from a few hundred signals: how busy this restaurant has been in the last fifteen minutes, how long this courier's last three pickups took, what rain does to this neighbourhood on a Friday. Nine teams own models that need those signals, and the hard part is not computing them. The hard part is that a signal computed one way in a training notebook and another way in a request path produces a model that scores 0.91 offline and performs like 0.74 in production — and the bug is nearly impossible to find afterwards, because both numbers are individually plausible and the training data no longer exists in the form that produced them. The second problem is organisational: all nine teams need "orders at this store in the last 30 minutes", each builds it differently, and none of them can find the others' version.

Hold every feature as a declarative definition in reviewed source control, and compile it into two execution plans — an offline plan and an online plan — from one artefact, refusing to publish anything that cannot express both. Materialise streaming features from one computation writing to both stores, and batch features into the offline store with the online store materialised down from it. Write two timestamps onto every offline value, event time and ingestion time, so a training set can be assembled as of any instant by joining the value that was readable then. Treat the online store as a derived, latest-value projection with weak durability and absolute correctness, rebuildable from the offline store and a seven-day stream log. Serve a vector inside a 25 ms client budget with per-feature status, so absence is always an explicit value with a reason code rather than a silent zero. Then prove the whole claim continuously: log a sample of served vectors, replay them through the offline path, and treat disagreement as a defect with a version stamp pointing at its cause.

## What it is, and what it is not

- **One definition compiled into two materialisation plans** — not Two pipelines that are supposed to agree
- **An online store that is derived and disposable** — not A fast database that is the system of record for features
- **Absence returned as a value with a reason code** — not A zero substituted where a value was missing
- **A training set pinned to snapshots, versions and a lookback** — not A query over whatever the feature tables hold today
- **Skew measured continuously as a control** — not A skew dashboard nobody opens
- **A catalogue with owners, adoption and lineage** — not A wiki page listing what exists

## The decisions that are the architecture

1. **One definition, compiled into two plans, or it does not publish** (ADR-01) — The compiler emits an offline plan and an online plan from a single artefact and fails the build if either cannot be expressed. That single gate is the parity guarantee, mechanised rather than reviewed.
2. **On-demand transformations run in the serving path, from the same artefact** (ADR-02) — The transformations most likely to skew are exactly the ones that cannot be precomputed. They are compiled from the shared definition and sandboxed inside a 2 ms budget, which buys parity where it is most at risk and accepts consumer code in the serving failure domain.
3. **Streaming features dual-write; batch features materialise down** (ADR-03) — Freshness where it is required, provenance where it is affordable. Two write topologies rather than one, with the version stamp and nightly replay as the detector for the partial-failure risk dual-write introduces.
4. **Skew is a continuous control, not a report** (ADR-04) — One per cent of served vectors are logged and replayed nightly through the offline path. That number is the only ongoing evidence that the two materialisations still agree, and the platform is exactly as trustworthy as it.
5. **Two timestamps on every offline value, or the feature is refused** (ADR-05) — Event time says when the fact became true; ingestion time says when the platform could first have known it. Without both, a point-in-time join is guesswork, and the platform refuses the feature rather than approximating it.
6. **The online store is derived, latest-value and disposable** (ADR-08) — The largest and hottest store in the platform has the weakest durability requirement and the strictest correctness requirement. It is rebuildable in 90 minutes, which makes online materialisation a cost lever and regional recovery a rebuild rather than a failover.
7. **Absence is a value; criticality belongs to the consumer** (ADR-09) — A missing feature returns a reason code and the declared default that training also used, and only a consumer-declared critical group fails the call. A fraud model may fail closed where a ranking model would rather rank slightly worse.
8. **The feature version is the unit of versioning and of pinning** (ADR-11) — A transformation, window or dtype change creates a new immutable version and leaves pinned consumers where they are. Training sets name the exact versions they were built from, which is what makes a regenerated set bit-identical.
9. **Definitions arrive only through reviewed source control** (ADR-12) — A feature that can be created by an API call is a feature with no review, no version history and no duplicate check. Registration is a merge, and near-duplicates are blocked there rather than discovered a year later.
10. **Backfill is isolated, budgeted and pre-emptible** (ADR-14) — Rewriting history is normal in this platform and must never be able to take serving with it. The cost is estimated before the run and the pool yields to live materialisation.

## Why this should still be right in ten years

Table formats, stream processors and key-value stores will all be replaced inside the life of this platform. These are the properties that should outlast them.

- **Compilation outlives the engines.** The contract is that one artefact produces both materialisations. Whether the offline engine is Spark, a warehouse, or something that does not exist yet, and whether the online store is a key-value store or a format nobody has shipped, the guarantee is unchanged and every existing definition survives the swap.
- **Two timestamps are a property of the world, not of a product.** The gap between when a fact became true and when a system could know it is created by physics and network partitions, not by a vendor. Any platform that wants reproducible training data will need both columns, whatever it is built on.
- **Derived state stays cheap to be wrong about.** As long as the online store is a projection, its technology is an implementation detail and its loss is an inconvenience. The decision that will age worst in most feature stores — betting correctness on a fast store — is the one this design refuses to make.
- **Explicit absence survives every model generation.** Reason codes and declared defaults are a contract between platform and consumer that does not care whether the consumer is a gradient-boosted tree, a transformer, or whatever replaces them.
- **Skew replay is a measurement, not a tool.** Replaying served inputs through the batch path to check agreement is a technique, not a product. It will still be the cheapest available proof of parity in ten years.
- **The organisational claim is the durable one.** The platform's real product is that nine teams share one definition of a signal. That value grows with the number of teams and models, and it is the last thing a replacement platform would be allowed to give up.

## Non-functional targets

Every figure below is a stated assumption for this design, chosen to be defensible and arguable rather than measured. A reviewer who changes one can follow it to the decision that depends on it.

| Quality | Target | How it is met | View |
|---|---|---|---|
| Serving availability | ≥ 99.95% monthly for the online path | Three AZs, stateless serving pods, a derived store that can lose a partition without losing truth | 15 |
| Online read latency | p50 3 ms, p99 15 ms, p99.9 40 ms server-side; p99 25 ms client-observed | Locally cached compiled plans, one batch read across groups, short-TTL cache for hot keys | 12 |
| Streaming freshness | event-to-readable p99 ≤ 5 s; windows ≤ 30 min readable p99 ≤ 30 s after close | Checkpointed stream processor dual-writing both stores from one computation | 13 |
| Batch freshness | hourly groups within 20 min of the hour; daily groups by 06:00 local p95 | Dependency-aware scheduling with partition-level retry, separate from the backfill pool | 13 |
| Read throughput | 320,000 vector reads/second peak, 3× burst for 10 minutes | On-demand key-value capacity, per-consumer quotas, request coalescing on hot keys | 08 |
| Offline join performance | 500M spine rows × 200 features p95 ≤ 45 min, ceiling 90 min | Snapshot-isolated as-of reads with partition pruning on entity type, group and event date | 14 |
| Point-in-time correctness | zero violations tolerated; regenerated sets bit-identical | Event and ingestion timestamps on every offline value; snapshot, version and lookback pinned in the manifest | 11 |
| Training–serving skew | ≤ 0.1% of sampled requests per feature beyond declared tolerance | 1% serving log replayed nightly through the offline path, attributed by definition version | 18 |
| Staleness detection | group marked stale within 60 s of exceeding 2× its freshness SLO; owner alerted ≤ 5 min | Freshness SLO per group evaluated continuously, marker surfaced in serving responses | 17 |
| Recovery | registry RPO 0 / RTO 1 h; offline RPO 15 min / RTO 4 h; full online rebuild ≤ 90 min; serving RTO 15 min cross-region | Small authoritative registry, re-derivable offline store, disposable online store rebuilt rather than replicated | 10 |
| Time to production | merged definition readable online ≤ 30 min for streaming, ≤ 1 cadence for batch | Compile-and-publish pipeline with the parity gate as the only blocking check | 16 |
| Cost | ≤ $0.40 per million feature-vector reads; zero-read online features reported at 30 days | Online materialisation flag per feature, cost attributed to producing and consuming teams per group | 17 |

## Scope

**In scope**

- Declarative feature definitions, a registry, and a compiler emitting an offline and an online plan from one artefact
- Batch and streaming materialisation, with contract checks that quarantine rather than publish a bad batch
- An offline store with event and ingestion timestamps, and point-in-time training set assembly with a reproducible manifest
- A low-latency online serving path with per-feature status, declared defaults and per-consumer quotas
- On-demand request-time transformations compiled from the shared definition artefact
- Freshness SLOs, staleness marking, skew replay and drift attribution
- A catalogue with owners, lineage, adoption evidence and duplicate detection
- Feature-group authorisation, sensitivity classification and entity-level erasure

**Explicitly out of scope**

- Model training, model serving and the model registry — the platform serves them and does not host them
- Experiment tracking, hyperparameter search and evaluation methodology
- Labelling and annotation
- The source systems that emit events, and their schemas
- Business intelligence reporting over the same underlying data
- Feature selection — which features a model should use is the data scientist's decision

## What a four-week prototype should prove

The prototype's job is to falsify the central claim — that one definition can be compiled into two materialisations that provably agree — on two features and one model, not to build a platform.

1. Compile one streaming feature and one batch feature from a single definition artefact into both an offline and an online plan, and show the build failing when only one plan can be expressed
2. Write both stores, then deliberately corrupt the online value for one entity and show nightly skew replay finding it and naming the definition version
3. Assemble a training set over a 10M-row spine and prove point-in-time correctness by showing that a value whose ingestion timestamp postdates its spine row is excluded
4. Regenerate the same training set from its manifest a week later and show it bit-identical after new data has landed
5. Delete the online store for one feature group and rebuild it from the offline store inside ten minutes with serving up throughout
6. Serve a vector with one group deliberately unavailable and show the reason code, the declared default, and the same default applied in the training path

- A source event schema changes mid-run: the batch is quarantined, the last good values keep serving, and the owner is alerted rather than the values being published
- The stream processor falls three minutes behind: the group is marked stale within 60 s and the reading model sees the age, not a fresh-looking number
- One feature group is made unreachable: the vector returns with a reason code and the declared default, and a consumer that declared the group critical fails its call instead
- A definition is edited to change a window: a new version is published, the pinned consumer stays on the old one, and both versions materialise side by side
- A thirteen-month backfill runs during the morning batch wave: it is pre-empted, restarts at partition grain, and no scheduled group misses its landing time

## Open risks, carried rather than hidden

| Risk | If it lands | Response |
|---|---|---|
| The parity gate is slower than writing a feature by hand, so teams route around it | Features get computed in model services again, the central claim quietly stops being true, and the platform becomes a catalogue of what some teams did | Treat time-to-production as a first-class NFR at 30 minutes and measure it per team; keep the parity gate the only blocking check in the pipeline; make the catalogue and backfill estimate good enough that using the platform is the fast path rather than the compliant one |
| The online store rebuild path is documented but never exercised | The 90-minute rebuild and the 15-minute cross-region RTO are both fiction, and the decision to treat the online store as disposable becomes the platform's largest single point of failure | Rebuild one production feature group on a schedule as a routine operation rather than a drill, and publish the measured time alongside the SLO so a regression is visible |
| On-demand transformations become a general-purpose compute surface in the serving path | Consumer-authored code grows past its 2 ms budget, and a bad transformation takes the serving tier down for every model, not just its own | Sandbox with a hard CPU and wall-clock ceiling enforced per call, cap the transformation surface to pure functions over request payload and fetched values, and hold on-demand features to phase 2 so the MVP proves parity before it accepts the risk |
| Group-granularity authorisation is too coarse for restricted features | Teams split groups to model access control, the group stops being a materialisation unit, and storage layout follows the permissions model | Keep sensitivity classification at the group and enforce it at registration; where a feature genuinely needs narrower access, accept the new group as the answer and price it, rather than adding per-feature ACLs to the read path |
| Nightly skew replay is too slow for a regression that appears at peak | A bad definition version serves wrong values for up to a day before the control that exists to catch it reports anything | Keep the per-consumer health view as the live signal for freshness and absence, raise the sampling rate per model on demand to 100%, and treat replay as the proof of parity rather than the alarm for it |

The reasoning behind every component and technology choice is in the [Architecture Decision Record](decision-record): 17 records across 6 areas, each with the alternatives that lost and what the choice costs.
