# Feature Store

**Solution Architecture v1.0 · Amazon Web Services · Data Platform Architecture · 2026-09 · 21 views · 17 architecture decision records**

When someone opens a delivery app at eight in the evening and sees "arriving in 32 minutes", a model produced that number from a few hundred signals: how busy this restaurant has been in the last fifteen minutes, how long this courier's last three pickups took, what rain does to this neighbourhood on a Friday, how often this user has cancelled. The same signals decide which stores appear first, whether a payment is waved through, and where surge lands. Nine teams own those models. This package is the feature store behind them — for an assumed consumer marketplace of 40M monthly active users, 1.1M partner stores, 180k couriers, 2.2M orders a day peaking at 4,500 a minute, 45 production models and 1,400 features in 120 groups across 6 entity types, built on S3 with Apache Iceberg for the offline store, DynamoDB for the online store, Kinesis with Managed Flink for streaming aggregates, EMR Serverless for batch and backfill, Aurora PostgreSQL for the registry, and gRPC services on EKS behind an internal load balancer.

The product is not the ability to compute a signal. It is **the ability to guarantee that the number a model trains on and the number it is served are the same number** — and to prove that guarantee every night rather than assert it.

The design rests on one rule: **a feature exists only as a compiled definition with two materialisations of one computation, and the online store is a rebuildable projection of the offline store — never an independently written system of record.**

The decisions that carry the design:

- **One definition artefact, compiled into both plans, or it does not publish.** The compiler emits an offline plan and an online plan from one source and fails the build if either cannot be expressed. Nothing reaches the registry without both, and materialisation engines execute a compiled plan rather than reading a definition. That single gate is the parity guarantee, mechanised rather than reviewed (ADR-01).
- **The transformations most likely to skew run inside the platform.** A distance computed at request time cannot be precomputed, and it is exactly where a Python implementation and a Go implementation historically diverge. On-demand features are declared in the same artefact, sandboxed inside 2 ms, and replayed from the same compiled expression during point-in-time assembly (ADR-02).
- **Streaming features dual-write; batch features materialise down.** Freshness where the SLO requires it, a single unambiguous source of truth where it is affordable. Two topologies with a stated rule, rather than one topology someone has to work around (ADR-03).
- **Skew is a continuous control, not a report.** One per cent of served vectors are logged and replayed nightly through the offline path as an as-of read at the request's own timestamp. That number is the only ongoing evidence the two materialisations still agree, and the platform is exactly as trustworthy as it (ADR-04).
- **Two timestamps on every offline value, or the feature is refused.** Event time says when the fact became true; ingestion time says when the platform could first have known it. Joins select on ingestion time, because a join on event time trains the model on information production will never have (ADR-05).
- **Snapshots belong to the format; lookback and tolerance belong to the join.** Iceberg gives snapshot isolation so a 45-minute training join reads a stable view while materialisation commits, and the snapshot id becomes a pinnable input. The engine owns the per-row as-of predicate, the declared lookback and reason-coded nulls (ADR-06).
- **Late data changes the next training set, never the last one.** A late event is accepted, timestamped honestly, and becomes visible to sets generated afterwards. Already-generated sets are never restated — reproducibility is worth more than retrospective accuracy (ADR-07).
- **The largest and hottest store is the disposable one.** The online store holds latest value only and is rebuildable in 90 minutes, which inverts its requirements to weak durability and absolute correctness, makes online materialisation a cost lever, and makes regional recovery a rebuild rather than a failover (ADR-08, ADR-15).
- **Absence is a value; the consumer chooses what it means.** Every response carries per-feature status, a reason code and the declared default that training also applied. Whether a missing feature degrades the model or fails the request is the model owner's declaration — a fraud model may fail closed where a ranking model would rather rank slightly worse. The silent zero is the most expensive bug this platform could ship (ADR-09).
- **Reuse is enforced where it is still cheap.** A feature is created by merging a definition, never by an API call, and a near-duplicate is blocked at registration rather than discovered in production a year later. Every group carries a named owner and a freshness SLO, and an unowned group is not materialised (ADR-12, ADR-13).
- **Version at the granularity people consume.** The feature version is the unit of versioning, of consumer pinning and of training-set provenance, which is what makes a regenerated set bit-identical and a skew verdict attributable to a cause (ADR-11).
- **Backfill is routine, so it gets its own capacity.** Rewriting history happens with every new feature. It runs in an isolated, budgeted, pre-emptible pool, and writes online only where a consumer actually reads online (ADR-14).
- **Classification is enforced at registration, not at read.** A feature cannot be derived from a restricted column into a group of lower classification, and the serving log inherits the highest class it sampled — otherwise the evidence trail becomes an unclassified replica of restricted data (ADR-16).
- **No shared keys anywhere.** Every service-to-service caller is a workload identity, authorisation is evaluated per feature group, and a denied group is refused by name rather than silently dropped from the vector (ADR-17).

The architecture one-pager (including why the design should still hold up in ten years, and the five risks that would change it) and the full decision record appear on the landing page of the diagram set, directly below the index of views. The same content is published as [docs/architecture-one-pager.md](docs/architecture-one-pager.md) (~13 min) and [docs/decision-record.md](docs/decision-record.md) (~59 min).

---

## What is here

| Path | Contents |
|---|---|
| `diagrams/index.html` | The landing page: 21 views in seven acts with every format linked, then the **architecture one-pager** and the **decision record** |
| `diagrams/*.html` | One self-contained page per view: the inlined diagram plus the reasoning cards, with copy / PNG / PDF export |
| `diagrams/svg/*.svg` | The same views as SVG with the diagram XML embedded; they re-open fully editable in diagrams.net |
| `diagrams/drawio/*.drawio` | draw.io source |
| `docs/architecture-one-pager.md` | The one-pager as markdown |
| `docs/decision-record.md` | The 17 decision records, the capability-to-technology table and the package glossary as markdown |
| `specs/part-a..b.json` | Diagram specifications, the source of truth for every view |
| `specs/manifest-a..b.json` | Acts, page titles, subtitles and reasoning cards |
| `specs/adr-onepager.json`, `specs/adr-records-a..b.json` | The one-pager, the decision records, the capability-to-technology table and the glossary |
| `scripts/build.sh` | Rebuilds every deliverable from the specs (Node only, no network) |
| `ask.md` | The requirement |

## The twenty-one views

| # | View | What it answers |
|---|---|---|
| 01 | System Context | Who reads a vector, the five sources the platform reads and never writes to, and why training and serving sit outside the boundary |
| 02 | High-Level Architecture | Six stages from a raw event to a trusted vector, with the compiler in the middle and "prove" as a stage rather than an afterthought |
| 03 | Actors and Journeys | Eight actors, three of them machines, two of which are expected to fail in specific designed-for ways |
| 04 | Journey — Ship a Feature | The journey that decides whether the platform gets used at all; the trough is backfill, not authoring |
| 05 | Journey — Features Go Stale at Peak | 19:40 on a Friday; the trough is attribution, and the unrecoverable moment is the silent zero |
| 06 | Layered Architecture | Seven layers, with definition below materialisation because materialisation reads a compiled plan and never a live definition |
| 07 | Containers and Components | One control plane, three data planes, four stores — because there are four different mutabilities to hold |
| 08 | Integration Surface | Five consumer surfaces, one write surface, no public serving endpoint, and nothing written back |
| 09 | Feature Data Flow | Emit to prove, with the contract guard between compute and materialisation and one loop closing left |
| 10 | Storage Zones by Durability | Four zones graded by what it takes to get the data back; the hottest store has the weakest durability requirement |
| 11 | Registry and Value Data Model | Eleven tables, and the two-timestamps-offline / one-timestamp-online asymmetry that is the point-in-time guarantee |
| 12 | Critical Flow — An ETA Vector Read | 25 ms client-observed, inside a 250 ms user request: cached plans, one batch read, explicit per-feature status |
| 13 | Materialisation — Four Paths | Batch, streaming, on-demand and backfill; the two blank cells are the architecture |
| 14 | Point-in-Time Training Set Assembly | Three pinned inputs, one join rule, a manifest that makes regeneration bit-identical |
| 15 | Deployment Architecture | Three AZs, a warm second region at 10%, and a 15-minute RTO met by rebuilding rather than failing state over |
| 16 | Definition Promotion Pipeline | One blocking gate — both plans compile — and a backfill budget estimated before merge rather than invoiced after |
| 17 | Observability | Five signal types across five stages, with data quality as a first-class row and one cell that pays for the platform |
| 18 | The Feature Lifecycle Loop | Define to revise, and the one field — the version stamp — that lets the loop close at all |
| 19 | Security Trust Zones | Five zones, classification enforced at registration, and the serving log inheriting the class it sampled |
| 20 | Identity and Access | A human proves purpose, a workload proves identity, and a denied group is refused by name |
| 21 | Failure Classes | Eight named failures with detection, immediate behaviour, recovery, and what the design makes impossible |

## Rebuilding

```bash
bash scripts/build.sh
```

Node 20+ and nothing else. No draw.io Desktop, no browser, no network. The build runs
generate → geometry gate → routing gate → editable SVG → plain SVG → draw.io → HTML pages and
index → one-pager and decision record → link check, and stops at the first failure.

**Verifier state at hand-over:** 21 files, 0 geometry errors, 0 geometry warnings, 166 edges
routed with 0 routing errors and 7 clutter warnings, 357 icons embedded, 336 nodes with 0
unresolved icons, and every relative link resolving. All seven remaining warnings are on view
01, where the `context` layout fans five nodes per side into one centre and adjacent relation
labels sit close to each other's lines. Reducing them further means removing an actor or a
source, which would cost more than the clutter does.
