Feature Store  ·  View 21 of 21  ·  Assurance

Failure Classes

Eight named failures, what detects each, what the platform does immediately, and what it must never do instead.

Editable source SVG draw.io All views
Detection Immediate behaviour Recovery If it goes wrong Stale upstream Age > 2× SLO in 60 s Mark stale, flag age Late batch lands Silent old value Streaming lag Lag SLO breach Stale, not fresh Replay from checkpoint Double-counted window Schema drift at ingest Type, null, range check Quarantine the batch Owner fixes source Bad batch published Partial vector Per-group read status Reason code + default Group returns Whole request failed Online partition loss Read error rate Fail over in SLO Rebuild ≤ 90 min Untested rebuild path Hot key Per-key read count Coalesce + short TTL Projection for the model Tail latency blowout Definition mismatch Version stamp compare Stale, block promotion Re-materialise group Undetectable skew Region loss Zonal then regional health AZ loss: no SLO impact RTO 15 min, rebuilt Cold second region Failure Classes — Detection, Behaviour, Recovery Every row's right-hand cell is the failure the platform is designed to make impossible, not the one it accepts. v 1.0 · owner Data Platform Architecture · date 2026-09

Decisions

  • Every row's right-hand cell is the failure the design makes impossible, not one it accepts. Read the matrix right-to-left to see what the architecture is actually for.
  • 'Flake' is not a root cause. A job is re-run only to confirm a failure that names a service the change does not touch, or one that died before any work began.
  • No failure class is handled by skipping, disabling or quietly defaulting a feature. Absence is always a value with a reason code.

Assumptions

  • Staleness detected within 60 s of exceeding 2× the freshness SLO; owner alerted within 5 min.
  • Serving continues from cached compiled plans for at least 60 minutes without the registry.

Risks

  • Three of the eight classes recover by replay from a 7-day stream log. A failure lasting longer than the retention window has no recovery path except re-derivation from the source systems, which the platform does not control.