Enterprise Metadata Management System  ·  View 22 of 22  ·  Assurance

Failure Modes and Recovery

What is allowed to fail, what is not, and the risks that are being accepted with their eyes open.

Editable source SVG draw.io All views
Tier 1 · read path · 99.9% · RTO 30 min · must not fail
Tier 1 · read path · 99.9% · RTO 30 min · must not fail
Search & read API
3 AZ · autoscaled
Search & read API...
Search Index
3 replicas
Search Index...
Read Cache
serves on index loss
Read Cache...
Stale-read banner
shows lag age
Stale-read banner...
Tier 2 · write and ingest · may lag, must not lose
Tier 2 · write and ingest · may lag, must not lose
Metadata Event Bus
7-day retention
Metadata Event Bus...
Aspect Store
RPO 5 min
Aspect Store...
Connector failure
quarantine + alert
Connector failure...
Poison payload
dead letter, skip
Poison payload...
Tier 3 · projections · disposable
Tier 3 · projections · disposable
Knowledge Graph
rebuild under 4 h
Knowledge Graph...
Index rebuild
replay change log
Index rebuild...
Payload Archive
replay beyond 7 days
Payload Archive...
Standing risks · owned, not hidden
Standing risks · owned, not hidden
Graph hot spot
wide lineage fan-out
Graph hot spot...
Source rate limits
harvest falls behind
Source rate limits...
Stale curation
recertification overdue
Stale curation...
Shadow catalogs
the silo returns
Shadow catalogs...
Backup vault
PITR 35 days
Backup vault...
Warm standby region
RTO 30 min
Warm standby region...
fall back to cache
fall back to cache
replay
replay
deep replay
deep replay
continuous backup
continuous backup
global replication
global replication
Failure Modes, Tiers and Recovery
Failure Modes, Tiers and Recovery
Application we own
Application we own
Data store
Data store
Risk / gap
Risk / gap
Queue / topic
Queue / topic
External / third party
External / third party
failure / alternate
failure / alternate
batch
batch
event / async
event / async
Adoption, not infrastructure, is the largest risk: a catalog nobody trusts is replaced by spreadsheets, which is why coverage and certification are on the operations dashboard in view 18.
Adoption, not infrastructure, is the largest risk: a catalog nobody trusts is replaced by spreadsheets, which is why coverage and certification are on the operations dashboard in view 18.
v 1.0 · owner Data & AI Architecture · date 2026-08
v 1.0 · owner Data & AI Architecture · date 2026-08
Text is not SVG - cannot display

Availability tiers

  • Tier 1, the read path, must not fail: multi-AZ, three index replicas, and a read cache that keeps hot profiles answerable through an index outage.
  • Tier 2, write and ingest, may lag but must not lose: the change log is the durable buffer and the payload archive is the deep replay source.
  • Tier 3, the projections, are disposable by design. Losing the graph costs a four-hour rebuild, not a restore.

Recovery objectives

  • RTO 30 min, RPO 5 min for the aspect store; PITR to any point in 35 days.
  • Projection rebuild under 4 hours from the change log, longer if it must replay from the archive.
  • DR rehearsal twice a year, including a full projection rebuild — the recovery path most likely to have rotted.

Accepted risks

  • Adoption, not infrastructure, is the largest risk. A catalog nobody trusts is replaced by spreadsheets, which is why coverage and certification are operational metrics in view 18.
  • Source rate limits can push harvest freshness past its target for large estates; the mitigation is prioritised scheduling for certified assets, and it is a stated limitation, not a solved problem.