Service Mesh Platform · View 12 of 31 · 4 · Data
Decisions
- Git is the source of intent; each cluster's etcd holds the applied copy. Losing a cluster's etcd costs a resync from Git. Losing Git would cost history, so Forgejo runs on a synchronous PostgreSQL standby with read-only mirrors at every site.
- Compiled xDS snapshots are never stored. istiod rebuilds them from the Kubernetes API in minutes, and a persisted snapshot would only be a second thing that could disagree with intent.
- Rendered, signed bundles are stored in Harbor. They are what a one-step revert pins to, so a revert never waits on a recompile.
Targets
- Intent RPO 0, RTO ≤ 15 min. CA material RPO 0, intermediate reissue ≤ 30 min. Snapshot rebuild ≤ 5 min. Telemetry RPO ≤ 5 min.
- Retention: metrics 15 days raw and 13 months rolled up, traces 7 days, access logs 30 days, issuance records and audit trail 13 months.
Risks
- Issuance records are written from SPIRE's audit log through the telemetry path, which drops under backpressure. A nightly reconciliation against SPIRE's datastore finds gaps; an Object Lock copy makes them immutable once found.