| Orchestrated instance state |
EKS EndpointSlice watch across 12 clusters, plus ECS task state change events |
Amazon Web Services |
Self-registration from each workload; Consul agents |
The orchestrator already holds reconciled desired state, so a missing instance is a diff rather than a silence |
ADR-03 |
| Non-orchestrated instance state |
EC2 Auto Scaling lifecycle hooks plus an authenticated registration API with leases |
Amazon Web Services |
Agent-based self-registration everywhere |
Covers fleets nothing watches without extending self-registration to the orchestrated majority |
ADR-16 |
| Desired state store |
Aurora PostgreSQL, 3-AZ, one writer per region |
Amazon Web Services |
DynamoDB global tables; etcd; Consul |
Small, exact, relational state with strong in-region consistency; RPO 0 matters more than write scale here |
ADR-02 |
| Observed state store |
DynamoDB on-demand with a 7-day TTL |
Amazon Web Services |
Timestream; a Prometheus-compatible TSDB; Cassandra |
Sized for 1.2 M writes/second where losing a window of samples is acceptable by design |
ADR-02 |
| Health signal ingest |
Kinesis Data Streams, sharded by service |
Amazon Web Services |
MSK; SQS; direct gRPC to the evaluator |
Shards give per-service isolation so one service's churn cannot delay another's evaluation |
ADR-04 |
| Active probing |
Per-AZ DaemonSet probers with bounded target assignment |
Open source on EKS |
Route 53 health checks for everything; a central prober fleet |
Sub-linear probe cost in instance count, and a prober failure scoped to one AZ rather than a region |
ADR-04 |
| Passive outcome reporting |
Envoy per-endpoint outcome counters reported to ingest |
Open source (Envoy) |
Application-level reporting; inference from service-mesh telemetry |
The only evidence class that measures what callers experience and that a workload cannot forge about itself |
ADR-04 |
| Eligibility evaluation |
Stateless evaluator on EKS, sharded by service, state in memory |
Open source on EKS |
Stream-processing framework; evaluation inside the proxies |
Computed state is disposable, so the evaluator can be restarted and rebuilt rather than recovered |
ADR-02 |
| View propagation |
Envoy-compatible xDS tier with incremental deltas and per-service versions |
Open source (xDS) on EKS |
Consul; short-TTL DNS only; gossip; long-poll |
The only mechanism whose propagation latency the platform owns, which a 5 s withdrawal budget requires |
ADR-10 |
| DNS compatibility surface |
Cloud Map plus Route 53 private zones on a short TTL |
Amazon Web Services |
CoreDNS only; no DNS surface at all |
Serves clients that cannot hold a subscription, with the weaker guarantee stated rather than hidden |
ADR-10 |
| Client data plane |
Envoy sidecar with a durable last-known-good cache, plus a thin resolver library |
Open source (Envoy) |
Client-side load balancing per language; a central routing proxy tier |
Routing authority has to live where the request is, and it has to be identical in every runtime |
ADR-01 |
| Runtime edge membership |
NLB and ALB target groups synchronised from the published views |
Amazon Web Services |
Target group health checks as the source of truth |
One eligibility semantics for edge-routed and sidecar-routed traffic, with no second opinion |
ADR-01 |
| Region health and failover signal |
Route 53 Application Recovery Controller readiness checks and routing controls |
Amazon Web Services |
Registry-derived region health; a custom quorum service |
A region's reachability must not be asserted by the system whose reachability is in question |
ADR-14 |
| Workload identity |
IRSA for pods, instance profiles for fleets, mutual TLS on every control-plane path |
Amazon Web Services |
SPIFFE/SPIRE; network-position trust |
Self-only health reporting is only enforceable if every signal carries an identity the platform can check |
ADR-15 |
| Audit and transition record |
Kinesis to S3 with Object Lock, 5-year retention |
Amazon Web Services |
CloudTrail only; the relational store |
Denials of service must be reconstructable long after the samples that caused them have expired |
ADR-15 |
| Policy and contract authoring |
Policy API with staged rollout, disruption budgets and one-step reversal |
Open source on EKS |
GitOps-applied CRDs; direct registry writes |
Health-check configuration can drain a fleet, so it needs the same gates as a code deploy |
ADR-08 |