Document 14 min read

Architecture One-Pager

Solution Architecture v1.0 · Amazon Web Services with an open-source Envoy/xDS data plane · Reliability Architecture · 2026-10 · 21 views · 16 architecture decision records

Health Check & Service Discovery · Solution Architecture v1.0 · Amazon Web Services with an open-source Envoy/xDS data plane · Reliability Architecture · 2026-10 · 21 views · 16 architecture decision records

The client's last-known-good view is authoritative for routing; the control plane is an advisory service that improves that view and is never in the request path.

When a team chat shows a spinner on send, when a shared document says "reconnecting", when the driver freezes on the map for eight seconds, the user is rarely looking at a crashed service. They are looking at a request routed to a replica that should not have been in rotation — one restarting, one whose connection pool was exhausted, one in a zone already declared unhealthy whose address was still cached in the caller. The opposite failure is less visible and more expensive: a bad health-check config, or a shared database wobbling, and a fleet takes itself out of rotation faster than it can be put back. Something has to answer two questions continuously for every one of twelve million internal calls a second — which instances of this service exist, and which should receive traffic — and that something is itself a distributed system that will have a bad day. The problem is not knowing which replicas are healthy. It is being allowed to be wrong about it, briefly and boundedly, without that being the same thing as an outage.

Reconcile instance existence from the orchestrator's own state rather than asking workloads to register themselves, and collect health as three independent evidence classes: active probes from outside, self-reports from inside, and the outcomes real callers actually observed. Fuse them into one eligibility decision per instance over a decaying score with asymmetric hysteresis, where silence is Unknown rather than unhealthy and a degraded shared dependency reduces weight rather than removing capacity. Assemble per-service versioned views with explicit topology priority tiers, and push incremental deltas to sixty thousand streaming subscriptions, budgeting withdrawal at five seconds and admission at thirty. Every client keeps the last view it received on local disk and routes from it — through a control-plane outage, for at least an hour, with no behaviour change but staleness. Removal is bounded in three independent places: a minimum healthy fraction that disregards health rather than draining a service, a cap on how fast a client's endpoint set may shrink, and a cap on how much of it any one client may eject. Recovery is treated as the dangerous event: slow start, jittered re-subscription and rate-capped shift-back are defaults.

What it is, and what it is not

  • A client cache that is authoritative for routing — not A registry consulted, or leased from, on the request path
  • Fail static on missing information — not Fail closed, which turns a blind observer into an outage
  • Liveness, readiness and dependency health as three questions — not One health boolean per instance
  • Unknown as a first-class state — not Unknown silently treated as unhealthy
  • Withdrawal and admission on separate budgets — not One propagation path with one latency target
  • Removal bounded by a minimum healthy fraction — not Every withdrawal honoured, however many arrive at once
  • Recovery protected by default — not Recovery as the moment everything is allowed back at once
  • Regional authority with explicit reconciliation — not A global registry every region depends on to route
  • Capability declared per service — not The weakest available health signal applied to the whole estate

The decisions that are the architecture

  1. The data plane is authoritative; the control plane is advisory (ADR-01) — A view already delivered keeps working. The control plane's only job is to improve it, which lets its availability target be 99.95% while resolution as callers experience it is five nines, and turns a registry outage into a staleness problem with a bounded blast radius rather than a company-wide outage.
  2. Computed state is disposable; only desired and observed state are durable (ADR-02) — Eligibility decisions and published views are rebuilt from the registry and the sample store, never restored. That is what makes a five-minute RTO a rebuild rather than a backup, and it is why a scoring bug is repaired by recomputing rather than by surgery on live routing state.
  3. Registration is reconciled, not self-reported (ADR-03) — The orchestrator already knows which pods and tasks exist; asking the workload to announce itself adds a failure mode that looks like a healthy fleet with missing members. Explicit registration exists only for the runtimes nothing watches.
  4. Three independent evidence classes, with passive outcomes primary where traffic allows (ADR-04) — A probe measures a path no caller uses, a self-report can be confidently wrong, and the outcomes real callers saw are both free and the only class a workload cannot forge about itself. Fusing all three is what makes grey failure visible.
  5. Unknown is a state, and silence never removes an instance (ADR-05) — A system with only healthy and unhealthy drains itself the first time its observation path breaks. Removal on missing information requires corroboration from a second evidence class.
  6. Dependency health reduces weight; it may not withdraw a shared dependency (ADR-06) — If every replica shares one database, honest unreadiness turns a partially functional service into a completely unavailable one. Unreadiness from a dependency is permitted only where the dependency is declared non-shared.
  7. Asymmetric hysteresis over a decaying score, with a hard-failure fast path (ADR-07) — Re-entry is strictly slower than removal, scoring responds to failure rate rather than the last sample, and a closed port skips the decay entirely — because damping exists to absorb noise, not to argue with a refused connection.
  8. The minimum healthy fraction disregards health rather than draining a service (ADR-08) — Below 50% eligible in a service's zone the platform returns the full endpoint set and says so. A bad health-check deploy is then a declared degradation instead of a self-inflicted outage.
  9. Flapping is quarantined with escalating backoff (ADR-09) — An oscillating instance destabilises every client's view, which is worse than either verdict. Quarantine is surfaced to the owning team as an event, because a quarantined instance is a bug report about a health check.
  10. Streaming subscriptions are the primary surface; DNS is compatibility (ADR-10) — Push with exact version tracking is the only way to hold a five-second withdrawal budget across sixty thousand clients. DNS stays, because a resolver that ignores TTLs is still a client that has to be served.
  11. Withdrawal and admission are budgeted separately (ADR-11) — Five seconds to remove, thirty to add, and under saturation the addition path is the one that sheds. Adding capacity late is cheaper than adding it wrong.
  12. Slow start is published by the control plane as a rising weight (ADR-12) — Uniform across every client and every heterogeneous data plane, which matters more than being traffic-proportional in an estate where the least-updated client library decides whether a client-side ramp exists at all.
  13. Recovery is rate-capped in three places (ADR-13) — Reconnection, traffic share and regional shift-back all have ceilings, because the moment everything is allowed back at once is how a recovered target is killed a second time.
  14. Regional authority, with an out-of-band failover signal (ADR-14) — A region routes with both peers and every global component unreachable. The signal that declares a region unhealthy deliberately does not depend on the registry it is talking about.
  15. The power to deny service is a privilege, and signals must be attributable (ADR-15) — An instance may report only about itself, a prober only about its assigned targets, and an unattributable signal is discarded. Marking an instance ineligible is rate-limited and audited because it is functionally a denial of service.
  16. One registry over heterogeneous runtimes, with capability declared per service (ADR-16) — A caller gets one resolution surface whether its callee is a pod, a fleet instance or somebody else's endpoint — and a service that cannot supply a readiness signal says so, rather than the whole estate being reduced to the weakest contract.

Why this should still be right in ten years

Orchestrators, mesh data planes, discovery protocols and managed DNS will all be replaced inside a decade. The parts of this design that should outlive them are the ones that are statements about authority and blast radius rather than about products.

  • Authority is not a protocol. "The thing holding the view routes; the thing computing it advises" is a claim about where authority lives. It survives replacing xDS with whatever follows it, Envoy with another data plane, and the registry store with anything, because none of those changes which copy of the truth is allowed to route.
  • Absence of evidence outlives every monitoring stack. A system that cannot distinguish "I know it is broken" from "I cannot see it" will eventually drain a healthy fleet, whatever is doing the observing. Unknown as a first-class state costs almost nothing and removes the largest self-inflicted risk in the design.
  • Removal must be bounded in more than one place. The cost of a wrong removal scales with fleet size; the cost of a wrong retention does not. That asymmetry is arithmetic, not fashion, so a minimum healthy fraction, a shrink cap and an ejection cap will still be the right three bounds when every component here has been replaced.
  • Recovery stays more dangerous than failure. Failure sheds load and recovery concentrates it, because demand does not disappear while something is down. Slow start, jitter and rate-capped shift-back are consequences of queueing, which no technology change repeals.
  • What will date fastest. The propagation mechanism is a current answer to a permanent question, and the numbers are a reference estate rather than a measurement. Expect both to be replaced; expect the four statements above to survive it.

Non-functional targets

Every target below is a stated assumption for this exercise, chosen so that a reviewer can disagree with one and follow it to the decision that depends on it.

Quality Target How it is met View
Resolution availability, as a caller experiences it ≥ 99.999% monthly Served from a durable per-client cache; the control plane is not consulted on the request path 02
Control plane availability ≥ 99.95% monthly per region Regional authority, 18 replicas across 3 AZs, no global request-path dependency 16
Static-serving window ≥ 60 minutes, behaviour unchanged Last-known-good cache on local disk, verified by a routine production removal drill 11
Resolution latency p99 ≤ 1 ms from cache; ≤ 50 ms cold In-process lookup against the cached view; cold path only on first subscription 08
Unplanned withdrawal propagation 99% of callers in 5 s, 99.9% in 15 s Incremental delta push over streaming subscriptions, withdrawal path never shed 13
Graceful drain visibility 99% of callers in 2 s, 99.9% in 5 s Drain-to-termination interval derived from measured propagation delay 04
Admission propagation 99% of callers within 30 s Deliberately slower than withdrawal; the addition path is the one that sheds under load 10
Detection of a failing replica 3 s p50, 10 s p99 from first failed request Passive caller outcomes fused with probe and self-report; hard failures skip the decay 15
Eligibility stability ≤ 4 changes per instance per 10 min Asymmetric hysteresis, decaying score, escalating quarantine; damping adds ≤ 5 s to p99 detection 15
Blast radius of a wrong removal ≥ 50% eligible per service per zone Minimum healthy fraction disregards health; client shrink cap 25% per 10 s; per-client ejection cap 21
Churn absorption 40,000 state changes/min for 10 min Evaluation sharded by service, propagation scaled by subscription count and change rate 06
Signal throughput 1.2 M results/second Bounded prober assignment, write-sized observed-state store with a 7-day TTL 17
Recovery RTO 5 min control plane; RPO 0 desired state Computed views rebuilt from desired and observed state on a rehearsed schedule; request path does not stop 11
Cost ≤ $0.012 per instance per day; probes ≤ 0.5% of internal bytes Passive signal preferred where traffic permits; per-service cost attribution exposed to owners 17

Scope

In scope

  • Registration and non-reusable instance identity across orchestrated and non-orchestrated runtimes, with leases
  • Health collection as three independent evidence classes: active probe, self-report, and caller-observed outcomes
  • Eligibility evaluation with hysteresis, flap quarantine, weights and the minimum healthy fraction guard
  • Per-service versioned views with topology priority tiers, published over streaming, DNS and a query API
  • Drain, disruption budgets, deploy ramps, and zone or region evacuation as a reversible preference shift
  • Recovery behaviour: slow start, reconnection caps, rate-capped shift-back, and the static-serving window
  • Signal attribution, resolution authorisation, and the audit trail behind every eligibility transition

Explicitly out of scope

  • Load-balancing algorithms beyond the endpoint set and weights the platform hands to them
  • The service mesh's policy and mTLS planes, which consume these views rather than producing them
  • The public API edge and anything facing an external client
  • The observability platform that stores the signals this plane exports
  • Application readiness logic itself — the platform defines the contract, the service decides what passing means

What a four-week prototype should prove

The prototype's job is to falsify the three claims the whole design rests on: that routing survives the control plane being switched off, that a withdrawal reaches a realistic client population inside five seconds while the system is busy, and that a deliberately broken health check degrades a service rather than draining it. Everything else in this package is a consequence of those three holding, and nothing else is worth four weeks.

  1. One EKS cluster, 40 services, 2,000 instances, Envoy sidecars with a persisted view cache
  2. One regional control plane: reconciler, prober DaemonSet, passive outcome ingest, evaluator, xDS tier
  3. One non-orchestrated fleet behind the registration API, and one declared third-party endpoint with no readiness signal
  4. Synthetic caller traffic with per-endpoint outcome reporting, and a load generator able to produce a 16× churn burst
  • Kill the control plane for 60 minutes under steady traffic: error rate must not move, and a restarted sidecar must route from disk without blocking
  • Fail one replica hard, and separately make one replica slow but alive: measure time to 99% of callers for each, and show the probe-green case is still caught
  • Push a health-check configuration that fails every replica of one service: the service must degrade with a declared 'health disregarded' view and an alert, not drain
  • Evacuate one zone during a 40,000-changes-per-minute burst: the withdrawal budget must hold while the addition path visibly sheds
  • Restart the xDS tier with 2,000 subscriptions attached: reconnection must stay under the 5%-per-10-s ceiling and no client view may regress
  • Partition the prober network from one AZ: those instances must become unknown and stay in rotation, not be withdrawn

Open risks, carried rather than hidden

Risk If it lands Response
The 60-minute static window is a claim nobody tests, so the first real control-plane outage is also the first test The design's central promise fails exactly when it is needed, and the blast radius is every service at once Routine production removal of the control plane from a sample of traffic, with static-serving minutes as a reported signal rather than an incident-only metric
Client library heterogeneity means the client-side guards do not actually exist everywhere Shrink caps and ejection caps are assumed by the control plane and absent in the oldest 5% of callers, which are the ones that will drain a service Report library version spread as a first-class signal, gate onboarding on a minimum version, and keep the control-plane-side fraction guard as the backstop that does not depend on clients
Passive signal is unavailable for low-traffic services, where probing remains the only evidence Grey failure stays invisible exactly where there is least redundancy to absorb it Per-service evidence weighting with a declared traffic threshold below which probe cadence is raised, and an explicit report of which services have single-class evidence
The xDS tier is stateful and holds 60,000 long-lived connections, so its own deploy is the platform's most dangerous routine operation A propagation-tier deploy becomes a company-wide reconnection storm Jittered re-subscription with a 5%-per-10-s ceiling, connection draining on tier rollout, and the static window as the fallback that makes the storm survivable
Regions will disagree about which region is healthy, and the disagreement arrives at the worst moment Two regions each route away from the other, or each declares itself the survivor Regional authority with read-only cross-region aggregates, an out-of-band failover signal independent of the registry, and reconciliation after partition that is explicit and logged rather than derived
A compromised caller can lie about someone else's health through the passive outcome channel A workload with no privileges degrades a service it merely calls Per-client ejection caps, outcome reports weighted by reporter population rather than volume, and attribution at ingest so the lying reporter is identifiable after the fact

The reasoning behind every component and technology choice is in the Architecture Decision Record: 16 records across 6 areas, each with the alternatives that lost and what the choice costs.