Health Check & Service Discovery

Solution Architecture v1.0 · Amazon Web Services with an open-source Envoy/xDS data plane · Reliability Architecture · 2026-10 · 21 views · 16 architecture decision records

When a team chat shows a spinner on send, when a shared document says "reconnecting", when the driver freezes on the map for eight seconds, the user is almost never looking at a crashed service. They are looking at a request routed to a replica that should not have been in rotation — one restarting, one whose connection pool was exhausted, one in a zone already declared unhealthy whose address was still cached in the caller. The opposite failure is less visible and more expensive: a bad health-check config, or a shared database wobbling, and a fleet takes itself out of rotation faster than it can be put back. This package is the internal health and discovery plane that answers two questions continuously for every service-to-service call in an assumed collaboration SaaS: 450 internal services, 60,000 instances, 3 active regions, 12 million internal requests a second at peak, 2,500 instance state changes a minute at steady state and 40,000 during a regional evacuation — on EKS, ECS and EC2 Auto Scaling for registration, Cloud Map and Route 53 for the DNS surface, an Envoy/xDS tier for propagation, Aurora for desired state, DynamoDB for observed state, and ARC for the out-of-band region signal.

21 views 21 HTML views21 SVG21 draw.io 2 documents Updated 2026-10-04
Architecture views

21 views, each in three formats.

Open a view to read it in full. Every SVG carries its diagram source inside it, so it opens in diagrams.net fully editable with no import step; the draw.io files are the same diagrams as plain source.

  1. 01
    System Context

    Three registration sources, one traffic path, and a list of things this plane is not permitted to need.

  2. 02
    High-Level Architecture

    Five stages, and a seam between the fourth and the fifth that the rest of the set is about.

  3. 03
    Actors and Their Journeys

    Four humans, three machines, and the one actor who can break this platform by doing their job correctly.

  4. 04
    Journey — Deploy Without Dropping Traffic

    The trough is not the drain. It is the moment the new replica gets a full share of traffic before it is warm.

  5. 05
    Journey — Chase Errors From One Replica

    Every probe is green and 4% of requests are failing. This is the journey the passive signal class exists for.

  6. 06
    Journey — Evacuate a Zone, and Bring It Back

    A 16× churn burst, a decision made under pressure, and a way back that has to be one reversible step.

  7. 07
    Layered Architecture

    Seven layers. Six of them are allowed to be unavailable for an hour.

  8. 08
    Platform Components

    Four planes drawn as two boxes, because the control plane fails as one thing and the data plane survives it.

  9. 09
    Interface Catalogue

    Four interfaces in, four out, and exactly one contract in each direction.

  10. 10
    Data Flow — Signal To Routing Change

    One probe result, one readiness report or one failed request, followed all the way to a changed endpoint set.

  11. 11
    State Classes

    Three classes ordered by what happens if the data is lost. Only one of them has an RPO worth arguing about.

  12. 12
    Registry Data Model

    Eight durable entities. What is missing from this view is the point of it.

  13. 13
    Critical Flow — A Replica Fails

    Fourteen messages, and the two of them that are the reason a user never sees this.

  14. 14
    Registration To Eligible

    Four runtimes with four different amounts of evidence, and one row that is identical for all of them.

  15. 15
    Eligibility Decision Path

    Four kinds of evidence, five verdicts, and two of the verdicts keep the instance in rotation.

  16. 16
    Deployment Architecture

    Three regions of the same shape, none of them depending on a global component to route traffic.

  17. 17
    Observability

    Six signal families across five stages, reduced to four alarms that map to four distinct actions.

  18. 18
    Instance Lifecycle

    Eight states, and no edge that returns an instance straight to Serving.

  19. 19
    Trust Zones

    Four zones, and the one control the whole design rests on sits at the ingest boundary.

  20. 20
    Identity and Authorisation Flow

    A privileged action the platform is allowed to refuse, and why the refusal is the control.

  21. 21
    Failure Classes

    Twelve classes, what catches each one, and what it costs when the bound fails.

Documents

The written architecture, on the page.

The view index above is the map; this is the argument. The one-pager and the decision record are part of the deliverable, so they are printed here in full — each also opens as its own page with a table of contents.

Document 1 of 2 · 14 min read

Architecture One-Pager

Health Check & Service Discovery · Solution Architecture v1.0 · Amazon Web Services with an open-source Envoy/xDS data plane · Reliability Architecture · 2026-10 · 21 views · 16 architecture decision records

The client's last-known-good view is authoritative for routing; the control plane is an advisory service that improves that view and is never in the request path.

When a team chat shows a spinner on send, when a shared document says "reconnecting", when the driver freezes on the map for eight seconds, the user is rarely looking at a crashed service. They are looking at a request routed to a replica that should not have been in rotation — one restarting, one whose connection pool was exhausted, one in a zone already declared unhealthy whose address was still cached in the caller. The opposite failure is less visible and more expensive: a bad health-check config, or a shared database wobbling, and a fleet takes itself out of rotation faster than it can be put back. Something has to answer two questions continuously for every one of twelve million internal calls a second — which instances of this service exist, and which should receive traffic — and that something is itself a distributed system that will have a bad day. The problem is not knowing which replicas are healthy. It is being allowed to be wrong about it, briefly and boundedly, without that being the same thing as an outage.

Reconcile instance existence from the orchestrator's own state rather than asking workloads to register themselves, and collect health as three independent evidence classes: active probes from outside, self-reports from inside, and the outcomes real callers actually observed. Fuse them into one eligibility decision per instance over a decaying score with asymmetric hysteresis, where silence is Unknown rather than unhealthy and a degraded shared dependency reduces weight rather than removing capacity. Assemble per-service versioned views with explicit topology priority tiers, and push incremental deltas to sixty thousand streaming subscriptions, budgeting withdrawal at five seconds and admission at thirty. Every client keeps the last view it received on local disk and routes from it — through a control-plane outage, for at least an hour, with no behaviour change but staleness. Removal is bounded in three independent places: a minimum healthy fraction that disregards health rather than draining a service, a cap on how fast a client's endpoint set may shrink, and a cap on how much of it any one client may eject. Recovery is treated as the dangerous event: slow start, jittered re-subscription and rate-capped shift-back are defaults.

What it is, and what it is not

  • A client cache that is authoritative for routing — not A registry consulted, or leased from, on the request path
  • Fail static on missing information — not Fail closed, which turns a blind observer into an outage
  • Liveness, readiness and dependency health as three questions — not One health boolean per instance
  • Unknown as a first-class state — not Unknown silently treated as unhealthy
  • Withdrawal and admission on separate budgets — not One propagation path with one latency target
  • Removal bounded by a minimum healthy fraction — not Every withdrawal honoured, however many arrive at once
  • Recovery protected by default — not Recovery as the moment everything is allowed back at once
  • Regional authority with explicit reconciliation — not A global registry every region depends on to route
  • Capability declared per service — not The weakest available health signal applied to the whole estate

The decisions that are the architecture

  1. The data plane is authoritative; the control plane is advisory (ADR-01) — A view already delivered keeps working. The control plane's only job is to improve it, which lets its availability target be 99.95% while resolution as callers experience it is five nines, and turns a registry outage into a staleness problem with a bounded blast radius rather than a company-wide outage.
  2. Computed state is disposable; only desired and observed state are durable (ADR-02) — Eligibility decisions and published views are rebuilt from the registry and the sample store, never restored. That is what makes a five-minute RTO a rebuild rather than a backup, and it is why a scoring bug is repaired by recomputing rather than by surgery on live routing state.
  3. Registration is reconciled, not self-reported (ADR-03) — The orchestrator already knows which pods and tasks exist; asking the workload to announce itself adds a failure mode that looks like a healthy fleet with missing members. Explicit registration exists only for the runtimes nothing watches.
  4. Three independent evidence classes, with passive outcomes primary where traffic allows (ADR-04) — A probe measures a path no caller uses, a self-report can be confidently wrong, and the outcomes real callers saw are both free and the only class a workload cannot forge about itself. Fusing all three is what makes grey failure visible.
  5. Unknown is a state, and silence never removes an instance (ADR-05) — A system with only healthy and unhealthy drains itself the first time its observation path breaks. Removal on missing information requires corroboration from a second evidence class.
  6. Dependency health reduces weight; it may not withdraw a shared dependency (ADR-06) — If every replica shares one database, honest unreadiness turns a partially functional service into a completely unavailable one. Unreadiness from a dependency is permitted only where the dependency is declared non-shared.
  7. Asymmetric hysteresis over a decaying score, with a hard-failure fast path (ADR-07) — Re-entry is strictly slower than removal, scoring responds to failure rate rather than the last sample, and a closed port skips the decay entirely — because damping exists to absorb noise, not to argue with a refused connection.
  8. The minimum healthy fraction disregards health rather than draining a service (ADR-08) — Below 50% eligible in a service's zone the platform returns the full endpoint set and says so. A bad health-check deploy is then a declared degradation instead of a self-inflicted outage.
  9. Flapping is quarantined with escalating backoff (ADR-09) — An oscillating instance destabilises every client's view, which is worse than either verdict. Quarantine is surfaced to the owning team as an event, because a quarantined instance is a bug report about a health check.
  10. Streaming subscriptions are the primary surface; DNS is compatibility (ADR-10) — Push with exact version tracking is the only way to hold a five-second withdrawal budget across sixty thousand clients. DNS stays, because a resolver that ignores TTLs is still a client that has to be served.
  11. Withdrawal and admission are budgeted separately (ADR-11) — Five seconds to remove, thirty to add, and under saturation the addition path is the one that sheds. Adding capacity late is cheaper than adding it wrong.
  12. Slow start is published by the control plane as a rising weight (ADR-12) — Uniform across every client and every heterogeneous data plane, which matters more than being traffic-proportional in an estate where the least-updated client library decides whether a client-side ramp exists at all.
  13. Recovery is rate-capped in three places (ADR-13) — Reconnection, traffic share and regional shift-back all have ceilings, because the moment everything is allowed back at once is how a recovered target is killed a second time.
  14. Regional authority, with an out-of-band failover signal (ADR-14) — A region routes with both peers and every global component unreachable. The signal that declares a region unhealthy deliberately does not depend on the registry it is talking about.
  15. The power to deny service is a privilege, and signals must be attributable (ADR-15) — An instance may report only about itself, a prober only about its assigned targets, and an unattributable signal is discarded. Marking an instance ineligible is rate-limited and audited because it is functionally a denial of service.
  16. One registry over heterogeneous runtimes, with capability declared per service (ADR-16) — A caller gets one resolution surface whether its callee is a pod, a fleet instance or somebody else's endpoint — and a service that cannot supply a readiness signal says so, rather than the whole estate being reduced to the weakest contract.

Why this should still be right in ten years

Orchestrators, mesh data planes, discovery protocols and managed DNS will all be replaced inside a decade. The parts of this design that should outlive them are the ones that are statements about authority and blast radius rather than about products.

  • Authority is not a protocol. "The thing holding the view routes; the thing computing it advises" is a claim about where authority lives. It survives replacing xDS with whatever follows it, Envoy with another data plane, and the registry store with anything, because none of those changes which copy of the truth is allowed to route.
  • Absence of evidence outlives every monitoring stack. A system that cannot distinguish "I know it is broken" from "I cannot see it" will eventually drain a healthy fleet, whatever is doing the observing. Unknown as a first-class state costs almost nothing and removes the largest self-inflicted risk in the design.
  • Removal must be bounded in more than one place. The cost of a wrong removal scales with fleet size; the cost of a wrong retention does not. That asymmetry is arithmetic, not fashion, so a minimum healthy fraction, a shrink cap and an ejection cap will still be the right three bounds when every component here has been replaced.
  • Recovery stays more dangerous than failure. Failure sheds load and recovery concentrates it, because demand does not disappear while something is down. Slow start, jitter and rate-capped shift-back are consequences of queueing, which no technology change repeals.
  • What will date fastest. The propagation mechanism is a current answer to a permanent question, and the numbers are a reference estate rather than a measurement. Expect both to be replaced; expect the four statements above to survive it.

Non-functional targets

Every target below is a stated assumption for this exercise, chosen so that a reviewer can disagree with one and follow it to the decision that depends on it.

Quality Target How it is met View
Resolution availability, as a caller experiences it ≥ 99.999% monthly Served from a durable per-client cache; the control plane is not consulted on the request path 02
Control plane availability ≥ 99.95% monthly per region Regional authority, 18 replicas across 3 AZs, no global request-path dependency 16
Static-serving window ≥ 60 minutes, behaviour unchanged Last-known-good cache on local disk, verified by a routine production removal drill 11
Resolution latency p99 ≤ 1 ms from cache; ≤ 50 ms cold In-process lookup against the cached view; cold path only on first subscription 08
Unplanned withdrawal propagation 99% of callers in 5 s, 99.9% in 15 s Incremental delta push over streaming subscriptions, withdrawal path never shed 13
Graceful drain visibility 99% of callers in 2 s, 99.9% in 5 s Drain-to-termination interval derived from measured propagation delay 04
Admission propagation 99% of callers within 30 s Deliberately slower than withdrawal; the addition path is the one that sheds under load 10
Detection of a failing replica 3 s p50, 10 s p99 from first failed request Passive caller outcomes fused with probe and self-report; hard failures skip the decay 15
Eligibility stability ≤ 4 changes per instance per 10 min Asymmetric hysteresis, decaying score, escalating quarantine; damping adds ≤ 5 s to p99 detection 15
Blast radius of a wrong removal ≥ 50% eligible per service per zone Minimum healthy fraction disregards health; client shrink cap 25% per 10 s; per-client ejection cap 21
Churn absorption 40,000 state changes/min for 10 min Evaluation sharded by service, propagation scaled by subscription count and change rate 06
Signal throughput 1.2 M results/second Bounded prober assignment, write-sized observed-state store with a 7-day TTL 17
Recovery RTO 5 min control plane; RPO 0 desired state Computed views rebuilt from desired and observed state on a rehearsed schedule; request path does not stop 11
Cost ≤ $0.012 per instance per day; probes ≤ 0.5% of internal bytes Passive signal preferred where traffic permits; per-service cost attribution exposed to owners 17

Scope

In scope

  • Registration and non-reusable instance identity across orchestrated and non-orchestrated runtimes, with leases
  • Health collection as three independent evidence classes: active probe, self-report, and caller-observed outcomes
  • Eligibility evaluation with hysteresis, flap quarantine, weights and the minimum healthy fraction guard
  • Per-service versioned views with topology priority tiers, published over streaming, DNS and a query API
  • Drain, disruption budgets, deploy ramps, and zone or region evacuation as a reversible preference shift
  • Recovery behaviour: slow start, reconnection caps, rate-capped shift-back, and the static-serving window
  • Signal attribution, resolution authorisation, and the audit trail behind every eligibility transition

Explicitly out of scope

  • Load-balancing algorithms beyond the endpoint set and weights the platform hands to them
  • The service mesh's policy and mTLS planes, which consume these views rather than producing them
  • The public API edge and anything facing an external client
  • The observability platform that stores the signals this plane exports
  • Application readiness logic itself — the platform defines the contract, the service decides what passing means

What a four-week prototype should prove

The prototype's job is to falsify the three claims the whole design rests on: that routing survives the control plane being switched off, that a withdrawal reaches a realistic client population inside five seconds while the system is busy, and that a deliberately broken health check degrades a service rather than draining it. Everything else in this package is a consequence of those three holding, and nothing else is worth four weeks.

  1. One EKS cluster, 40 services, 2,000 instances, Envoy sidecars with a persisted view cache
  2. One regional control plane: reconciler, prober DaemonSet, passive outcome ingest, evaluator, xDS tier
  3. One non-orchestrated fleet behind the registration API, and one declared third-party endpoint with no readiness signal
  4. Synthetic caller traffic with per-endpoint outcome reporting, and a load generator able to produce a 16× churn burst
  • Kill the control plane for 60 minutes under steady traffic: error rate must not move, and a restarted sidecar must route from disk without blocking
  • Fail one replica hard, and separately make one replica slow but alive: measure time to 99% of callers for each, and show the probe-green case is still caught
  • Push a health-check configuration that fails every replica of one service: the service must degrade with a declared 'health disregarded' view and an alert, not drain
  • Evacuate one zone during a 40,000-changes-per-minute burst: the withdrawal budget must hold while the addition path visibly sheds
  • Restart the xDS tier with 2,000 subscriptions attached: reconnection must stay under the 5%-per-10-s ceiling and no client view may regress
  • Partition the prober network from one AZ: those instances must become unknown and stay in rotation, not be withdrawn

Open risks, carried rather than hidden

Risk If it lands Response
The 60-minute static window is a claim nobody tests, so the first real control-plane outage is also the first test The design's central promise fails exactly when it is needed, and the blast radius is every service at once Routine production removal of the control plane from a sample of traffic, with static-serving minutes as a reported signal rather than an incident-only metric
Client library heterogeneity means the client-side guards do not actually exist everywhere Shrink caps and ejection caps are assumed by the control plane and absent in the oldest 5% of callers, which are the ones that will drain a service Report library version spread as a first-class signal, gate onboarding on a minimum version, and keep the control-plane-side fraction guard as the backstop that does not depend on clients
Passive signal is unavailable for low-traffic services, where probing remains the only evidence Grey failure stays invisible exactly where there is least redundancy to absorb it Per-service evidence weighting with a declared traffic threshold below which probe cadence is raised, and an explicit report of which services have single-class evidence
The xDS tier is stateful and holds 60,000 long-lived connections, so its own deploy is the platform's most dangerous routine operation A propagation-tier deploy becomes a company-wide reconnection storm Jittered re-subscription with a 5%-per-10-s ceiling, connection draining on tier rollout, and the static window as the fallback that makes the storm survivable
Regions will disagree about which region is healthy, and the disagreement arrives at the worst moment Two regions each route away from the other, or each declares itself the survivor Regional authority with read-only cross-region aggregates, an out-of-band failover signal independent of the registry, and reconciliation after partition that is explicit and logged rather than derived
A compromised caller can lie about someone else's health through the passive outcome channel A workload with no privileges degrades a service it merely calls Per-client ejection caps, outcome reports weighted by reporter population rather than volume, and attribution at ingest so the lying reporter is identifiable after the fact

The reasoning behind every component and technology choice is in the Architecture Decision Record: 16 records across 6 areas, each with the alternatives that lost and what the choice costs.

Document 2 of 2 · 60 min read

Architecture Decision Record

Health Check & Service Discovery · Solution Architecture v1.0 · Amazon Web Services with an open-source Envoy/xDS data plane · Reliability Architecture · 2026-10 · 21 views · 16 architecture decision records

The argument these decisions serve is summarised in the Architecture One-Pager.

Sixteen decisions make up this architecture. Everything else across the twenty-one views is a consequence of one of them, and each record carries the question that forced it, how it is realised on AWS, the alternatives including the ones that are right for a different organisation, and the conditions under which the choice should be reversed.

Status of this document. This is a design, not a report on a running system. Every rate, latency, volume, retention and cost figure in this package is a stated assumption, chosen to be defensible for a widely used collaboration SaaS at roughly 25 million daily active users and sized from a reference estate of 450 internal services, 60,000 instances, 3 active regions and 12 million internal service-to-service requests a second at peak. They are written as numbers so that they can be argued with and corrected, which vagueness does not allow.

How to read a record

  • Question: The forcing question: why a decision was needed at all.
  • Context: The requirement, the scale and the constraint that make it hard.
  • Decision: What this architecture does, stated so it can be checked.
  • How it is realised on AWS: The concrete mechanism: which service or package, configured how, in which subscription.
  • Options weighed: Chosen, rejected, deferred, or right elsewhere, with the reason for each.
  • Consequences: What the choice buys and what it costs, both kept visible.
  • Choose differently when: The conditions that would flip the decision for your system.
  • Why it holds up over time: What keeps the decision right as scale, staff and technology change.
  • Lesson: The principle that transfers beyond this platform.

Decision map

Authority and blast radius: Where routing authority lives, what is durable, and how wrong the platform is allowed to be.

  • ADR-01 · The client's last-known-good view is authoritative for routing; the control plane is advisory
  • ADR-02 · Computed state is disposable; only desired and observed state are durable

Evidence and what counts as health: How the platform learns an instance is in trouble, and what it refuses to infer from silence.

  • ADR-03 · Registration is reconciled from the orchestrator, not self-reported by the workload
  • ADR-04 · Three independent evidence classes, with passive caller outcomes primary where traffic allows
  • ADR-05 · Unknown is a first-class state, and silence alone never removes an instance
  • ADR-06 · Dependency health reduces weight; it may not withdraw an instance over a shared dependency

Decision stability: Turning a stream of noisy samples into a verdict that does not oscillate, and bounding what a wrong verdict costs.

  • ADR-07 · Asymmetric hysteresis over a decaying score, with a fast path for hard failures
  • ADR-08 · The minimum healthy fraction disregards health rather than draining a service
  • ADR-09 · A flapping instance is quarantined with escalating backoff, and its owner is told

Propagation and convergence: Getting a changed view to sixty thousand clients, and the asymmetry between removing and adding.

  • ADR-10 · Streaming subscriptions are the primary surface; DNS is a compatibility surface
  • ADR-11 · Withdrawal and admission are separately budgeted, and admission is the path that sheds

Recovery: The decisions that treat coming back as more dangerous than going down.

  • ADR-12 · Slow start is published by the control plane as a rising weight
  • ADR-13 · Recovery is rate-capped in three independent places

Trust and heterogeneity: Who is allowed to deny service to whom, and how runtimes that cannot answer the question are represented.

  • ADR-14 · Regional registry authority, with a failover signal that does not depend on the registry
  • ADR-15 · The power to deny service is a privilege, and every signal must be attributable
  • ADR-16 · One registry over heterogeneous runtimes, with capability declared per service

Technology by capability

Amazon Web Services was chosen for this exercise deliberately. It is tied for the least-used stack across this practice's existing use cases and has not been used since 2026-09, while the two most recent use cases both ran on Google Cloud — the rotation this library exists for. It also happens to fit the topic honestly: health checking and discovery here is a multi-runtime problem, and AWS supplies genuinely distinct primitives for containers, tasks, VM fleets, managed functions and the DNS surface, which is the heterogeneity the design has to absorb rather than wish away. The data plane is deliberately open source, because the one component that must be identical across every runtime is the one the platform ships into every workload.

Capability Choice Origin Credible alternative Why this one Record
Orchestrated instance state EKS EndpointSlice watch across 12 clusters, plus ECS task state change events Amazon Web Services Self-registration from each workload; Consul agents The orchestrator already holds reconciled desired state, so a missing instance is a diff rather than a silence ADR-03
Non-orchestrated instance state EC2 Auto Scaling lifecycle hooks plus an authenticated registration API with leases Amazon Web Services Agent-based self-registration everywhere Covers fleets nothing watches without extending self-registration to the orchestrated majority ADR-16
Desired state store Aurora PostgreSQL, 3-AZ, one writer per region Amazon Web Services DynamoDB global tables; etcd; Consul Small, exact, relational state with strong in-region consistency; RPO 0 matters more than write scale here ADR-02
Observed state store DynamoDB on-demand with a 7-day TTL Amazon Web Services Timestream; a Prometheus-compatible TSDB; Cassandra Sized for 1.2 M writes/second where losing a window of samples is acceptable by design ADR-02
Health signal ingest Kinesis Data Streams, sharded by service Amazon Web Services MSK; SQS; direct gRPC to the evaluator Shards give per-service isolation so one service's churn cannot delay another's evaluation ADR-04
Active probing Per-AZ DaemonSet probers with bounded target assignment Open source on EKS Route 53 health checks for everything; a central prober fleet Sub-linear probe cost in instance count, and a prober failure scoped to one AZ rather than a region ADR-04
Passive outcome reporting Envoy per-endpoint outcome counters reported to ingest Open source (Envoy) Application-level reporting; inference from service-mesh telemetry The only evidence class that measures what callers experience and that a workload cannot forge about itself ADR-04
Eligibility evaluation Stateless evaluator on EKS, sharded by service, state in memory Open source on EKS Stream-processing framework; evaluation inside the proxies Computed state is disposable, so the evaluator can be restarted and rebuilt rather than recovered ADR-02
View propagation Envoy-compatible xDS tier with incremental deltas and per-service versions Open source (xDS) on EKS Consul; short-TTL DNS only; gossip; long-poll The only mechanism whose propagation latency the platform owns, which a 5 s withdrawal budget requires ADR-10
DNS compatibility surface Cloud Map plus Route 53 private zones on a short TTL Amazon Web Services CoreDNS only; no DNS surface at all Serves clients that cannot hold a subscription, with the weaker guarantee stated rather than hidden ADR-10
Client data plane Envoy sidecar with a durable last-known-good cache, plus a thin resolver library Open source (Envoy) Client-side load balancing per language; a central routing proxy tier Routing authority has to live where the request is, and it has to be identical in every runtime ADR-01
Runtime edge membership NLB and ALB target groups synchronised from the published views Amazon Web Services Target group health checks as the source of truth One eligibility semantics for edge-routed and sidecar-routed traffic, with no second opinion ADR-01
Region health and failover signal Route 53 Application Recovery Controller readiness checks and routing controls Amazon Web Services Registry-derived region health; a custom quorum service A region's reachability must not be asserted by the system whose reachability is in question ADR-14
Workload identity IRSA for pods, instance profiles for fleets, mutual TLS on every control-plane path Amazon Web Services SPIFFE/SPIRE; network-position trust Self-only health reporting is only enforceable if every signal carries an identity the platform can check ADR-15
Audit and transition record Kinesis to S3 with Object Lock, 5-year retention Amazon Web Services CloudTrail only; the relational store Denials of service must be reconstructable long after the samples that caused them have expired ADR-15
Policy and contract authoring Policy API with staged rollout, disruption budgets and one-step reversal Open source on EKS GitOps-applied CRDs; direct registry writes Health-check configuration can drain a fleet, so it needs the same gates as a code deploy ADR-08

The decisions, and the alternatives that lost

Authority and blast radius

Where routing authority lives, what is durable, and how wrong the platform is allowed to be.

ADR-01 · The client's last-known-good view is authoritative for routing; the control plane is advisory

Status: Accepted · Shown on views: 02, 08, 11

When a caller needs an endpoint and the discovery system is unreachable, what is the correct behaviour?

Context. The obvious design makes the registry the single truth and has callers consult it, or hold a short lease from it, per resolution. It gives one answer that is always current, and it means every one of twelve million internal requests a second depends on one distributed system being available. That system's availability then becomes the ceiling on every product's availability, and its worst day is everybody's worst day. The alternative is to accept that routing runs on a view that may be stale, and to make staleness the normal operating condition rather than an error — which moves the hard problems into the client and makes propagation delay the number the design lives or dies by.

Decision. The view a client already holds is authoritative for routing. The control plane computes and publishes improvements to that view and is never consulted synchronously on a request. Every client persists its last-known-good view locally and continues routing from it, unchanged, through a total control-plane outage for at least 60 minutes, degrading only by staleness. A resolution that cannot be served from cache or stream fails to a documented default; it never blocks.

How it is realised on AWS. Envoy sidecars hold incremental xDS subscriptions against the regional control plane on EKS and write each applied view to a local volume. On stream loss the sidecar keeps serving the persisted view and re-subscribes with jittered backoff. The thin resolver library used by sidecar-less callers implements the same cache-first contract. NLB and ALB target groups are synchronised from the same views, so edge-routed traffic inherits the behaviour without taking a dependency of its own.

Option Verdict Reasoning
Client-held last-known-good view is authoritative; control plane advisory Chosen Decouples caller availability from control-plane availability entirely; pays in staleness and in guards that must exist twice.
Registry consulted or leased per resolution Rejected Exactly consistent and makes the discovery plane a hard dependency of every request in the company.
Short-TTL DNS only, no persistence Rejected Looks like caching but is not: a resolver that honours a 5 s TTL has no view at all 5 s into an outage.
Central routing proxies holding the view on the caller's behalf Right elsewhere Right where client libraries cannot be changed; adds a network hop and a new tier that must itself be more available than the control plane.

What it buys

  • The control plane can be deployed, patched, scaled and drilled like an ordinary service, because its unavailability is a staleness event rather than an outage.
  • A registry failure, a propagation-tier deploy and a cross-region partition all collapse into the same bounded failure mode.
  • Resolution latency is an in-process lookup at p99 ≤ 1 ms, which removes discovery from every latency budget in the estate.

What it costs

  • Every safety mechanism that bounds a wrong decision must exist in the client as well as the control plane — the shrink cap, the ejection cap, the refusal of an empty set.
  • Correctness now depends on the least-updated client library in the estate, which is a fleet-management problem rather than an architecture one.
  • Withdrawal propagation delay becomes the single most important number in the design, and the only honest measure of how wrong the data plane is permitted to be.

Choose differently when. If the estate were small enough that one control plane could credibly be more available than every service calling it, or if routing to a withdrawn endpoint were unrecoverable rather than retryable — a payment captured twice, an irreversible write — then exact consistency is worth the dependency and the registry should be the authority. The decision should also be revisited if client library heterogeneity becomes unmanageable, because this design silently delegates safety to those libraries.

Why it holds up over time. This is a statement about where authority lives, not about a protocol. It survives replacing xDS, Envoy, the orchestrator and the registry store, because none of those changes which copy of the truth is allowed to route.

Lesson. Decide early whether a consumer of your system is allowed to be briefly wrong or must always be right. If briefly wrong is acceptable, put the truth in the consumer and make your own availability somebody else's non-event.

ADR-02 · Computed state is disposable; only desired and observed state are durable

Status: Accepted · Shown on views: 11, 12, 16

Which of the platform's own data has to survive a total loss of the control plane, and which should be thrown away rather than restored?

Context. A discovery plane holds three quite different kinds of data and the temptation is to treat them alike. Service definitions, instance rows, health contracts and policy are small, exact and irreplaceable: nothing can recompute who owns a service or what its readiness contract says. Probe results and caller outcomes are enormous, arrive at over a million a second, and matter only in aggregate over the last few minutes. Eligibility decisions and published views are derived from the other two by a pure function. Storing the third class durably invites repairing it, and a platform that repairs its routing state in place has no recovery story at all.

Decision. Desired state is the system of record with RPO 0, held in a replicated store that is strongly consistent within a region. Observed state is recomputable with RPO 60 s, held in a store sized for write volume rather than durability, where losing a window of samples reduces confidence without corrupting a verdict. Computed state — eligibility, weights and versioned views — is held in memory only, has no RPO, and is rebuilt rather than restored. Every eligibility transition and policy change is additionally emitted to a durable append-only log for audit and incident reconstruction.

How it is realised on AWS. Registry and policy on Aurora PostgreSQL across three AZs. Observed state in DynamoDB with a 7-day TTL and on-demand capacity sized for 1.2 M writes a second. Eligibility and assembled views live in the control-plane pods' memory and are reconstructed on start from Aurora plus the last few minutes of DynamoDB samples. Transitions stream to Kinesis and land in S3 with Object Lock for the 5-year audit retention.

Option Verdict Reasoning
Three state classes with three durability contracts Chosen Makes RTO a rebuild throughput question and removes any temptation to repair routing state.
One durable store for everything Rejected Simplest to operate and makes the sample write rate the sizing constraint for the store holding the service catalogue.
Persist computed eligibility for faster restart Rejected Buys seconds of startup time and reintroduces repairable routing state, which is the failure mode being designed out.
Event-sourced control plane with a replayable decision log Deferred Attractive for reconstructing exactly what the platform believed at a past instant; the transition log gives most of it without the operational weight.

What it buys

  • RTO 5 minutes is a rebuild measurement rather than a restore hope, and the rebuild is exercised on a schedule instead of during an incident.
  • A scoring or fusion bug is repaired by recomputing from retained samples, with no state surgery.
  • The sample store can be cheap and lossy, which is what makes a million writes a second affordable.

What it costs

  • The rebuild path must be fast enough to meet the RTO at full scale, which is a continuing throughput commitment and not a one-off test.
  • Losing up to 60 s of samples briefly reduces confidence in eligibility, which interacts with the unknown state and has to be reasoned about together.
  • Two stores with different consistency models mean two operational runbooks.

Choose differently when. At a tenth of this scale the sample volume no longer forces a separate store, and one strongly consistent store for everything is simpler and defensible. Persisting computed state becomes right if control-plane restarts ever need to be measured in seconds rather than minutes — for example if the platform were ever placed on the request path, which ADR-01 forbids.

Why it holds up over time. The classification by what-happens-if-lost outlives any particular store. Swapping Aurora, DynamoDB or Kinesis changes nothing about which data may be thrown away.

Lesson. Sort your state by what happens if it disappears, not by which component produced it. Most of what a control plane holds should be unrecoverable-by-design and recomputed, because that is what makes recovery routine.

Evidence and what counts as health

How the platform learns an instance is in trouble, and what it refuses to infer from silence.

ADR-03 · Registration is reconciled from the orchestrator, not self-reported by the workload

Status: Accepted · Shown on views: 08, 14, 12

How does the platform learn that an instance exists?

Context. The classical service-discovery pattern has each process register itself on start and deregister on stop. It is simple, it works in a demo, and it has two failure modes that look nothing alike. A process that crashes before deregistering leaves a zombie endpoint that is only removed when its lease expires. A process that starts but fails to register is worse: the fleet looks healthy and is short of members, and nothing in the system is in a position to notice, because the only thing that knew the instance should exist was the instance. Meanwhile the orchestrator already holds exactly this information, as reconciled desired state, for the majority of the estate.

Decision. For orchestrated workloads, instance existence is derived by continuous reconciliation against the orchestrator's own state; the workload does not register itself. An authenticated registration API exists only for runtimes nothing watches. Every instance receives a registry-unique identity that is stable for its lifetime and never reused, so a restarted process is a new instance rather than a continuation. Ungraceful termination is caught by lease expiry within 30 s; graceful termination deregisters before the process stops accepting connections.

How it is realised on AWS. A reconciler watches EndpointSlices across 12 EKS clusters and ECS task state change events, and consumes EC2 Auto Scaling lifecycle hooks for VM fleets. Divergence between orchestrator state and registry state is reported as a reconciliation diff, which is also how registration leaks are found. The registration API is reserved for declared external endpoints and for fleets outside Auto Scaling, authenticated by instance profile or IRSA.

Option Verdict Reasoning
Reconcile from orchestrator state; explicit registration for the remainder Chosen Makes a missing instance detectable as a diff rather than invisible, and removes a whole class of startup failure.
Self-registration by every workload Rejected One mechanism for the whole estate, at the cost of a failure mode where the platform cannot know what it has not been told.
Sidecar registers on the application's behalf Rejected Moves the problem into the sidecar without removing it; the sidecar can also fail to register.
Orchestrator-native discovery only, no separate registry Right elsewhere Right for a single-cluster estate; it cannot represent a VM fleet, a managed function or a third-party endpoint, which is most of the hard part here.

What it buys

  • A missing instance appears as a reconciliation diff instead of as a quietly smaller fleet.
  • Non-reusable identity makes a stale cached endpoint detectably stale rather than ambiguously reassigned.
  • The registration API's attack surface shrinks to the small non-orchestrated tail.

What it costs

  • One watch integration per runtime type, each with its own failure modes and its own back-pressure behaviour.
  • Reconciliation lag becomes a signal that has to be watched, because it is now the path by which existence is learned.
  • A heterogeneous estate ends up with two registration mechanisms, which has to be documented so callers do not care.

Choose differently when. If the entire estate ran on one orchestrator with one API, reconciliation and native discovery converge and a separate registry earns less. Self-registration becomes the better choice if workloads routinely run somewhere nothing watches — a developer laptop, a customer-managed environment — where there is no authority to reconcile against.

Why it holds up over time. The principle is to prefer the authority that already knows over the party that would have to be asked. That holds as orchestrators change.

Lesson. If some other system already holds the fact you need, read it rather than asking the subject to tell you. Self-reported existence cannot distinguish 'not there' from 'never spoke up'.

ADR-04 · Three independent evidence classes, with passive caller outcomes primary where traffic allows

Status: Accepted · Shown on views: 10, 15, 05

Who should be believed about whether an instance can serve: the instance, a prober, or the callers?

Context. Each source is wrong in a different and characteristic way. A self-report knows the most about internal state and a broken process can report healthy with complete confidence, while a compromised one can simply lie. An external prober is independent, costs N×M traffic at sixty thousand instances, and measures a path no real caller uses — a replica can be perfectly healthy for the prober and broken for everyone else. The outcomes real callers observed are the only evidence that measures what actually matters, cost nothing because the traffic was happening anyway, and go silent on a low-traffic service and cannot distinguish a bad callee from a bad caller. Grey failure — brownout, exhausted pool, long GC pause — is assumed here to be the most common real failure, and it is precisely the class a probe cannot see.

Decision. The platform collects three separately addressable evidence classes — active probe, self-report, and passive caller-observed outcomes — and fuses them into one eligibility decision. The absence of one class is never read as a negative result from another. Where a service's traffic volume makes passive signal statistically sufficient, passive outcomes are the primary evidence and probing drops to a confirmation cadence; below that threshold probing remains primary and the service is reported as having single-class evidence.

How it is realised on AWS. Probers run as a per-AZ DaemonSet with bounded target assignment, so no instance is probed by every prober. Envoy sidecars report per-endpoint outcome counters to a Kinesis-backed ingest at 1.2 M results a second combined. Self-reports ride the Instance API under the workload's own identity. Fusion weights are configured per service, with the traffic threshold for passive-primary declared in the health contract.

Option Verdict Reasoning
All three classes fused, passive-primary above a traffic threshold Chosen The only arrangement that sees grey failure, survives a prober-network partition, and still covers low-traffic services.
Active probing only Rejected Universal and uniform, and blind to exactly the failure mode that is most common and most user-visible.
Self-report only Rejected Cheapest and most informed, and unfalsifiable — a confidently wrong process is indistinguishable from a healthy one.
Passive observation only Rejected Free and accurate where traffic exists; a quiet service becomes permanently unknown, which is not an answer.
Caller-side health checking with no central evaluation Right elsewhere Right for a small estate with homogeneous clients; at 450 services it means 450 differing opinions and no shared record.

What it buys

  • Grey failure is detectable, which is the difference between this design and one that only finds dead processes.
  • A prober-network partition degrades confidence instead of draining a healthy fleet, because a second class still speaks.
  • Probe traffic stays under 0.5% of internal bytes, because the expensive class is used where the free one cannot reach.

What it costs

  • Three ingest paths, three failure modes, and a fusion function whose weights are another thing to get wrong.
  • Per-service weighting is a configuration surface that most teams will leave at its default, so the default has to be right.
  • Passive signal requires every client library to report outcomes, which re-couples correctness to client fleet management.

Choose differently when. If every service in the estate carried enough traffic for passive signal to be statistically sufficient, active probing becomes a confirmation mechanism that could arguably be dropped. If client libraries cannot be trusted to report outcomes, probe-primary with self-report corroboration is the honest fallback — and grey failure then has to be caught by the callers' own error budgets instead.

Why it holds up over time. The classification of evidence by how it can be wrong is independent of any mechanism. Whatever replaces probes and sidecars, the three ways of being wrong remain three.

Lesson. Where a measurement matters, prefer evidence produced by the party that experiences the outcome over evidence produced by the party being measured or by a synthetic observer standing in for them.

ADR-05 · Unknown is a first-class state, and silence alone never removes an instance

Status: Accepted · Shown on views: 15, 10, 21

What should the platform conclude when an instance stops producing any health signal at all?

Context. A two-valued health model has to map silence onto one of its two values, and both choices are wrong in an important case. Mapping silence to unhealthy means that the first time the observation path breaks — a prober network partition, an ingest outage, a bad agent rollout — the platform concludes that the entire fleet is dead and drains a service that is working perfectly. This is not a theoretical failure: the observation path is a separate system from the thing observed, and it has its own availability. Mapping silence to healthy is the opposite error and keeps dead instances in rotation indefinitely. Neither is a statement about the instance; both are statements about the observer.

Decision. Eligibility has four values: eligible, degraded with a weight, ineligible, and unknown. An instance whose signals have gone silent because the observation path failed is classified unknown and is not removed from rotation on that basis. Moving an unknown instance to ineligible requires corroboration from a second, independent evidence class — most often the caller-observed outcomes, which do not depend on the prober network. Lease expiry remains an independent path to removal for an instance the orchestrator also no longer reports.

How it is realised on AWS. The evaluator distinguishes 'no sample within the contract interval' from 'sample reporting failure', and tracks per-class silence separately. Prober reachability is itself a monitored signal, so a prober-side outage is attributable rather than being inferred from a thousand simultaneously silent instances. The unknown count per service is exported and alertable.

Option Verdict Reasoning
Four-state eligibility with unknown as a distinct value Chosen Separates a statement about the instance from a statement about the observer, which a boolean cannot do.
Silence treated as unhealthy Rejected Conventional, simple, and converts any observation-path failure into a self-inflicted outage of every service at once.
Silence treated as healthy Rejected Safe against observer failure and keeps dead endpoints in rotation until something else notices.
Silence resolved by a quorum of probers Deferred Reduces the chance of correlated observer failure without eliminating it; the second evidence class achieves more for less machinery.

What it buys

  • An ingest or prober outage degrades confidence instead of draining the estate, which is the single largest self-inflicted risk this design removes.
  • The unknown count is a direct measure of observation coverage, which turns a blind spot into a reported number.
  • Operators get a state that matches what they actually know, which makes an override decision better informed.

What it costs

  • Unknown instances keep receiving traffic, so a genuinely dead instance may stay in rotation longer than a two-state model would allow.
  • Four states multiply the paths through the evaluator and through every client that interprets a view.
  • The design depends on a second evidence class existing — which is exactly what a low-traffic service lacks.

Choose differently when. If the observation path could be made more available than the workloads it observes — probing from within the same pod, for example, with no network between them — silence becomes genuinely informative about the instance and the unknown state earns less. For a service where serving a dead endpoint is more costly than serving none at all, mapping silence to unhealthy is defensible and should be a per-service choice.

Why it holds up over time. The distinction between 'I know it is broken' and 'I cannot see it' is permanent. Any monitoring or discovery system that collapses the two will eventually drain a healthy fleet.

Lesson. Never let a system infer a fact about a subject from the failure of its own instrument. Give absence of evidence its own name, and require corroboration before acting on it.

ADR-06 · Dependency health reduces weight; it may not withdraw an instance over a shared dependency

Status: Accepted · Shown on views: 15, 21, 03

If a service's database is degraded, should the service report itself unready?

Context. Reporting unready is the honest answer and it is what most readiness probes end up doing, because wiring a dependency check into a readiness endpoint is a two-line change that feels responsible. It is correct when the dependency is per-instance: that replica cannot work, others can, route around it. It is catastrophic when the dependency is shared, which it usually is. Every replica checks the same database, every replica fails the same check, every replica withdraws within the same few seconds, and a service that could still serve the eighty per cent of its requests that never touch that database serves nothing at all. The cascade then propagates outward, because the callers of that service now have no healthy endpoints and may withdraw in turn.

Decision. Dependency health is expressible at three authority levels, declared per service in the health contract: informational, weight reduction, or unreadiness. The default is weight reduction. Unreadiness from a dependency failure is permitted only where the dependency is declared non-shared across the service's replicas. The minimum healthy fraction remains the backstop regardless of the declared level, so even a misdeclared shared dependency cannot drain a service.

How it is realised on AWS. The health contract carries a dependency_authority field and a per-dependency shared flag. Self-reports name the dependency and the observed degradation; the evaluator applies the declared authority. A contract change that raises a shared dependency to unreadiness authority is rejected at the policy API rather than at runtime.

Option Verdict Reasoning
Weight reduction by default, unreadiness only for declared non-shared dependencies Chosen Keeps partial function available, and makes the dangerous configuration something a team has to assert rather than inherit.
Dependency failure always sets unreadiness Rejected Honest per instance and a correlated-withdrawal generator per service; this is the single most common way a fleet drains itself.
Dependency health informational only Rejected Safe against cascades and discards real information about which replicas are least able to serve.
Per-endpoint rather than per-instance health Deferred The right long-term answer — withdraw the routes that need the database, keep the rest — and it needs route-level health in the data plane, which is a larger change than this platform.

What it buys

  • A shared-dependency wobble degrades a service rather than removing it, which is almost always what the user would choose.
  • The dangerous configuration is opt-in and policy-checked, so the default is the safe one.
  • Partial function stays available, which matters because most requests to most services do not touch the degraded dependency.

What it costs

  • Capacity stays in rotation that will fail the subset of requests that do touch the dependency, so callers see errors they could have avoided.
  • Teams must classify their dependencies as shared or not, and will sometimes get it wrong.
  • Per-instance granularity is a compromise: the honest unit is the route, and this design does not have it.

Choose differently when. Where a dependency is genuinely per-instance — a local cache, an attached volume, a sidecar — unreadiness is correct and should be declared. Where failing a request is unrecoverable for the user rather than retryable, withdrawing is better than serving, and the authority level should be raised deliberately with the fraction guard understood as the only remaining backstop.

Why it holds up over time. The distinction between shared and per-instance dependency failure is a property of topology, not of technology, and it will decide this question for as long as replicas share anything.

Lesson. Before wiring a dependency check into a readiness probe, ask what happens when every replica fails that check at the same instant. If the answer is total unavailability, the check belongs in a weight, not in a boolean.

Decision stability

Turning a stream of noisy samples into a verdict that does not oscillate, and bounding what a wrong verdict costs.

ADR-07 · Asymmetric hysteresis over a decaying score, with a fast path for hard failures

Status: Accepted · Shown on views: 15, 18, 17

How many failures should it take to remove an instance, and how many successes to let it back?

Context. Every damping mechanism buys stability by paying detection latency, and the exchange rate is not obvious. Consecutive-failure thresholds are simple and multiply the detection budget by the threshold: a three-strike rule on a one-second probe makes a three-second detection a nine-second one. A decaying score responds to failure rate rather than to the last sample, which is closer to what matters and introduces tuning that nobody can reason about at three in the morning. Symmetric thresholds are the worst of the options: an instance that is marginal will oscillate, and an oscillating instance destabilises every client's view, which is worse than either verdict being applied consistently.

Decision. Eligibility is computed from a decaying score per instance per evidence class, with asymmetric thresholds: re-entry is strictly slower than removal. A hard failure — connection refused, process gone, lease expired — bypasses the decay entirely and removes immediately, because damping exists to absorb ambiguity and a closed port is not ambiguous. Damping is bounded: it may add no more than 5 s to the p99 detection budget for a genuine failure, and that bound is the constraint the score's parameters are fitted to.

How it is realised on AWS. The evaluator holds an exponentially weighted failure rate per instance per class, with configurable half-life, and a separate hard-failure channel that short-circuits to ineligible. Thresholds live in the health contract with platform-enforced floors. The 5 s damping budget is measured as the difference between first-failed-sample and published-withdrawal for the hard and soft paths respectively, and reported per service.

Option Verdict Reasoning
Decaying score, asymmetric thresholds, hard-failure fast path Chosen Responds to rate rather than to the last sample, and does not make an obviously dead instance wait for statistics.
Consecutive-failure thresholds only Rejected Trivial to explain and multiplies the detection budget by the threshold, which breaks the 10 s p99 commitment.
Symmetric hysteresis Rejected Simpler to configure and guarantees oscillation on marginal instances, which is the failure this is meant to prevent.
Machine-learned anomaly scoring over the sample stream Deferred Might catch failure patterns a rate misses; an unexplainable verdict in a routing decision is a worse trade than a missed subtlety.

What it buys

  • Marginal instances stop oscillating, so client views stay stable and the delta churn volume stays bounded.
  • A dead process is removed in 10 s p99 regardless of the damping configuration, because it never enters the damped path.
  • The damping cost is an explicit, measured 5 s rather than an emergent property of three thresholds multiplying.

What it costs

  • The score's half-life is a parameter that few teams will understand, so the platform default carries most of the weight.
  • A slowly-degrading instance crosses the threshold later than a consecutive-failure rule would, by design.
  • Two paths through the evaluator — damped and fast — is two code paths that must agree about everything else.

Choose differently when. If the dominant failure mode in the estate turned out to be hard rather than grey, the decaying score earns little and consecutive-failure thresholds would be the simpler, more explicable choice. If detection latency ever becomes more valuable than stability for a specific service — a payment path, say — the thresholds should be tightened for that service and the resulting flapping accepted and quarantined.

Why it holds up over time. Asymmetry between removal and re-entry is a permanent property of any system where a wrong removal is cheaper to reverse than a wrong retention is to discover. The specific scoring function is replaceable.

Lesson. Make damping an explicit, measured cost against a detection budget rather than an implicit consequence of thresholds. And never make an obviously dead thing wait for statistical confidence.

ADR-08 · The minimum healthy fraction disregards health rather than draining a service

Status: Accepted · Shown on views: 15, 21, 06

What should happen when most of a service's replicas report unhealthy at the same moment?

Context. Correlated mass withdrawal is the characteristic failure of a competent health-checking system. It arrives from a bad deploy, a shared dependency outage, or — most often — a health-check configuration change that is wrong in a way nobody noticed, because health-check config is rarely treated as production code. The platform then does exactly what it was asked to do, removes almost every replica, and the service's availability goes to zero because its health checking worked. Honouring every withdrawal is the honest behaviour and it means a two-line YAML change can take down a service more reliably than deleting its deployment.

Decision. Each service carries a minimum healthy fraction per zone, defaulting to 50%. When the eligible fraction would fall below it, the platform disregards health and returns the full endpoint set, declares in the response that it has done so, and raises it as an alert. A removal that would take the last eligible instance of a service in a region raises an incident-grade signal rather than completing silently. The guard sits downstream of evaluation, so a bug in scoring or fusion cannot bypass it.

How it is realised on AWS. The guard is evaluated in the view assembler, per service per zone, after eligibility and before delta encoding. Views assembled under the guard carry a degraded-confidence marker that clients surface in their own telemetry, and the fraction-guard trip is one of the platform's four alarms.

Option Verdict Reasoning
Disregard health below a per-service, per-zone fraction, and declare it Chosen Converts a self-inflicted outage into a declared degradation, which is nearly always the better failure.
Honour every withdrawal Rejected Maximally honest and makes the health-checking system the most effective outage generator in the estate.
A single global fraction for all services Rejected One number to reason about, and wrong for both a three-replica service and a three-thousand-replica one.
Require a human decision above a blast-radius threshold Right elsewhere Right where serving a known-bad instance is unsafe — a payment or data-mutating path — and wrong at three in the morning for a read path.
Graduated response: reduce weights before disregarding health Deferred A smoother curve and worth having; the cliff is easier to reason about during an incident, which is when it fires.

What it buys

  • A bad health-check rollout degrades a service instead of removing it, which is the most likely mass-withdrawal cause.
  • The declaration means callers and operators can tell 'the fleet is broken' from 'the health check is broken'.
  • The guard's position downstream of evaluation means it survives bugs in evaluation.

What it costs

  • Traffic is sent to instances known to be failing, which is worse than a clean failure if those instances corrupt data or hold connections.
  • A genuinely broken fleet is kept in rotation, so the user sees errors rather than a fast failure.
  • 50% is a cliff, and a service sitting near it will trip and untrip in ways that are confusing to watch.

Choose differently when. For a service where a failing request is unrecoverable for the user or leaves inconsistent state, draining is better than serving and the fraction should be lowered or the guard replaced by a human gate. If the platform could reliably distinguish a broken fleet from a broken health check, it could choose differently in each case and the blunt guard would no longer be needed.

Why it holds up over time. The principle — no single signal may remove a service's capacity — outlives any particular threshold, and applies to every control system with an actuator larger than its sensor's reliability.

Lesson. Bound what your control loop is allowed to do, not just what it is likely to decide. A correct system executing a wrong instruction is the failure mode that most often reaches users.

ADR-09 · A flapping instance is quarantined with escalating backoff, and its owner is told

Status: Accepted · Shown on views: 15, 18, 17

What should happen to an instance that keeps changing its mind about whether it is healthy?

Context. Hysteresis reduces oscillation; it does not eliminate it. An instance that fails under load, recovers when traffic is withdrawn, and fails again when traffic returns will oscillate indefinitely, and each transition is a delta pushed to every subscribed client for that service. At sixty thousand subscriptions, a handful of flapping instances across the estate is a measurable fraction of propagation capacity spent on noise, and the churn makes every client's view less stable than either verdict would be. The instance is also, almost always, a bug report: either its health check is wrong or it genuinely cannot sustain its share, and both are facts its owner needs.

Decision. An instance that changes eligibility more than four times in ten minutes is quarantined: held ineligible for an exponentially increasing period, starting at one minute and capped at thirty. Quarantine is surfaced to the owning team as an event with the transition history that produced it, and appears in the service's flap report. Quarantine never applies to a whole service: the fraction guard takes precedence, so quarantining cannot be a path to draining a fleet.

How it is realised on AWS. The evaluator counts transitions per instance over a sliding ten-minute window and writes a quarantine_until timestamp on the eligibility row. Quarantine events are published to the owning team's notification channel and aggregated into a per-service flap rate, which is a reported signal rather than an alarm.

Option Verdict Reasoning
Escalating quarantine with owner notification Chosen Stops the churn, and treats the flap as the defect report it usually is rather than hiding it.
Longer hysteresis instead of quarantine Rejected Suppresses the symptom across the board and pays detection latency on every instance to fix a problem a few have.
Quarantine with manual release only Rejected Safest against premature return and turns every flap into human work at whatever hour it happens.
Permanent removal after repeated flapping Right elsewhere Right where instances are cheap and immediately replaceable; here it risks removing capacity the service needs for a condition that may be load-induced.

What it buys

  • Delta churn from noisy instances is bounded, which protects propagation capacity for the changes that matter.
  • Client views stay stable, so a flapping instance stops being everyone's problem.
  • The owning team receives a specific, evidenced bug report instead of a vague reliability complaint.

What it costs

  • A genuinely recovered instance can be held out for up to thirty minutes, which is lost capacity.
  • Quarantine can mask a load-induced failure as an instance defect, sending the owner looking in the wrong place.
  • Another piece of per-instance state to hold, report and reason about during an incident.

Choose differently when. If capacity were tight enough that holding a recovered instance out for minutes mattered more than view stability, the cap should fall and the escalation flatten. In an estate where instances are replaced in seconds, terminating the flapping instance and letting the orchestrator supply a new one is simpler and better.

Why it holds up over time. Quarantine of an oscillating participant is a standard answer in any system that aggregates noisy votes into a shared decision, and it will outlive this platform's mechanisms.

Lesson. When a component's instability becomes everyone else's instability, isolate it and raise it as a defect. Suppressing oscillation globally to accommodate a few participants taxes all of them.

Propagation and convergence

Getting a changed view to sixty thousand clients, and the asymmetry between removing and adding.

ADR-10 · Streaming subscriptions are the primary surface; DNS is a compatibility surface

Status: Accepted · Shown on views: 09, 08, 13

How does a changed view reach sixty thousand clients inside five seconds?

Context. Three mechanisms are available and they are not close substitutes. Short-TTL DNS is universally consumable, needs no client library, and is bounded by resolver behaviour nobody controls — including resolvers, language runtimes and sidecars that cache beyond the TTL or ignore it entirely. A five-second withdrawal budget cannot be held by a mechanism whose propagation time is a property of other people's software. Streaming subscriptions give sub-second push with exact version tracking, and require a stateful tier holding sixty thousand long-lived connections, which must survive its own deploys and becomes the most delicate routine operation in the platform. Gossip removes the central tier and makes convergence time and message ordering probabilistic, which is difficult to reason about precisely when reasoning matters most.

Decision. Incremental streaming subscriptions are the primary propagation surface, carrying deltas with monotonically increasing view versions. A DNS zone is maintained as a compatibility surface with identical eligibility semantics, for clients that cannot hold a subscription, and its weaker propagation guarantee is stated rather than hidden. A request/response resolution API exists for tooling and target-group synchronisation and is not for request-path use. Gossip is not used.

How it is realised on AWS. An Envoy-compatible xDS tier on EKS holds the subscriptions, sending incremental resource updates with a version per service and falling back to a full view only when a client's version is too old to reconcile. Cloud Map and Route 53 are populated from the same assembled views on a short TTL. Clients re-subscribe with jitter capped at 5% of a population per 10 s window, and the tier drains connections on its own rollout.

Option Verdict Reasoning
Streaming primary, DNS as compatibility, API for tooling Chosen The only option that holds a 5 s budget while still serving clients that cannot subscribe.
DNS everywhere with aggressive TTLs Rejected No client library to ship, and the propagation guarantee belongs to resolvers the platform does not operate.
Gossip between data planes Rejected Removes the stateful tier, and replaces an explainable latency with a probabilistic one and no version ordering.
Long-poll or periodic full-view fetch Rejected Simple and stateless on the server; at 60,000 clients and a 5 s budget it is a continuous full-view fan-out.
Gossip within a zone, streaming across zones Right elsewhere Attractive where zone-local convergence dominates and the client fleet is homogeneous enough to trust with a membership protocol.

What it buys

  • Exact per-client version tracking, which is what makes monotonic application and incremental deltas possible at all.
  • Withdrawal propagation is a property of the platform rather than of every client's resolver.
  • DNS clients are still served, with a stated weaker guarantee rather than a silent one.

What it costs

  • A stateful tier holding 60,000 connections is the platform's most dangerous routine deploy, mitigated by jitter and the static window rather than removed.
  • A client library is now required for the primary path, which re-couples correctness to fleet management.
  • Two surfaces with different propagation guarantees must be kept semantically identical, which is a continuing test burden.

Choose differently when. If the client fleet were homogeneous and small enough that a full-view fetch per interval were affordable, the stateful tier is unnecessary complexity. If the withdrawal budget could be relaxed to tens of seconds — because local ejection carries the fast path entirely — DNS becomes sufficient and the whole tier can go.

Why it holds up over time. xDS is a current answer to a permanent question. What should outlive it is the requirement for versioned, incremental, monotonic delivery, which any replacement must also provide.

Lesson. Do not accept a propagation guarantee that depends on software you do not operate. If you have a latency budget, you need a mechanism whose latency you own.

ADR-11 · Withdrawal and admission are separately budgeted, and admission is the path that sheds

Status: Accepted · Shown on views: 10, 13, 06

When propagation capacity runs short, which changes get delayed?

Context. A single propagation path with a single latency target has to be sized for the worst case, which here is a regional evacuation generating forty thousand instance state changes a minute against a steady state of two and a half thousand. Sizing for sixteen times steady state is expensive and still arbitrary. But the two kinds of change are not equally urgent, and the asymmetry is large: a withdrawal that arrives late means traffic continues to a failing endpoint and users see errors, while an addition that arrives late means the service runs at slightly less capacity than it has. Treating them identically spends capacity on the cheaper one at precisely the moment the expensive one matters.

Decision. Withdrawal and admission are separately budgeted: 99% of callers see an unplanned withdrawal within 5 s and 99.9% within 15 s, while a newly eligible instance need only reach 99% of callers within 30 s. Under saturation the platform lengthens the batching interval for additions and holds the withdrawal path at its budget, and declares that it has done so. Views are monotonic per client, so a delayed addition can never be applied after a later view that already contains it.

How it is realised on AWS. The view assembler maintains two delta queues per service with different batching windows, and the xDS tier drains the withdrawal queue first. Shed episodes are a reported signal. Version comparability across control-plane replicas comes from a region-wide monotonic counter, so a client that reconnects to a different replica cannot regress.

Option Verdict Reasoning
Separate budgets, shed additions, hold withdrawals Chosen Spends scarce propagation capacity on the change whose lateness users feel.
One propagation path, one budget Rejected Simpler to build and operate, and must be sized for a 16× burst to protect the urgent case.
Priority by service tier rather than by change type Rejected Protects important services and lets a tier-1 addition delay a tier-2 withdrawal, which is the wrong axis.
Push withdrawals, poll for additions Deferred A cleaner expression of the same asymmetry; it means two mechanisms to keep semantically identical for no gain the queues do not already give.

What it buys

  • A 16× churn burst becomes a capacity question about additions rather than a correctness question about withdrawals.
  • The platform can be sized for steady-state additions and burst withdrawals, which is materially cheaper.
  • Monotonic versions mean shedding cannot produce an out-of-order view, only a later one.

What it costs

  • During a burst a service runs with less capacity visible to callers than it actually has, for up to the shed interval.
  • Two queues with different behaviour is a subtlety that shows up in every propagation-lag investigation.
  • The admission budget is deliberately loose, which makes slow start's effectiveness partly dependent on propagation timing.

Choose differently when. If adding capacity late were itself a user-visible failure — an autoscaling-driven service where new capacity is the response to load — the asymmetry narrows and both paths need tight budgets. The decision is also wrong if local ejection were strong enough to make withdrawal propagation uncritical, in which case both budgets could relax.

Why it holds up over time. The asymmetry between removing and adding is a property of what users experience, not of any transport, and it holds for any system that publishes membership.

Lesson. When two kinds of update share a channel, ask which one's lateness the user feels. Then give them different budgets and make the cheap one yield.

Recovery

The decisions that treat coming back as more dangerous than going down.

ADR-12 · Slow start is published by the control plane as a rising weight

Status: Accepted · Shown on views: 04, 14, 18

Who decides how much traffic a newly admitted instance receives in its first minute?

Context. A cold replica has empty caches, unwarmed connection pools and a just-in-time compiler that has not seen production traffic. Giving it an equal share of a service's load the instant it passes readiness is a reliable way to make it fail readiness again, which is the thundering-herd failure in its smallest form. The ramp can live in two places. In the client, it is immediate and proportional to that client's own traffic, and it is only as correct as the least-updated client library in the estate — which, across 450 services and many languages, means it is absent somewhere. In the control plane, published as a rising weight, it is uniform across every client and every heterogeneous data plane, and it is only as smooth as the push interval.

Decision. Slow start is owned by the control plane and expressed as a weight that rises as a function of time in rotation and observed success rate, over an interval of at least 60 s. Clients apply the published weight; they do not compute their own ramp. The same mechanism admits a new deployment version, a recovered instance and a newly started instance, so there is one ramp behaviour rather than three.

How it is realised on AWS. Eligibility rows carry a weight that the evaluator raises on a schedule after admission, and the view assembler publishes weight changes as ordinary deltas. Because admission propagation is budgeted at 30 s, the ramp is published in steps rather than continuously — which is accepted, since the ramp's purpose is to bound the first minute, not to be smooth.

Option Verdict Reasoning
Control-plane-published weights Chosen Uniform across a heterogeneous client fleet, which matters more here than proportional smoothness.
Client-side ramp from a declared start time Rejected Immediate and traffic-proportional, and correct only where the client library is current — which cannot be assumed.
Control plane publishes the policy, clients execute it Deferred The best of both and it still depends on client libraries implementing it; worth revisiting once version spread is under control.
No ramp; rely on readiness to gate admission Right elsewhere Right for stateless services with no warm-up cost, and wrong for anything holding a cache or a pool.

What it buys

  • Every client ramps identically, including DNS clients and target-group-routed traffic that have no ramp logic at all.
  • One mechanism covers deploys, restarts and recoveries, so there is one behaviour to understand and test.
  • The ramp is visible in the published view, which makes it debuggable from the caller's side.

What it costs

  • The ramp is as coarse as the admission propagation interval, so it is a staircase rather than a curve.
  • Weight changes consume propagation capacity for every ramping instance, which at ~900 deploys a day is continuous traffic.
  • A service whose warm-up exceeds 60 s is still hurt, and few teams will tune the interval.

Choose differently when. Once client library versions are uniform enough to trust, moving the ramp into the client removes the propagation cost and makes it proportional to actual load, which is strictly better. For a genuinely stateless service the ramp is pure cost and should be configurable to zero.

Why it holds up over time. The decision is really about where to put behaviour when client heterogeneity is the binding constraint, and that constraint recurs in every platform with a client library.

Lesson. When a safety behaviour must hold everywhere, put it where you control it — even at the cost of doing it less well than the place that has better information.

ADR-13 · Recovery is rate-capped in three independent places

Status: Accepted · Shown on views: 06, 14, 18

What stops a recovered component from being destroyed by the traffic that was waiting for it?

Context. Recovery is the more dangerous event, and the reason is arithmetic rather than subtle. While something is down, demand does not disappear: clients retry, queues fill, and sixty thousand subscriptions wait to reconnect. The instant the thing recovers, all of that arrives simultaneously against a component that is cold. A recovered availability zone receives the traffic of the two zones that were carrying it. A restarted propagation tier receives every client's re-subscription at once. A recovered instance receives its full share before its first cache is warm. Each of these is a different mechanism arriving at the same failure, which is why one cap is not enough.

Decision. Three independent rate caps are platform defaults rather than tuning. No more than 5% of a client population may reconnect to the control plane in any 10 s window, enforced by jittered randomised backoff. A recovered or new instance reaches full weight over at least 60 s. Traffic shifted back into a recovered zone or region is rate-capped, and where the failover was declared manually the shift-back must be explicitly completed rather than happening automatically.

How it is realised on AWS. The xDS tier rejects re-subscriptions above its admission rate with a retry hint, and client libraries apply full jitter on backoff. Weight ramps come from ADR-12. Zone and region shift-back is a policy-API operation that moves topology preference in capped steps, with the final step requiring confirmation for a manual failover.

Option Verdict Reasoning
Three independent caps as defaults Chosen Each mechanism arrives at the same failure by a different route, so each needs its own ceiling.
One global admission control at the control plane Rejected One place to reason about, and it cannot cap traffic shifting between zones, which never touches the control plane.
Client-side backoff only Rejected Standard practice and dependent on every client library doing it correctly, which is the assumption this design avoids making.
Automatic shift-back on recovery Rejected Faster return to full capacity and removes the human judgement that the recovery is real.
Queue-based admission with explicit draining of the backlog Deferred More precise than rate caps for the reconnection case; the caps are sufficient and much simpler to operate.

What it buys

  • A control-plane restart does not become a company-wide reconnection storm, which makes routine deploys of the tier survivable.
  • A recovered zone is reloaded gradually, so the recovery does not immediately produce a second incident.
  • Manual failovers end with a human confirming the return, which is where the judgement belongs.

What it costs

  • Full recovery takes longer than it needs to when the recovered component is genuinely healthy.
  • A manual shift-back that nobody completes leaves a region drained indefinitely, which needs its own alert.
  • Three caps in three places is three sets of parameters, and only one of them is visible from any single view.

Choose differently when. If recovery capacity were provably abundant — a stateless service with no warm-up behind an elastic pool — the caps cost availability for no benefit. Automatic shift-back is correct where failover is itself automatic and cheap to reverse, because then no human was in the loop to confirm anything.

Why it holds up over time. That recovery concentrates demand is a permanent consequence of queueing, and any replacement for these mechanisms will need its own ceilings.

Lesson. Design the recovery path before the failure path. Failure sheds load; recovery concentrates it, and concentrated load against a cold component is how one incident becomes two.

Trust and heterogeneity

Who is allowed to deny service to whom, and how runtimes that cannot answer the question are represented.

ADR-14 · Regional registry authority, with a failover signal that does not depend on the registry

Status: Accepted · Shown on views: 16, 06, 21

Should there be one global registry, or one per region with federation between them?

Context. A single global registry gives one consistent answer, makes cross-region resolution trivial, and makes the global control plane a dependency of every region's ability to route — which contradicts ADR-01 at the regional scale instead of the request scale. Independent regional registries with asynchronous replication keep each region alive through a partition, and guarantee that the regions will sometimes disagree about which region is healthy. That disagreement arrives precisely when the answer matters most, during a partition, and a system that asks the registry whether a region is reachable is asking the wrong component: the registry's own reachability is the thing in question.

Decision. Each region operates an independent control plane with its own registry authority, able to serve all of its own resolution with no dependency on a global component. Cross-region resolution is served from an asynchronously replicated, read-only aggregate of the peers, and replication loss degrades cross-region resolution only. The signal that declares a region unhealthy is out of band and has no dependency on the registry it describes. Reconciliation after a partition is explicit and logged, never derived from timestamps.

How it is realised on AWS. Three independent regional stacks in us-east-1, eu-west-1 and ap-south-1, each with its own Aurora registry and its own control plane on EKS. Cross-region state replicates asynchronously into read-only aggregates. Route 53 Application Recovery Controller readiness checks and routing controls provide the out-of-band region health and failover signal, deliberately chosen because its data plane is separate from the platform's own.

Option Verdict Reasoning
Regional authority with read-only cross-region aggregates and an out-of-band failover signal Chosen A region routes with every peer and every global component unreachable, which is the property that matters.
Single global registry with regional read caches Rejected One consistent answer, and the global plane becomes a dependency of every region's routing.
Regional authority with a global read-only aggregate as the single cross-region view Rejected Nearly the chosen design, with one global component whose loss removes all cross-region resolution at once.
Strongly consistent multi-region registry Rejected Removes disagreement entirely and makes every registry write pay a cross-region round trip, and a partition stop writes.
Single-region control plane for the whole estate Right elsewhere Right for an estate genuinely in one region; here it makes one region's bad day global.

What it buys

  • A regional partition is a cross-region resolution degradation rather than a routing outage anywhere.
  • The failover decision does not depend on the system whose reachability is in question.
  • Each region's control plane can be deployed and drilled independently, which is how the static-serving claim gets tested.

What it costs

  • Regions will disagree about reality after a partition, and the reconciliation is explicit work rather than an automatic merge.
  • Three stacks to operate, upgrade and keep semantically identical.
  • Cross-region resolution is always slightly stale, so a cross-region failover routes on an older view than a local one.

Choose differently when. If the estate's traffic were genuinely global rather than regionally served — every request touching services in two regions — the staleness of a cross-region aggregate becomes a correctness problem and strong consistency starts to earn its cost. A single region of operation makes the whole federation question moot.

Why it holds up over time. Not asking a component about its own reachability is a permanent rule. The specific choice of regions and replication mechanism is not.

Lesson. Never let a system be the source of truth about whether it can be reached. Put the reachability signal somewhere with a different failure domain, even if it is less convenient.

ADR-15 · The power to deny service is a privilege, and every signal must be attributable

Status: Accepted · Shown on views: 19, 20, 10

Who is allowed to mark an instance unhealthy, and what stops that being an attack?

Context. The discovery plane's central capability is removing endpoints from rotation. Stated plainly, it is an authorised denial-of-service mechanism operating continuously across the entire estate, and most implementations treat its inputs as telemetry rather than as security-relevant assertions. If any workload can report health about any instance, then a single compromised pod can remove a service it merely calls. If an operator can always force an instance ineligible, the break-glass path is a denial-of-service path with a badge. And the one evidence class that cannot be forged about oneself — the outcomes callers observed — can still be forged about somebody else.

Decision. Every registration, health report and resolution request is authenticated with workload identity rather than network position. An instance may report health only about itself; a prober only about the instances assigned to it. A signal that cannot be attributed to a registered identity is discarded, not down-weighted. Marking an instance ineligible by hand is privileged, rate-limited per actor per service, requires a justification, and is audited whether or not it succeeds. A registered address must be proved to belong to the identity registering it. Resolution is authorised per caller-callee pair, so the registry is not a map of the estate for any compromised workload.

How it is realised on AWS. IRSA and instance profiles supply workload identity; all control-plane paths use mutual authentication. Prober assignment is held in the registry, so a prober report outside its assignment is rejected at ingest. Overrides go through the policy API, which checks authorisation and the disruption budget before anything is applied, and writes the attempt to the 5-year immutable audit store. Caller outcome reports are weighted by reporter population rather than by report volume, so one noisy reporter cannot dominate.

Option Verdict Reasoning
Identity-attributed signals, self-only reporting, privileged and audited overrides Chosen Treats the removal capability as what it is, and makes the common attack — report about a neighbour — structurally impossible.
Trust signals from any authenticated workload in the mesh Rejected Simple, and gives every workload the ability to deny service to every other.
Network-position-based trust Rejected Requires no identity plumbing and fails completely the first time anything inside the perimeter is compromised.
Unrestricted operator override Rejected Fast under pressure and makes the break-glass credential the most powerful outage tool in the company.
Signed health assertions with per-report non-repudiation Deferred Stronger than identity-scoped attribution, and the per-signal cost at 1.2 M results a second does not currently buy enough.

What it buys

  • A compromised workload cannot remove a service it merely calls, which is the attack this capability otherwise invites.
  • Every denial is attributable after the fact, including the ones that were refused.
  • Per-pair resolution authorisation means a compromised workload cannot enumerate the estate.

What it costs

  • Identity plumbing on every path, including the high-volume passive ingest, where it is the per-sample cost.
  • Prober assignment becomes registry state that must be correct, or legitimate probe reports are rejected.
  • A refused override under incident pressure is a frustrating experience, and the refusal has to explain itself well to be accepted.

Choose differently when. In a single-tenant estate with no meaningful internal threat model, per-pair resolution authorisation is overhead and a flat mesh identity would do. If incident response ever showed that refused overrides were materially extending outages, the budget check should be made overridable with a second authoriser rather than removed.

Why it holds up over time. Reasoning about a capability by what it can destroy rather than by what it is called is a permanent discipline, and it survives every change of identity technology.

Lesson. Name your system's capabilities by their effect. A health-check input is an authorisation to deny service, and it should be secured like one.

ADR-16 · One registry over heterogeneous runtimes, with capability declared per service

Status: Accepted · Shown on views: 14, 01, 09

How should workloads that cannot answer a readiness question be represented?

Context. The estate is not uniform and will not become uniform. Alongside twelve orchestrated clusters there are long-lived VM fleets, managed functions behind an invoke path, remaining data-centre services, and third-party endpoints that publish no health signal of any kind. Forcing all of them into one abstraction gives callers a single resolution surface and one mental model, at the price of an abstraction whose health contract is only as rich as its weakest member. Giving the heterogeneous tail a second mechanism keeps the orchestrated path clean and forces every caller to know how its callee is deployed, which is exactly the coupling a discovery plane exists to remove.

Decision. All workloads are services in one registry with one resolution surface. Each service declares a capability profile stating which health signals its runtime can actually supply and which are unavailable, and the platform never reduces a richer service's contract to the weakest profile in the estate. A service with no readiness signal is represented as eligible with health unknown, stated in the view rather than implied, and is governed by passive caller outcomes alone.

How it is realised on AWS. The service row carries a capability_profile; the health contract's required fields are validated against it, so a contract demanding readiness from a third-party endpoint is rejected at authoring time. Views carry the capability per endpoint so a client can tell an unprobed endpoint from a probed healthy one. Non-orchestrated fleets renew leases; orchestrated ones do not need to.

Option Verdict Reasoning
One registry, capability declared per service Chosen Callers get one surface and the weak cases stay honest rather than being hidden behind a uniform contract.
One registry with a synthetic health adapter per runtime Rejected Makes every endpoint look probed, which is a lie the caller cannot detect and will route on.
Separate external-dependency registry with its own semantics Rejected Keeps the orchestrated path pristine and makes every caller know its callee's deployment model.
Lowest-common-denominator contract for the whole estate Rejected Uniform and simple, and discards the readiness and dependency signals that most of the estate can actually supply.
Orchestrator-native discovery plus a side channel for the rest Right elsewhere Right where the non-orchestrated tail is small and shrinking; here it is neither.

What it buys

  • A caller resolves by name without knowing whether its callee is a pod, a fleet instance or somebody else's endpoint.
  • Capability is visible in the view, so 'health unknown' is a fact the client can act on rather than an assumption.
  • The orchestrated majority keeps its full contract instead of being levelled down.

What it costs

  • Clients must handle an endpoint whose health is unknown, and some will route to it as though it were verified.
  • Capability profiles are configuration that can be wrong, and a wrong profile silently weakens or over-promises a contract.
  • Four registration paths behind one surface is more platform code than a single-runtime design needs.

Choose differently when. If the non-orchestrated tail shrank to nothing, orchestrator-native discovery becomes sufficient and this registry is a layer with no job. If callers proved unable to handle unknown-health endpoints correctly, a separate registry for them — forcing an explicit decision at the call site — would be the safer design despite the coupling.

Why it holds up over time. Declaring capability rather than assuming it is how any abstraction over heterogeneous providers stays honest, and that outlives the particular runtimes.

Lesson. When unifying things that are not the same, make the differences declarable and visible. An abstraction that hides what it cannot deliver moves the failure to the caller, who has less information than you did.

Every package used, in one table

These terms are used precisely in this package. Several are used loosely in the wider literature on service discovery and health checking, and the loose readings are what make two discovery designs disagree with each other while appearing to say the same thing.

Package What it is What it does here Considered instead
Liveness Whether the process is running and its port is open. Decides whether to restart an instance; it is the orchestrator's question, not the router's. "Health", which collapses it with readiness and makes a restarting replica indistinguishable from a busy one.
Readiness The instance's own declaration that it is willing to receive traffic now. The instance's vote in its own eligibility; one of three evidence classes. "Up", which hides that the instance is asserting something it may be wrong about.
Dependency health The instance's view of the downstreams it needs in order to do useful work. Reduces weight by default; may set unreadiness only for a declared non-shared dependency. Folding it into readiness, which is how a shared database wobble becomes a total outage.
Eligibility The platform's single verdict per instance: eligible, degraded with a weight, ineligible, or unknown. The only thing published to clients; everything upstream is evidence for it. "Healthy/unhealthy", which has no way to say that the observer is the broken part.
Unknown The state of an instance whose evidence has gone silent because the observation path failed. Keeps the instance in rotation and requires corroboration before removal. Treating silence as unhealthy, which drains a healthy fleet the first time a prober network partitions.
View A per-service, versioned endpoint set with weights, zones, versions and topology priority tiers. The unit of publication, of monotonic application at the client, and of staleness. "The endpoint list", which hides the version and therefore hides how old it is.
Last-known-good cache The durable per-client copy of the most recent view it successfully applied. The authority for routing, and the reason a control-plane outage is a staleness event. "Cache", which implies an optimisation rather than the system of record for routing.
Fail static Continuing to act on the last known state when new information cannot be obtained. The platform's required behaviour on control-plane unavailability. "Fail open" and "fail closed", both of which describe a decision rather than the refusal to make a new one.
Withdrawal An instance becoming ineligible, and the propagation of that fact to callers. The urgent path, budgeted at 5 s to 99% of callers and never shed. "Deregistration", which is a different event: the instance ceasing to exist.
Minimum healthy fraction The proportion of a service's instances per zone below which health is disregarded. The backstop that turns correlated mass withdrawal into a declared degradation. "Panic threshold", which describes the mechanism's mood rather than its contract.
Flapping An instance oscillating between eligibility states faster than its damping absorbs. Triggers escalating quarantine and an owner-visible defect report. "Noise", which suggests the right response is more filtering rather than isolation.
Slow start A rising published weight that brings a newly admitted instance to full share over at least 60 s. The ramp that stops a cold replica being killed by its own admission. "Warm-up", which usually names an application behaviour rather than a routing weight.
Outlier ejection A client removing an endpoint that is failing for that client specifically. The fast path: it acts before the control plane has an opinion, under a per-client cap. "Circuit breaking", which is about protecting the caller rather than correcting the endpoint set.
Failure domain A region, zone or cluster, modelled explicitly so correlated failure is expressible. Carries topology priority tiers and makes evacuation a one-step preference shift. "Location", which carries no statement about what fails together.
Capability profile A per-service declaration of which health signals its runtime can actually supply. Keeps a third-party endpoint honest without levelling the whole estate to its contract. Assuming a uniform health contract, which makes an unprobed endpoint look verified.
The package

Everything as it was delivered.

These files are served exactly as they were produced — the diagram pages keep their own house style because that is the artifact, not a rendering of it.