Health Check & Service Discovery

Architecture Views

21 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

Read this set in order. It opens with the boundary and the people, because the whole argument rests on one asymmetry: the control plane may fail for an hour and traffic must not notice, while a replica that starts failing must stop receiving traffic in five seconds. Every later view either honours that asymmetry or contradicts it. Every number in the set is a stated assumption, sized for 450 services, 60,000 instances and 12 million internal requests a second.

Context and scope

What is inside the boundary, who feeds it, and what it is not allowed to depend on.

People and journeys

Who this is for, and the three moments where getting it wrong is visible to a user.
03 The people who change traffic Service owner 450 services, ~120 teams Goal — Ship twice a day without anyone noticing, and get told if my readiness check is lying about me. Core journeys Deploy without dropping traffic Declare a health contract Read my own flap report SRE on call follow-the-sun, 3 regions Goal — Decide in two minutes whether this is one bad replica, one bad zone, or my own health check. Core journeys Chase errors from one replica Evacuate a zone or region Override an eligibility decision The people who are owed an answer Security engineer quarterly review Goal — Know that nothing but a service itself can mark that service unhealthy, and prove it from a log. Core journeys Audit a denial of service Review resolution authorisation Platform engineer owns this plane Goal — Prove the request path survives my own control plane being switched off, before it is switched off for me. Core journeys Run a control-plane removal drill Onboard a non-orchestrated fleet Machines in the cast Calling service 12 M req/s peak Goal — Get an endpoint list I can route on right now, even if nothing answers me. Core journeys Resolve a callee by name Eject a bad endpoint locally Deploy pipeline ~900 deploys/day Goal — Be told no when draining this instance would take the service below its floor. Core journeys Request a drain Ramp a new version Third-party endpoint health unknown Goal — Be routed to without pretending I publish a readiness signal. Core journeys Be declared with no probe Who This Is For, and What They Get To Do Person or role Journey / task Application we own Security / platform External / third party v 1.0 · owner Reliability Architecture Actors and Their Journeys Four humans, three machines, and the one actor who can break this platform by doing their job correctly. HTML page SVG draw.io

Structure

The parts, the planes, and the one seam that separates advisory from authoritative.

Data

What is stored, what is recomputed, and what is deliberately never durable.

Runtime

What actually happens when a replica fails, when one starts, and how a verdict is reached.

Operations

Where the platform runs, what is watched, and the loop an instance travels.

Assurance

Who may deny service to whom, and what catches each failure class.

Architecture One-Pager

The problem, the shape of the answer, the decisions that carry it, and what a prototype should prove.

The client's last-known-good view is authoritative for routing; the control plane is an advisory service that improves that view and is never in the request path.

When a team chat shows a spinner on send, when a shared document says "reconnecting", when the driver freezes on the map for eight seconds, the user is rarely looking at a crashed service. They are looking at a request routed to a replica that should not have been in rotation — one restarting, one whose connection pool was exhausted, one in a zone already declared unhealthy whose address was still cached in the caller. The opposite failure is less visible and more expensive: a bad health-check config, or a shared database wobbling, and a fleet takes itself out of rotation faster than it can be put back. Something has to answer two questions continuously for every one of twelve million internal calls a second — which instances of this service exist, and which should receive traffic — and that something is itself a distributed system that will have a bad day. The problem is not knowing which replicas are healthy. It is being allowed to be wrong about it, briefly and boundedly, without that being the same thing as an outage.

Reconcile instance existence from the orchestrator's own state rather than asking workloads to register themselves, and collect health as three independent evidence classes: active probes from outside, self-reports from inside, and the outcomes real callers actually observed. Fuse them into one eligibility decision per instance over a decaying score with asymmetric hysteresis, where silence is Unknown rather than unhealthy and a degraded shared dependency reduces weight rather than removing capacity. Assemble per-service versioned views with explicit topology priority tiers, and push incremental deltas to sixty thousand streaming subscriptions, budgeting withdrawal at five seconds and admission at thirty. Every client keeps the last view it received on local disk and routes from it — through a control-plane outage, for at least an hour, with no behaviour change but staleness. Removal is bounded in three independent places: a minimum healthy fraction that disregards health rather than draining a service, a cap on how fast a client's endpoint set may shrink, and a cap on how much of it any one client may eject. Recovery is treated as the dangerous event: slow start, jittered re-subscription and rate-capped shift-back are defaults.

What it is, and what it is not

A client cache that is authoritative for routingA registry consulted, or leased from, on the request path
Fail static on missing informationFail closed, which turns a blind observer into an outage
Liveness, readiness and dependency health as three questionsOne health boolean per instance
Unknown as a first-class stateUnknown silently treated as unhealthy
Withdrawal and admission on separate budgetsOne propagation path with one latency target
Removal bounded by a minimum healthy fractionEvery withdrawal honoured, however many arrive at once
Recovery protected by defaultRecovery as the moment everything is allowed back at once
Regional authority with explicit reconciliationA global registry every region depends on to route
Capability declared per serviceThe weakest available health signal applied to the whole estate

The decisions that are the architecture

01The data plane is authoritative; the control plane is advisory

A view already delivered keeps working. The control plane's only job is to improve it, which lets its availability target be 99.95% while resolution as callers experience it is five nines, and turns a registry outage into a staleness problem with a bounded blast radius rather than a company-wide outage.

ADR-01

02Computed state is disposable; only desired and observed state are durable

Eligibility decisions and published views are rebuilt from the registry and the sample store, never restored. That is what makes a five-minute RTO a rebuild rather than a backup, and it is why a scoring bug is repaired by recomputing rather than by surgery on live routing state.

ADR-02

03Registration is reconciled, not self-reported

The orchestrator already knows which pods and tasks exist; asking the workload to announce itself adds a failure mode that looks like a healthy fleet with missing members. Explicit registration exists only for the runtimes nothing watches.

ADR-03

04Three independent evidence classes, with passive outcomes primary where traffic allows

A probe measures a path no caller uses, a self-report can be confidently wrong, and the outcomes real callers saw are both free and the only class a workload cannot forge about itself. Fusing all three is what makes grey failure visible.

ADR-04

05Unknown is a state, and silence never removes an instance

A system with only healthy and unhealthy drains itself the first time its observation path breaks. Removal on missing information requires corroboration from a second evidence class.

ADR-05

06Dependency health reduces weight; it may not withdraw a shared dependency

If every replica shares one database, honest unreadiness turns a partially functional service into a completely unavailable one. Unreadiness from a dependency is permitted only where the dependency is declared non-shared.

ADR-06

07Asymmetric hysteresis over a decaying score, with a hard-failure fast path

Re-entry is strictly slower than removal, scoring responds to failure rate rather than the last sample, and a closed port skips the decay entirely — because damping exists to absorb noise, not to argue with a refused connection.

ADR-07

08The minimum healthy fraction disregards health rather than draining a service

Below 50% eligible in a service's zone the platform returns the full endpoint set and says so. A bad health-check deploy is then a declared degradation instead of a self-inflicted outage.

ADR-08

09Flapping is quarantined with escalating backoff

An oscillating instance destabilises every client's view, which is worse than either verdict. Quarantine is surfaced to the owning team as an event, because a quarantined instance is a bug report about a health check.

ADR-09

10Streaming subscriptions are the primary surface; DNS is compatibility

Push with exact version tracking is the only way to hold a five-second withdrawal budget across sixty thousand clients. DNS stays, because a resolver that ignores TTLs is still a client that has to be served.

ADR-10

11Withdrawal and admission are budgeted separately

Five seconds to remove, thirty to add, and under saturation the addition path is the one that sheds. Adding capacity late is cheaper than adding it wrong.

ADR-11

12Slow start is published by the control plane as a rising weight

Uniform across every client and every heterogeneous data plane, which matters more than being traffic-proportional in an estate where the least-updated client library decides whether a client-side ramp exists at all.

ADR-12

13Recovery is rate-capped in three places

Reconnection, traffic share and regional shift-back all have ceilings, because the moment everything is allowed back at once is how a recovered target is killed a second time.

ADR-13

14Regional authority, with an out-of-band failover signal

A region routes with both peers and every global component unreachable. The signal that declares a region unhealthy deliberately does not depend on the registry it is talking about.

ADR-14

15The power to deny service is a privilege, and signals must be attributable

An instance may report only about itself, a prober only about its assigned targets, and an unattributable signal is discarded. Marking an instance ineligible is rate-limited and audited because it is functionally a denial of service.

ADR-15

16One registry over heterogeneous runtimes, with capability declared per service

A caller gets one resolution surface whether its callee is a pod, a fleet instance or somebody else's endpoint — and a service that cannot supply a readiness signal says so, rather than the whole estate being reduced to the weakest contract.

ADR-16

Why this should still be right in ten years

Orchestrators, mesh data planes, discovery protocols and managed DNS will all be replaced inside a decade. The parts of this design that should outlive them are the ones that are statements about authority and blast radius rather than about products.

Authority is not a protocol

"The thing holding the view routes; the thing computing it advises" is a claim about where authority lives. It survives replacing xDS with whatever follows it, Envoy with another data plane, and the registry store with anything, because none of those changes which copy of the truth is allowed to route.

Absence of evidence outlives every monitoring stack

A system that cannot distinguish "I know it is broken" from "I cannot see it" will eventually drain a healthy fleet, whatever is doing the observing. Unknown as a first-class state costs almost nothing and removes the largest self-inflicted risk in the design.

Removal must be bounded in more than one place

The cost of a wrong removal scales with fleet size; the cost of a wrong retention does not. That asymmetry is arithmetic, not fashion, so a minimum healthy fraction, a shrink cap and an ejection cap will still be the right three bounds when every component here has been replaced.

Recovery stays more dangerous than failure

Failure sheds load and recovery concentrates it, because demand does not disappear while something is down. Slow start, jitter and rate-capped shift-back are consequences of queueing, which no technology change repeals.

What will date fastest

The propagation mechanism is a current answer to a permanent question, and the numbers are a reference estate rather than a measurement. Expect both to be replaced; expect the four statements above to survive it.

Non-functional targets

Every target below is a stated assumption for this exercise, chosen so that a reviewer can disagree with one and follow it to the decision that depends on it.

QualityTargetHow it is metView
Resolution availability, as a caller experiences it ≥ 99.999% monthly Served from a durable per-client cache; the control plane is not consulted on the request path 02
Control plane availability ≥ 99.95% monthly per region Regional authority, 18 replicas across 3 AZs, no global request-path dependency 16
Static-serving window ≥ 60 minutes, behaviour unchanged Last-known-good cache on local disk, verified by a routine production removal drill 11
Resolution latency p99 ≤ 1 ms from cache; ≤ 50 ms cold In-process lookup against the cached view; cold path only on first subscription 08
Unplanned withdrawal propagation 99% of callers in 5 s, 99.9% in 15 s Incremental delta push over streaming subscriptions, withdrawal path never shed 13
Graceful drain visibility 99% of callers in 2 s, 99.9% in 5 s Drain-to-termination interval derived from measured propagation delay 04
Admission propagation 99% of callers within 30 s Deliberately slower than withdrawal; the addition path is the one that sheds under load 10
Detection of a failing replica 3 s p50, 10 s p99 from first failed request Passive caller outcomes fused with probe and self-report; hard failures skip the decay 15
Eligibility stability ≤ 4 changes per instance per 10 min Asymmetric hysteresis, decaying score, escalating quarantine; damping adds ≤ 5 s to p99 detection 15
Blast radius of a wrong removal ≥ 50% eligible per service per zone Minimum healthy fraction disregards health; client shrink cap 25% per 10 s; per-client ejection cap 21
Churn absorption 40,000 state changes/min for 10 min Evaluation sharded by service, propagation scaled by subscription count and change rate 06
Signal throughput 1.2 M results/second Bounded prober assignment, write-sized observed-state store with a 7-day TTL 17
Recovery RTO 5 min control plane; RPO 0 desired state Computed views rebuilt from desired and observed state on a rehearsed schedule; request path does not stop 11
Cost ≤ $0.012 per instance per day; probes ≤ 0.5% of internal bytes Passive signal preferred where traffic permits; per-service cost attribution exposed to owners 17

Scope

In scope

  • Registration and non-reusable instance identity across orchestrated and non-orchestrated runtimes, with leases
  • Health collection as three independent evidence classes: active probe, self-report, and caller-observed outcomes
  • Eligibility evaluation with hysteresis, flap quarantine, weights and the minimum healthy fraction guard
  • Per-service versioned views with topology priority tiers, published over streaming, DNS and a query API
  • Drain, disruption budgets, deploy ramps, and zone or region evacuation as a reversible preference shift
  • Recovery behaviour: slow start, reconnection caps, rate-capped shift-back, and the static-serving window
  • Signal attribution, resolution authorisation, and the audit trail behind every eligibility transition

Explicitly out of scope

  • Load-balancing algorithms beyond the endpoint set and weights the platform hands to them
  • The service mesh's policy and mTLS planes, which consume these views rather than producing them
  • The public API edge and anything facing an external client
  • The observability platform that stores the signals this plane exports
  • Application readiness logic itself — the platform defines the contract, the service decides what passing means

What a four-week prototype should prove

The prototype's job is to falsify the three claims the whole design rests on: that routing survives the control plane being switched off, that a withdrawal reaches a realistic client population inside five seconds while the system is busy, and that a deliberately broken health check degrades a service rather than draining it. Everything else in this package is a consequence of those three holding, and nothing else is worth four weeks.

  1. One EKS cluster, 40 services, 2,000 instances, Envoy sidecars with a persisted view cache
  2. One regional control plane: reconciler, prober DaemonSet, passive outcome ingest, evaluator, xDS tier
  3. One non-orchestrated fleet behind the registration API, and one declared third-party endpoint with no readiness signal
  4. Synthetic caller traffic with per-endpoint outcome reporting, and a load generator able to produce a 16× churn burst
  • Kill the control plane for 60 minutes under steady traffic: error rate must not move, and a restarted sidecar must route from disk without blocking
  • Fail one replica hard, and separately make one replica slow but alive: measure time to 99% of callers for each, and show the probe-green case is still caught
  • Push a health-check configuration that fails every replica of one service: the service must degrade with a declared 'health disregarded' view and an alert, not drain
  • Evacuate one zone during a 40,000-changes-per-minute burst: the withdrawal budget must hold while the addition path visibly sheds
  • Restart the xDS tier with 2,000 subscriptions attached: reconnection must stay under the 5%-per-10-s ceiling and no client view may regress
  • Partition the prober network from one AZ: those instances must become unknown and stay in rotation, not be withdrawn

Open risks, carried rather than hidden

RiskIf it landsResponse
The 60-minute static window is a claim nobody tests, so the first real control-plane outage is also the first test The design's central promise fails exactly when it is needed, and the blast radius is every service at once Routine production removal of the control plane from a sample of traffic, with static-serving minutes as a reported signal rather than an incident-only metric
Client library heterogeneity means the client-side guards do not actually exist everywhere Shrink caps and ejection caps are assumed by the control plane and absent in the oldest 5% of callers, which are the ones that will drain a service Report library version spread as a first-class signal, gate onboarding on a minimum version, and keep the control-plane-side fraction guard as the backstop that does not depend on clients
Passive signal is unavailable for low-traffic services, where probing remains the only evidence Grey failure stays invisible exactly where there is least redundancy to absorb it Per-service evidence weighting with a declared traffic threshold below which probe cadence is raised, and an explicit report of which services have single-class evidence
The xDS tier is stateful and holds 60,000 long-lived connections, so its own deploy is the platform's most dangerous routine operation A propagation-tier deploy becomes a company-wide reconnection storm Jittered re-subscription with a 5%-per-10-s ceiling, connection draining on tier rollout, and the static window as the fallback that makes the storm survivable
Regions will disagree about which region is healthy, and the disagreement arrives at the worst moment Two regions each route away from the other, or each declares itself the survivor Regional authority with read-only cross-region aggregates, an out-of-band failover signal independent of the registry, and reconciliation after partition that is explicit and logged rather than derived
A compromised caller can lie about someone else's health through the passive outcome channel A workload with no privileges degrades a service it merely calls Per-client ejection caps, outcome reports weighted by reporter population rather than volume, and attribution at ingest so the lying reporter is identifiable after the fact

Architecture Decision Record

Why every component and every technology on these 21 views is what it is, and what each choice costs.

Sixteen decisions make up this architecture. Everything else across the twenty-one views is a consequence of one of them, and each record carries the question that forced it, how it is realised on AWS, the alternatives including the ones that are right for a different organisation, and the conditions under which the choice should be reversed.

Status of this document. This is a design, not a report on a running system. Every rate, latency, volume, retention and cost figure in this package is a stated assumption, chosen to be defensible for a widely used collaboration SaaS at roughly 25 million daily active users and sized from a reference estate of 450 internal services, 60,000 instances, 3 active regions and 12 million internal service-to-service requests a second at peak. They are written as numbers so that they can be argued with and corrected, which vagueness does not allow.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

Authority and blast radius 2

Where routing authority lives, what is durable, and how wrong the platform is allowed to be.

ADR-01The client's last-known-good view is authoritative for routing; the control plane is advisory ADR-02Computed state is disposable; only desired and observed state are durable

Evidence and what counts as health 4

How the platform learns an instance is in trouble, and what it refuses to infer from silence.

ADR-03Registration is reconciled from the orchestrator, not self-reported by the workload ADR-04Three independent evidence classes, with passive caller outcomes primary where traffic allows ADR-05Unknown is a first-class state, and silence alone never removes an instance ADR-06Dependency health reduces weight; it may not withdraw an instance over a shared dependency

Decision stability 3

Turning a stream of noisy samples into a verdict that does not oscillate, and bounding what a wrong verdict costs.

ADR-07Asymmetric hysteresis over a decaying score, with a fast path for hard failures ADR-08The minimum healthy fraction disregards health rather than draining a service ADR-09A flapping instance is quarantined with escalating backoff, and its owner is told

Propagation and convergence 2

Getting a changed view to sixty thousand clients, and the asymmetry between removing and adding.

ADR-10Streaming subscriptions are the primary surface; DNS is a compatibility surface ADR-11Withdrawal and admission are separately budgeted, and admission is the path that sheds

Recovery 2

The decisions that treat coming back as more dangerous than going down.

ADR-12Slow start is published by the control plane as a rising weight ADR-13Recovery is rate-capped in three independent places

Trust and heterogeneity 3

Who is allowed to deny service to whom, and how runtimes that cannot answer the question are represented.

ADR-14Regional registry authority, with a failover signal that does not depend on the registry ADR-15The power to deny service is a privilege, and every signal must be attributable ADR-16One registry over heterogeneous runtimes, with capability declared per service

Technology by capability

Amazon Web Services was chosen for this exercise deliberately. It is tied for the least-used stack across this practice's existing use cases and has not been used since 2026-09, while the two most recent use cases both ran on Google Cloud — the rotation this library exists for. It also happens to fit the topic honestly: health checking and discovery here is a multi-runtime problem, and AWS supplies genuinely distinct primitives for containers, tasks, VM fleets, managed functions and the DNS surface, which is the heterogeneity the design has to absorb rather than wish away. The data plane is deliberately open source, because the one component that must be identical across every runtime is the one the platform ships into every workload.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Orchestrated instance state EKS EndpointSlice watch across 12 clusters, plus ECS task state change events Amazon Web Services Self-registration from each workload; Consul agents The orchestrator already holds reconciled desired state, so a missing instance is a diff rather than a silence ADR-03
Non-orchestrated instance state EC2 Auto Scaling lifecycle hooks plus an authenticated registration API with leases Amazon Web Services Agent-based self-registration everywhere Covers fleets nothing watches without extending self-registration to the orchestrated majority ADR-16
Desired state store Aurora PostgreSQL, 3-AZ, one writer per region Amazon Web Services DynamoDB global tables; etcd; Consul Small, exact, relational state with strong in-region consistency; RPO 0 matters more than write scale here ADR-02
Observed state store DynamoDB on-demand with a 7-day TTL Amazon Web Services Timestream; a Prometheus-compatible TSDB; Cassandra Sized for 1.2 M writes/second where losing a window of samples is acceptable by design ADR-02
Health signal ingest Kinesis Data Streams, sharded by service Amazon Web Services MSK; SQS; direct gRPC to the evaluator Shards give per-service isolation so one service's churn cannot delay another's evaluation ADR-04
Active probing Per-AZ DaemonSet probers with bounded target assignment Open source on EKS Route 53 health checks for everything; a central prober fleet Sub-linear probe cost in instance count, and a prober failure scoped to one AZ rather than a region ADR-04
Passive outcome reporting Envoy per-endpoint outcome counters reported to ingest Open source (Envoy) Application-level reporting; inference from service-mesh telemetry The only evidence class that measures what callers experience and that a workload cannot forge about itself ADR-04
Eligibility evaluation Stateless evaluator on EKS, sharded by service, state in memory Open source on EKS Stream-processing framework; evaluation inside the proxies Computed state is disposable, so the evaluator can be restarted and rebuilt rather than recovered ADR-02
View propagation Envoy-compatible xDS tier with incremental deltas and per-service versions Open source (xDS) on EKS Consul; short-TTL DNS only; gossip; long-poll The only mechanism whose propagation latency the platform owns, which a 5 s withdrawal budget requires ADR-10
DNS compatibility surface Cloud Map plus Route 53 private zones on a short TTL Amazon Web Services CoreDNS only; no DNS surface at all Serves clients that cannot hold a subscription, with the weaker guarantee stated rather than hidden ADR-10
Client data plane Envoy sidecar with a durable last-known-good cache, plus a thin resolver library Open source (Envoy) Client-side load balancing per language; a central routing proxy tier Routing authority has to live where the request is, and it has to be identical in every runtime ADR-01
Runtime edge membership NLB and ALB target groups synchronised from the published views Amazon Web Services Target group health checks as the source of truth One eligibility semantics for edge-routed and sidecar-routed traffic, with no second opinion ADR-01
Region health and failover signal Route 53 Application Recovery Controller readiness checks and routing controls Amazon Web Services Registry-derived region health; a custom quorum service A region's reachability must not be asserted by the system whose reachability is in question ADR-14
Workload identity IRSA for pods, instance profiles for fleets, mutual TLS on every control-plane path Amazon Web Services SPIFFE/SPIRE; network-position trust Self-only health reporting is only enforceable if every signal carries an identity the platform can check ADR-15
Audit and transition record Kinesis to S3 with Object Lock, 5-year retention Amazon Web Services CloudTrail only; the relational store Denials of service must be reconstructable long after the samples that caused them have expired ADR-15
Policy and contract authoring Policy API with staged rollout, disruption budgets and one-step reversal Open source on EKS GitOps-applied CRDs; direct registry writes Health-check configuration can drain a fleet, so it needs the same gates as a code deploy ADR-08

The decisions, and the alternatives that lost

Authority and blast radiusWhere routing authority lives, what is durable, and how wrong the platform is allowed to be.

ADR-01

The client's last-known-good view is authoritative for routing; the control plane is advisory

Accepted

When a caller needs an endpoint and the discovery system is unreachable, what is the correct behaviour?

Context
The obvious design makes the registry the single truth and has callers consult it, or hold a short lease from it, per resolution. It gives one answer that is always current, and it means every one of twelve million internal requests a second depends on one distributed system being available. That system's availability then becomes the ceiling on every product's availability, and its worst day is everybody's worst day. The alternative is to accept that routing runs on a view that may be stale, and to make staleness the normal operating condition rather than an error — which moves the hard problems into the client and makes propagation delay the number the design lives or dies by.
Decision
The view a client already holds is authoritative for routing. The control plane computes and publishes improvements to that view and is never consulted synchronously on a request. Every client persists its last-known-good view locally and continues routing from it, unchanged, through a total control-plane outage for at least 60 minutes, degrading only by staleness. A resolution that cannot be served from cache or stream fails to a documented default; it never blocks.
How it is realised on AWS
Envoy sidecars hold incremental xDS subscriptions against the regional control plane on EKS and write each applied view to a local volume. On stream loss the sidecar keeps serving the persisted view and re-subscribes with jittered backoff. The thin resolver library used by sidecar-less callers implements the same cache-first contract. NLB and ALB target groups are synchronised from the same views, so edge-routed traffic inherits the behaviour without taking a dependency of its own.
Options weighed
  • ChosenClient-held last-known-good view is authoritative; control plane advisory: Decouples caller availability from control-plane availability entirely; pays in staleness and in guards that must exist twice.
  • RejectedRegistry consulted or leased per resolution: Exactly consistent and makes the discovery plane a hard dependency of every request in the company.
  • RejectedShort-TTL DNS only, no persistence: Looks like caching but is not: a resolver that honours a 5 s TTL has no view at all 5 s into an outage.
  • Right elsewhereCentral routing proxies holding the view on the caller's behalf: Right where client libraries cannot be changed; adds a network hop and a new tier that must itself be more available than the control plane.
Consequences
What it buys
  • The control plane can be deployed, patched, scaled and drilled like an ordinary service, because its unavailability is a staleness event rather than an outage.
  • A registry failure, a propagation-tier deploy and a cross-region partition all collapse into the same bounded failure mode.
  • Resolution latency is an in-process lookup at p99 ≤ 1 ms, which removes discovery from every latency budget in the estate.
What it costs
  • Every safety mechanism that bounds a wrong decision must exist in the client as well as the control plane — the shrink cap, the ejection cap, the refusal of an empty set.
  • Correctness now depends on the least-updated client library in the estate, which is a fleet-management problem rather than an architecture one.
  • Withdrawal propagation delay becomes the single most important number in the design, and the only honest measure of how wrong the data plane is permitted to be.
Choose differently when
If the estate were small enough that one control plane could credibly be more available than every service calling it, or if routing to a withdrawn endpoint were unrecoverable rather than retryable — a payment captured twice, an irreversible write — then exact consistency is worth the dependency and the registry should be the authority. The decision should also be revisited if client library heterogeneity becomes unmanageable, because this design silently delegates safety to those libraries.
Why it holds up over time
This is a statement about where authority lives, not about a protocol. It survives replacing xDS, Envoy, the orchestrator and the registry store, because none of those changes which copy of the truth is allowed to route.
LessonDecide early whether a consumer of your system is allowed to be briefly wrong or must always be right. If briefly wrong is acceptable, put the truth in the consumer and make your own availability somebody else's non-event.
Shown on views02 08 11
ADR-02

Computed state is disposable; only desired and observed state are durable

Accepted

Which of the platform's own data has to survive a total loss of the control plane, and which should be thrown away rather than restored?

Context
A discovery plane holds three quite different kinds of data and the temptation is to treat them alike. Service definitions, instance rows, health contracts and policy are small, exact and irreplaceable: nothing can recompute who owns a service or what its readiness contract says. Probe results and caller outcomes are enormous, arrive at over a million a second, and matter only in aggregate over the last few minutes. Eligibility decisions and published views are derived from the other two by a pure function. Storing the third class durably invites repairing it, and a platform that repairs its routing state in place has no recovery story at all.
Decision
Desired state is the system of record with RPO 0, held in a replicated store that is strongly consistent within a region. Observed state is recomputable with RPO 60 s, held in a store sized for write volume rather than durability, where losing a window of samples reduces confidence without corrupting a verdict. Computed state — eligibility, weights and versioned views — is held in memory only, has no RPO, and is rebuilt rather than restored. Every eligibility transition and policy change is additionally emitted to a durable append-only log for audit and incident reconstruction.
How it is realised on AWS
Registry and policy on Aurora PostgreSQL across three AZs. Observed state in DynamoDB with a 7-day TTL and on-demand capacity sized for 1.2 M writes a second. Eligibility and assembled views live in the control-plane pods' memory and are reconstructed on start from Aurora plus the last few minutes of DynamoDB samples. Transitions stream to Kinesis and land in S3 with Object Lock for the 5-year audit retention.
Options weighed
  • ChosenThree state classes with three durability contracts: Makes RTO a rebuild throughput question and removes any temptation to repair routing state.
  • RejectedOne durable store for everything: Simplest to operate and makes the sample write rate the sizing constraint for the store holding the service catalogue.
  • RejectedPersist computed eligibility for faster restart: Buys seconds of startup time and reintroduces repairable routing state, which is the failure mode being designed out.
  • DeferredEvent-sourced control plane with a replayable decision log: Attractive for reconstructing exactly what the platform believed at a past instant; the transition log gives most of it without the operational weight.
Consequences
What it buys
  • RTO 5 minutes is a rebuild measurement rather than a restore hope, and the rebuild is exercised on a schedule instead of during an incident.
  • A scoring or fusion bug is repaired by recomputing from retained samples, with no state surgery.
  • The sample store can be cheap and lossy, which is what makes a million writes a second affordable.
What it costs
  • The rebuild path must be fast enough to meet the RTO at full scale, which is a continuing throughput commitment and not a one-off test.
  • Losing up to 60 s of samples briefly reduces confidence in eligibility, which interacts with the unknown state and has to be reasoned about together.
  • Two stores with different consistency models mean two operational runbooks.
Choose differently when
At a tenth of this scale the sample volume no longer forces a separate store, and one strongly consistent store for everything is simpler and defensible. Persisting computed state becomes right if control-plane restarts ever need to be measured in seconds rather than minutes — for example if the platform were ever placed on the request path, which ADR-01 forbids.
Why it holds up over time
The classification by what-happens-if-lost outlives any particular store. Swapping Aurora, DynamoDB or Kinesis changes nothing about which data may be thrown away.
LessonSort your state by what happens if it disappears, not by which component produced it. Most of what a control plane holds should be unrecoverable-by-design and recomputed, because that is what makes recovery routine.
Shown on views11 12 16

Evidence and what counts as healthHow the platform learns an instance is in trouble, and what it refuses to infer from silence.

ADR-03

Registration is reconciled from the orchestrator, not self-reported by the workload

Accepted

How does the platform learn that an instance exists?

Context
The classical service-discovery pattern has each process register itself on start and deregister on stop. It is simple, it works in a demo, and it has two failure modes that look nothing alike. A process that crashes before deregistering leaves a zombie endpoint that is only removed when its lease expires. A process that starts but fails to register is worse: the fleet looks healthy and is short of members, and nothing in the system is in a position to notice, because the only thing that knew the instance should exist was the instance. Meanwhile the orchestrator already holds exactly this information, as reconciled desired state, for the majority of the estate.
Decision
For orchestrated workloads, instance existence is derived by continuous reconciliation against the orchestrator's own state; the workload does not register itself. An authenticated registration API exists only for runtimes nothing watches. Every instance receives a registry-unique identity that is stable for its lifetime and never reused, so a restarted process is a new instance rather than a continuation. Ungraceful termination is caught by lease expiry within 30 s; graceful termination deregisters before the process stops accepting connections.
How it is realised on AWS
A reconciler watches EndpointSlices across 12 EKS clusters and ECS task state change events, and consumes EC2 Auto Scaling lifecycle hooks for VM fleets. Divergence between orchestrator state and registry state is reported as a reconciliation diff, which is also how registration leaks are found. The registration API is reserved for declared external endpoints and for fleets outside Auto Scaling, authenticated by instance profile or IRSA.
Options weighed
  • ChosenReconcile from orchestrator state; explicit registration for the remainder: Makes a missing instance detectable as a diff rather than invisible, and removes a whole class of startup failure.
  • RejectedSelf-registration by every workload: One mechanism for the whole estate, at the cost of a failure mode where the platform cannot know what it has not been told.
  • RejectedSidecar registers on the application's behalf: Moves the problem into the sidecar without removing it; the sidecar can also fail to register.
  • Right elsewhereOrchestrator-native discovery only, no separate registry: Right for a single-cluster estate; it cannot represent a VM fleet, a managed function or a third-party endpoint, which is most of the hard part here.
Consequences
What it buys
  • A missing instance appears as a reconciliation diff instead of as a quietly smaller fleet.
  • Non-reusable identity makes a stale cached endpoint detectably stale rather than ambiguously reassigned.
  • The registration API's attack surface shrinks to the small non-orchestrated tail.
What it costs
  • One watch integration per runtime type, each with its own failure modes and its own back-pressure behaviour.
  • Reconciliation lag becomes a signal that has to be watched, because it is now the path by which existence is learned.
  • A heterogeneous estate ends up with two registration mechanisms, which has to be documented so callers do not care.
Choose differently when
If the entire estate ran on one orchestrator with one API, reconciliation and native discovery converge and a separate registry earns less. Self-registration becomes the better choice if workloads routinely run somewhere nothing watches — a developer laptop, a customer-managed environment — where there is no authority to reconcile against.
Why it holds up over time
The principle is to prefer the authority that already knows over the party that would have to be asked. That holds as orchestrators change.
LessonIf some other system already holds the fact you need, read it rather than asking the subject to tell you. Self-reported existence cannot distinguish 'not there' from 'never spoke up'.
Shown on views08 14 12
ADR-04

Three independent evidence classes, with passive caller outcomes primary where traffic allows

Accepted

Who should be believed about whether an instance can serve: the instance, a prober, or the callers?

Context
Each source is wrong in a different and characteristic way. A self-report knows the most about internal state and a broken process can report healthy with complete confidence, while a compromised one can simply lie. An external prober is independent, costs N×M traffic at sixty thousand instances, and measures a path no real caller uses — a replica can be perfectly healthy for the prober and broken for everyone else. The outcomes real callers observed are the only evidence that measures what actually matters, cost nothing because the traffic was happening anyway, and go silent on a low-traffic service and cannot distinguish a bad callee from a bad caller. Grey failure — brownout, exhausted pool, long GC pause — is assumed here to be the most common real failure, and it is precisely the class a probe cannot see.
Decision
The platform collects three separately addressable evidence classes — active probe, self-report, and passive caller-observed outcomes — and fuses them into one eligibility decision. The absence of one class is never read as a negative result from another. Where a service's traffic volume makes passive signal statistically sufficient, passive outcomes are the primary evidence and probing drops to a confirmation cadence; below that threshold probing remains primary and the service is reported as having single-class evidence.
How it is realised on AWS
Probers run as a per-AZ DaemonSet with bounded target assignment, so no instance is probed by every prober. Envoy sidecars report per-endpoint outcome counters to a Kinesis-backed ingest at 1.2 M results a second combined. Self-reports ride the Instance API under the workload's own identity. Fusion weights are configured per service, with the traffic threshold for passive-primary declared in the health contract.
Options weighed
  • ChosenAll three classes fused, passive-primary above a traffic threshold: The only arrangement that sees grey failure, survives a prober-network partition, and still covers low-traffic services.
  • RejectedActive probing only: Universal and uniform, and blind to exactly the failure mode that is most common and most user-visible.
  • RejectedSelf-report only: Cheapest and most informed, and unfalsifiable — a confidently wrong process is indistinguishable from a healthy one.
  • RejectedPassive observation only: Free and accurate where traffic exists; a quiet service becomes permanently unknown, which is not an answer.
  • Right elsewhereCaller-side health checking with no central evaluation: Right for a small estate with homogeneous clients; at 450 services it means 450 differing opinions and no shared record.
Consequences
What it buys
  • Grey failure is detectable, which is the difference between this design and one that only finds dead processes.
  • A prober-network partition degrades confidence instead of draining a healthy fleet, because a second class still speaks.
  • Probe traffic stays under 0.5% of internal bytes, because the expensive class is used where the free one cannot reach.
What it costs
  • Three ingest paths, three failure modes, and a fusion function whose weights are another thing to get wrong.
  • Per-service weighting is a configuration surface that most teams will leave at its default, so the default has to be right.
  • Passive signal requires every client library to report outcomes, which re-couples correctness to client fleet management.
Choose differently when
If every service in the estate carried enough traffic for passive signal to be statistically sufficient, active probing becomes a confirmation mechanism that could arguably be dropped. If client libraries cannot be trusted to report outcomes, probe-primary with self-report corroboration is the honest fallback — and grey failure then has to be caught by the callers' own error budgets instead.
Why it holds up over time
The classification of evidence by how it can be wrong is independent of any mechanism. Whatever replaces probes and sidecars, the three ways of being wrong remain three.
LessonWhere a measurement matters, prefer evidence produced by the party that experiences the outcome over evidence produced by the party being measured or by a synthetic observer standing in for them.
Shown on views10 15 05
ADR-05

Unknown is a first-class state, and silence alone never removes an instance

Accepted

What should the platform conclude when an instance stops producing any health signal at all?

Context
A two-valued health model has to map silence onto one of its two values, and both choices are wrong in an important case. Mapping silence to unhealthy means that the first time the observation path breaks — a prober network partition, an ingest outage, a bad agent rollout — the platform concludes that the entire fleet is dead and drains a service that is working perfectly. This is not a theoretical failure: the observation path is a separate system from the thing observed, and it has its own availability. Mapping silence to healthy is the opposite error and keeps dead instances in rotation indefinitely. Neither is a statement about the instance; both are statements about the observer.
Decision
Eligibility has four values: eligible, degraded with a weight, ineligible, and unknown. An instance whose signals have gone silent because the observation path failed is classified unknown and is not removed from rotation on that basis. Moving an unknown instance to ineligible requires corroboration from a second, independent evidence class — most often the caller-observed outcomes, which do not depend on the prober network. Lease expiry remains an independent path to removal for an instance the orchestrator also no longer reports.
How it is realised on AWS
The evaluator distinguishes 'no sample within the contract interval' from 'sample reporting failure', and tracks per-class silence separately. Prober reachability is itself a monitored signal, so a prober-side outage is attributable rather than being inferred from a thousand simultaneously silent instances. The unknown count per service is exported and alertable.
Options weighed
  • ChosenFour-state eligibility with unknown as a distinct value: Separates a statement about the instance from a statement about the observer, which a boolean cannot do.
  • RejectedSilence treated as unhealthy: Conventional, simple, and converts any observation-path failure into a self-inflicted outage of every service at once.
  • RejectedSilence treated as healthy: Safe against observer failure and keeps dead endpoints in rotation until something else notices.
  • DeferredSilence resolved by a quorum of probers: Reduces the chance of correlated observer failure without eliminating it; the second evidence class achieves more for less machinery.
Consequences
What it buys
  • An ingest or prober outage degrades confidence instead of draining the estate, which is the single largest self-inflicted risk this design removes.
  • The unknown count is a direct measure of observation coverage, which turns a blind spot into a reported number.
  • Operators get a state that matches what they actually know, which makes an override decision better informed.
What it costs
  • Unknown instances keep receiving traffic, so a genuinely dead instance may stay in rotation longer than a two-state model would allow.
  • Four states multiply the paths through the evaluator and through every client that interprets a view.
  • The design depends on a second evidence class existing — which is exactly what a low-traffic service lacks.
Choose differently when
If the observation path could be made more available than the workloads it observes — probing from within the same pod, for example, with no network between them — silence becomes genuinely informative about the instance and the unknown state earns less. For a service where serving a dead endpoint is more costly than serving none at all, mapping silence to unhealthy is defensible and should be a per-service choice.
Why it holds up over time
The distinction between 'I know it is broken' and 'I cannot see it' is permanent. Any monitoring or discovery system that collapses the two will eventually drain a healthy fleet.
LessonNever let a system infer a fact about a subject from the failure of its own instrument. Give absence of evidence its own name, and require corroboration before acting on it.
Shown on views15 10 21
ADR-06

Dependency health reduces weight; it may not withdraw an instance over a shared dependency

Accepted

If a service's database is degraded, should the service report itself unready?

Context
Reporting unready is the honest answer and it is what most readiness probes end up doing, because wiring a dependency check into a readiness endpoint is a two-line change that feels responsible. It is correct when the dependency is per-instance: that replica cannot work, others can, route around it. It is catastrophic when the dependency is shared, which it usually is. Every replica checks the same database, every replica fails the same check, every replica withdraws within the same few seconds, and a service that could still serve the eighty per cent of its requests that never touch that database serves nothing at all. The cascade then propagates outward, because the callers of that service now have no healthy endpoints and may withdraw in turn.
Decision
Dependency health is expressible at three authority levels, declared per service in the health contract: informational, weight reduction, or unreadiness. The default is weight reduction. Unreadiness from a dependency failure is permitted only where the dependency is declared non-shared across the service's replicas. The minimum healthy fraction remains the backstop regardless of the declared level, so even a misdeclared shared dependency cannot drain a service.
How it is realised on AWS
The health contract carries a dependency_authority field and a per-dependency shared flag. Self-reports name the dependency and the observed degradation; the evaluator applies the declared authority. A contract change that raises a shared dependency to unreadiness authority is rejected at the policy API rather than at runtime.
Options weighed
  • ChosenWeight reduction by default, unreadiness only for declared non-shared dependencies: Keeps partial function available, and makes the dangerous configuration something a team has to assert rather than inherit.
  • RejectedDependency failure always sets unreadiness: Honest per instance and a correlated-withdrawal generator per service; this is the single most common way a fleet drains itself.
  • RejectedDependency health informational only: Safe against cascades and discards real information about which replicas are least able to serve.
  • DeferredPer-endpoint rather than per-instance health: The right long-term answer — withdraw the routes that need the database, keep the rest — and it needs route-level health in the data plane, which is a larger change than this platform.
Consequences
What it buys
  • A shared-dependency wobble degrades a service rather than removing it, which is almost always what the user would choose.
  • The dangerous configuration is opt-in and policy-checked, so the default is the safe one.
  • Partial function stays available, which matters because most requests to most services do not touch the degraded dependency.
What it costs
  • Capacity stays in rotation that will fail the subset of requests that do touch the dependency, so callers see errors they could have avoided.
  • Teams must classify their dependencies as shared or not, and will sometimes get it wrong.
  • Per-instance granularity is a compromise: the honest unit is the route, and this design does not have it.
Choose differently when
Where a dependency is genuinely per-instance — a local cache, an attached volume, a sidecar — unreadiness is correct and should be declared. Where failing a request is unrecoverable for the user rather than retryable, withdrawing is better than serving, and the authority level should be raised deliberately with the fraction guard understood as the only remaining backstop.
Why it holds up over time
The distinction between shared and per-instance dependency failure is a property of topology, not of technology, and it will decide this question for as long as replicas share anything.
LessonBefore wiring a dependency check into a readiness probe, ask what happens when every replica fails that check at the same instant. If the answer is total unavailability, the check belongs in a weight, not in a boolean.
Shown on views15 21 03

Decision stabilityTurning a stream of noisy samples into a verdict that does not oscillate, and bounding what a wrong verdict costs.

ADR-07

Asymmetric hysteresis over a decaying score, with a fast path for hard failures

Accepted

How many failures should it take to remove an instance, and how many successes to let it back?

Context
Every damping mechanism buys stability by paying detection latency, and the exchange rate is not obvious. Consecutive-failure thresholds are simple and multiply the detection budget by the threshold: a three-strike rule on a one-second probe makes a three-second detection a nine-second one. A decaying score responds to failure rate rather than to the last sample, which is closer to what matters and introduces tuning that nobody can reason about at three in the morning. Symmetric thresholds are the worst of the options: an instance that is marginal will oscillate, and an oscillating instance destabilises every client's view, which is worse than either verdict being applied consistently.
Decision
Eligibility is computed from a decaying score per instance per evidence class, with asymmetric thresholds: re-entry is strictly slower than removal. A hard failure — connection refused, process gone, lease expired — bypasses the decay entirely and removes immediately, because damping exists to absorb ambiguity and a closed port is not ambiguous. Damping is bounded: it may add no more than 5 s to the p99 detection budget for a genuine failure, and that bound is the constraint the score's parameters are fitted to.
How it is realised on AWS
The evaluator holds an exponentially weighted failure rate per instance per class, with configurable half-life, and a separate hard-failure channel that short-circuits to ineligible. Thresholds live in the health contract with platform-enforced floors. The 5 s damping budget is measured as the difference between first-failed-sample and published-withdrawal for the hard and soft paths respectively, and reported per service.
Options weighed
  • ChosenDecaying score, asymmetric thresholds, hard-failure fast path: Responds to rate rather than to the last sample, and does not make an obviously dead instance wait for statistics.
  • RejectedConsecutive-failure thresholds only: Trivial to explain and multiplies the detection budget by the threshold, which breaks the 10 s p99 commitment.
  • RejectedSymmetric hysteresis: Simpler to configure and guarantees oscillation on marginal instances, which is the failure this is meant to prevent.
  • DeferredMachine-learned anomaly scoring over the sample stream: Might catch failure patterns a rate misses; an unexplainable verdict in a routing decision is a worse trade than a missed subtlety.
Consequences
What it buys
  • Marginal instances stop oscillating, so client views stay stable and the delta churn volume stays bounded.
  • A dead process is removed in 10 s p99 regardless of the damping configuration, because it never enters the damped path.
  • The damping cost is an explicit, measured 5 s rather than an emergent property of three thresholds multiplying.
What it costs
  • The score's half-life is a parameter that few teams will understand, so the platform default carries most of the weight.
  • A slowly-degrading instance crosses the threshold later than a consecutive-failure rule would, by design.
  • Two paths through the evaluator — damped and fast — is two code paths that must agree about everything else.
Choose differently when
If the dominant failure mode in the estate turned out to be hard rather than grey, the decaying score earns little and consecutive-failure thresholds would be the simpler, more explicable choice. If detection latency ever becomes more valuable than stability for a specific service — a payment path, say — the thresholds should be tightened for that service and the resulting flapping accepted and quarantined.
Why it holds up over time
Asymmetry between removal and re-entry is a permanent property of any system where a wrong removal is cheaper to reverse than a wrong retention is to discover. The specific scoring function is replaceable.
LessonMake damping an explicit, measured cost against a detection budget rather than an implicit consequence of thresholds. And never make an obviously dead thing wait for statistical confidence.
Shown on views15 18 17
ADR-08

The minimum healthy fraction disregards health rather than draining a service

Accepted

What should happen when most of a service's replicas report unhealthy at the same moment?

Context
Correlated mass withdrawal is the characteristic failure of a competent health-checking system. It arrives from a bad deploy, a shared dependency outage, or — most often — a health-check configuration change that is wrong in a way nobody noticed, because health-check config is rarely treated as production code. The platform then does exactly what it was asked to do, removes almost every replica, and the service's availability goes to zero because its health checking worked. Honouring every withdrawal is the honest behaviour and it means a two-line YAML change can take down a service more reliably than deleting its deployment.
Decision
Each service carries a minimum healthy fraction per zone, defaulting to 50%. When the eligible fraction would fall below it, the platform disregards health and returns the full endpoint set, declares in the response that it has done so, and raises it as an alert. A removal that would take the last eligible instance of a service in a region raises an incident-grade signal rather than completing silently. The guard sits downstream of evaluation, so a bug in scoring or fusion cannot bypass it.
How it is realised on AWS
The guard is evaluated in the view assembler, per service per zone, after eligibility and before delta encoding. Views assembled under the guard carry a degraded-confidence marker that clients surface in their own telemetry, and the fraction-guard trip is one of the platform's four alarms.
Options weighed
  • ChosenDisregard health below a per-service, per-zone fraction, and declare it: Converts a self-inflicted outage into a declared degradation, which is nearly always the better failure.
  • RejectedHonour every withdrawal: Maximally honest and makes the health-checking system the most effective outage generator in the estate.
  • RejectedA single global fraction for all services: One number to reason about, and wrong for both a three-replica service and a three-thousand-replica one.
  • Right elsewhereRequire a human decision above a blast-radius threshold: Right where serving a known-bad instance is unsafe — a payment or data-mutating path — and wrong at three in the morning for a read path.
  • DeferredGraduated response: reduce weights before disregarding health: A smoother curve and worth having; the cliff is easier to reason about during an incident, which is when it fires.
Consequences
What it buys
  • A bad health-check rollout degrades a service instead of removing it, which is the most likely mass-withdrawal cause.
  • The declaration means callers and operators can tell 'the fleet is broken' from 'the health check is broken'.
  • The guard's position downstream of evaluation means it survives bugs in evaluation.
What it costs
  • Traffic is sent to instances known to be failing, which is worse than a clean failure if those instances corrupt data or hold connections.
  • A genuinely broken fleet is kept in rotation, so the user sees errors rather than a fast failure.
  • 50% is a cliff, and a service sitting near it will trip and untrip in ways that are confusing to watch.
Choose differently when
For a service where a failing request is unrecoverable for the user or leaves inconsistent state, draining is better than serving and the fraction should be lowered or the guard replaced by a human gate. If the platform could reliably distinguish a broken fleet from a broken health check, it could choose differently in each case and the blunt guard would no longer be needed.
Why it holds up over time
The principle — no single signal may remove a service's capacity — outlives any particular threshold, and applies to every control system with an actuator larger than its sensor's reliability.
LessonBound what your control loop is allowed to do, not just what it is likely to decide. A correct system executing a wrong instruction is the failure mode that most often reaches users.
Shown on views15 21 06
ADR-09

A flapping instance is quarantined with escalating backoff, and its owner is told

Accepted

What should happen to an instance that keeps changing its mind about whether it is healthy?

Context
Hysteresis reduces oscillation; it does not eliminate it. An instance that fails under load, recovers when traffic is withdrawn, and fails again when traffic returns will oscillate indefinitely, and each transition is a delta pushed to every subscribed client for that service. At sixty thousand subscriptions, a handful of flapping instances across the estate is a measurable fraction of propagation capacity spent on noise, and the churn makes every client's view less stable than either verdict would be. The instance is also, almost always, a bug report: either its health check is wrong or it genuinely cannot sustain its share, and both are facts its owner needs.
Decision
An instance that changes eligibility more than four times in ten minutes is quarantined: held ineligible for an exponentially increasing period, starting at one minute and capped at thirty. Quarantine is surfaced to the owning team as an event with the transition history that produced it, and appears in the service's flap report. Quarantine never applies to a whole service: the fraction guard takes precedence, so quarantining cannot be a path to draining a fleet.
How it is realised on AWS
The evaluator counts transitions per instance over a sliding ten-minute window and writes a quarantine_until timestamp on the eligibility row. Quarantine events are published to the owning team's notification channel and aggregated into a per-service flap rate, which is a reported signal rather than an alarm.
Options weighed
  • ChosenEscalating quarantine with owner notification: Stops the churn, and treats the flap as the defect report it usually is rather than hiding it.
  • RejectedLonger hysteresis instead of quarantine: Suppresses the symptom across the board and pays detection latency on every instance to fix a problem a few have.
  • RejectedQuarantine with manual release only: Safest against premature return and turns every flap into human work at whatever hour it happens.
  • Right elsewherePermanent removal after repeated flapping: Right where instances are cheap and immediately replaceable; here it risks removing capacity the service needs for a condition that may be load-induced.
Consequences
What it buys
  • Delta churn from noisy instances is bounded, which protects propagation capacity for the changes that matter.
  • Client views stay stable, so a flapping instance stops being everyone's problem.
  • The owning team receives a specific, evidenced bug report instead of a vague reliability complaint.
What it costs
  • A genuinely recovered instance can be held out for up to thirty minutes, which is lost capacity.
  • Quarantine can mask a load-induced failure as an instance defect, sending the owner looking in the wrong place.
  • Another piece of per-instance state to hold, report and reason about during an incident.
Choose differently when
If capacity were tight enough that holding a recovered instance out for minutes mattered more than view stability, the cap should fall and the escalation flatten. In an estate where instances are replaced in seconds, terminating the flapping instance and letting the orchestrator supply a new one is simpler and better.
Why it holds up over time
Quarantine of an oscillating participant is a standard answer in any system that aggregates noisy votes into a shared decision, and it will outlive this platform's mechanisms.
LessonWhen a component's instability becomes everyone else's instability, isolate it and raise it as a defect. Suppressing oscillation globally to accommodate a few participants taxes all of them.
Shown on views15 18 17

Propagation and convergenceGetting a changed view to sixty thousand clients, and the asymmetry between removing and adding.

ADR-10

Streaming subscriptions are the primary surface; DNS is a compatibility surface

Accepted

How does a changed view reach sixty thousand clients inside five seconds?

Context
Three mechanisms are available and they are not close substitutes. Short-TTL DNS is universally consumable, needs no client library, and is bounded by resolver behaviour nobody controls — including resolvers, language runtimes and sidecars that cache beyond the TTL or ignore it entirely. A five-second withdrawal budget cannot be held by a mechanism whose propagation time is a property of other people's software. Streaming subscriptions give sub-second push with exact version tracking, and require a stateful tier holding sixty thousand long-lived connections, which must survive its own deploys and becomes the most delicate routine operation in the platform. Gossip removes the central tier and makes convergence time and message ordering probabilistic, which is difficult to reason about precisely when reasoning matters most.
Decision
Incremental streaming subscriptions are the primary propagation surface, carrying deltas with monotonically increasing view versions. A DNS zone is maintained as a compatibility surface with identical eligibility semantics, for clients that cannot hold a subscription, and its weaker propagation guarantee is stated rather than hidden. A request/response resolution API exists for tooling and target-group synchronisation and is not for request-path use. Gossip is not used.
How it is realised on AWS
An Envoy-compatible xDS tier on EKS holds the subscriptions, sending incremental resource updates with a version per service and falling back to a full view only when a client's version is too old to reconcile. Cloud Map and Route 53 are populated from the same assembled views on a short TTL. Clients re-subscribe with jitter capped at 5% of a population per 10 s window, and the tier drains connections on its own rollout.
Options weighed
  • ChosenStreaming primary, DNS as compatibility, API for tooling: The only option that holds a 5 s budget while still serving clients that cannot subscribe.
  • RejectedDNS everywhere with aggressive TTLs: No client library to ship, and the propagation guarantee belongs to resolvers the platform does not operate.
  • RejectedGossip between data planes: Removes the stateful tier, and replaces an explainable latency with a probabilistic one and no version ordering.
  • RejectedLong-poll or periodic full-view fetch: Simple and stateless on the server; at 60,000 clients and a 5 s budget it is a continuous full-view fan-out.
  • Right elsewhereGossip within a zone, streaming across zones: Attractive where zone-local convergence dominates and the client fleet is homogeneous enough to trust with a membership protocol.
Consequences
What it buys
  • Exact per-client version tracking, which is what makes monotonic application and incremental deltas possible at all.
  • Withdrawal propagation is a property of the platform rather than of every client's resolver.
  • DNS clients are still served, with a stated weaker guarantee rather than a silent one.
What it costs
  • A stateful tier holding 60,000 connections is the platform's most dangerous routine deploy, mitigated by jitter and the static window rather than removed.
  • A client library is now required for the primary path, which re-couples correctness to fleet management.
  • Two surfaces with different propagation guarantees must be kept semantically identical, which is a continuing test burden.
Choose differently when
If the client fleet were homogeneous and small enough that a full-view fetch per interval were affordable, the stateful tier is unnecessary complexity. If the withdrawal budget could be relaxed to tens of seconds — because local ejection carries the fast path entirely — DNS becomes sufficient and the whole tier can go.
Why it holds up over time
xDS is a current answer to a permanent question. What should outlive it is the requirement for versioned, incremental, monotonic delivery, which any replacement must also provide.
LessonDo not accept a propagation guarantee that depends on software you do not operate. If you have a latency budget, you need a mechanism whose latency you own.
Shown on views09 08 13
ADR-11

Withdrawal and admission are separately budgeted, and admission is the path that sheds

Accepted

When propagation capacity runs short, which changes get delayed?

Context
A single propagation path with a single latency target has to be sized for the worst case, which here is a regional evacuation generating forty thousand instance state changes a minute against a steady state of two and a half thousand. Sizing for sixteen times steady state is expensive and still arbitrary. But the two kinds of change are not equally urgent, and the asymmetry is large: a withdrawal that arrives late means traffic continues to a failing endpoint and users see errors, while an addition that arrives late means the service runs at slightly less capacity than it has. Treating them identically spends capacity on the cheaper one at precisely the moment the expensive one matters.
Decision
Withdrawal and admission are separately budgeted: 99% of callers see an unplanned withdrawal within 5 s and 99.9% within 15 s, while a newly eligible instance need only reach 99% of callers within 30 s. Under saturation the platform lengthens the batching interval for additions and holds the withdrawal path at its budget, and declares that it has done so. Views are monotonic per client, so a delayed addition can never be applied after a later view that already contains it.
How it is realised on AWS
The view assembler maintains two delta queues per service with different batching windows, and the xDS tier drains the withdrawal queue first. Shed episodes are a reported signal. Version comparability across control-plane replicas comes from a region-wide monotonic counter, so a client that reconnects to a different replica cannot regress.
Options weighed
  • ChosenSeparate budgets, shed additions, hold withdrawals: Spends scarce propagation capacity on the change whose lateness users feel.
  • RejectedOne propagation path, one budget: Simpler to build and operate, and must be sized for a 16× burst to protect the urgent case.
  • RejectedPriority by service tier rather than by change type: Protects important services and lets a tier-1 addition delay a tier-2 withdrawal, which is the wrong axis.
  • DeferredPush withdrawals, poll for additions: A cleaner expression of the same asymmetry; it means two mechanisms to keep semantically identical for no gain the queues do not already give.
Consequences
What it buys
  • A 16× churn burst becomes a capacity question about additions rather than a correctness question about withdrawals.
  • The platform can be sized for steady-state additions and burst withdrawals, which is materially cheaper.
  • Monotonic versions mean shedding cannot produce an out-of-order view, only a later one.
What it costs
  • During a burst a service runs with less capacity visible to callers than it actually has, for up to the shed interval.
  • Two queues with different behaviour is a subtlety that shows up in every propagation-lag investigation.
  • The admission budget is deliberately loose, which makes slow start's effectiveness partly dependent on propagation timing.
Choose differently when
If adding capacity late were itself a user-visible failure — an autoscaling-driven service where new capacity is the response to load — the asymmetry narrows and both paths need tight budgets. The decision is also wrong if local ejection were strong enough to make withdrawal propagation uncritical, in which case both budgets could relax.
Why it holds up over time
The asymmetry between removing and adding is a property of what users experience, not of any transport, and it holds for any system that publishes membership.
LessonWhen two kinds of update share a channel, ask which one's lateness the user feels. Then give them different budgets and make the cheap one yield.
Shown on views10 13 06

RecoveryThe decisions that treat coming back as more dangerous than going down.

ADR-12

Slow start is published by the control plane as a rising weight

Accepted

Who decides how much traffic a newly admitted instance receives in its first minute?

Context
A cold replica has empty caches, unwarmed connection pools and a just-in-time compiler that has not seen production traffic. Giving it an equal share of a service's load the instant it passes readiness is a reliable way to make it fail readiness again, which is the thundering-herd failure in its smallest form. The ramp can live in two places. In the client, it is immediate and proportional to that client's own traffic, and it is only as correct as the least-updated client library in the estate — which, across 450 services and many languages, means it is absent somewhere. In the control plane, published as a rising weight, it is uniform across every client and every heterogeneous data plane, and it is only as smooth as the push interval.
Decision
Slow start is owned by the control plane and expressed as a weight that rises as a function of time in rotation and observed success rate, over an interval of at least 60 s. Clients apply the published weight; they do not compute their own ramp. The same mechanism admits a new deployment version, a recovered instance and a newly started instance, so there is one ramp behaviour rather than three.
How it is realised on AWS
Eligibility rows carry a weight that the evaluator raises on a schedule after admission, and the view assembler publishes weight changes as ordinary deltas. Because admission propagation is budgeted at 30 s, the ramp is published in steps rather than continuously — which is accepted, since the ramp's purpose is to bound the first minute, not to be smooth.
Options weighed
  • ChosenControl-plane-published weights: Uniform across a heterogeneous client fleet, which matters more here than proportional smoothness.
  • RejectedClient-side ramp from a declared start time: Immediate and traffic-proportional, and correct only where the client library is current — which cannot be assumed.
  • DeferredControl plane publishes the policy, clients execute it: The best of both and it still depends on client libraries implementing it; worth revisiting once version spread is under control.
  • Right elsewhereNo ramp; rely on readiness to gate admission: Right for stateless services with no warm-up cost, and wrong for anything holding a cache or a pool.
Consequences
What it buys
  • Every client ramps identically, including DNS clients and target-group-routed traffic that have no ramp logic at all.
  • One mechanism covers deploys, restarts and recoveries, so there is one behaviour to understand and test.
  • The ramp is visible in the published view, which makes it debuggable from the caller's side.
What it costs
  • The ramp is as coarse as the admission propagation interval, so it is a staircase rather than a curve.
  • Weight changes consume propagation capacity for every ramping instance, which at ~900 deploys a day is continuous traffic.
  • A service whose warm-up exceeds 60 s is still hurt, and few teams will tune the interval.
Choose differently when
Once client library versions are uniform enough to trust, moving the ramp into the client removes the propagation cost and makes it proportional to actual load, which is strictly better. For a genuinely stateless service the ramp is pure cost and should be configurable to zero.
Why it holds up over time
The decision is really about where to put behaviour when client heterogeneity is the binding constraint, and that constraint recurs in every platform with a client library.
LessonWhen a safety behaviour must hold everywhere, put it where you control it — even at the cost of doing it less well than the place that has better information.
Shown on views04 14 18
ADR-13

Recovery is rate-capped in three independent places

Accepted

What stops a recovered component from being destroyed by the traffic that was waiting for it?

Context
Recovery is the more dangerous event, and the reason is arithmetic rather than subtle. While something is down, demand does not disappear: clients retry, queues fill, and sixty thousand subscriptions wait to reconnect. The instant the thing recovers, all of that arrives simultaneously against a component that is cold. A recovered availability zone receives the traffic of the two zones that were carrying it. A restarted propagation tier receives every client's re-subscription at once. A recovered instance receives its full share before its first cache is warm. Each of these is a different mechanism arriving at the same failure, which is why one cap is not enough.
Decision
Three independent rate caps are platform defaults rather than tuning. No more than 5% of a client population may reconnect to the control plane in any 10 s window, enforced by jittered randomised backoff. A recovered or new instance reaches full weight over at least 60 s. Traffic shifted back into a recovered zone or region is rate-capped, and where the failover was declared manually the shift-back must be explicitly completed rather than happening automatically.
How it is realised on AWS
The xDS tier rejects re-subscriptions above its admission rate with a retry hint, and client libraries apply full jitter on backoff. Weight ramps come from ADR-12. Zone and region shift-back is a policy-API operation that moves topology preference in capped steps, with the final step requiring confirmation for a manual failover.
Options weighed
  • ChosenThree independent caps as defaults: Each mechanism arrives at the same failure by a different route, so each needs its own ceiling.
  • RejectedOne global admission control at the control plane: One place to reason about, and it cannot cap traffic shifting between zones, which never touches the control plane.
  • RejectedClient-side backoff only: Standard practice and dependent on every client library doing it correctly, which is the assumption this design avoids making.
  • RejectedAutomatic shift-back on recovery: Faster return to full capacity and removes the human judgement that the recovery is real.
  • DeferredQueue-based admission with explicit draining of the backlog: More precise than rate caps for the reconnection case; the caps are sufficient and much simpler to operate.
Consequences
What it buys
  • A control-plane restart does not become a company-wide reconnection storm, which makes routine deploys of the tier survivable.
  • A recovered zone is reloaded gradually, so the recovery does not immediately produce a second incident.
  • Manual failovers end with a human confirming the return, which is where the judgement belongs.
What it costs
  • Full recovery takes longer than it needs to when the recovered component is genuinely healthy.
  • A manual shift-back that nobody completes leaves a region drained indefinitely, which needs its own alert.
  • Three caps in three places is three sets of parameters, and only one of them is visible from any single view.
Choose differently when
If recovery capacity were provably abundant — a stateless service with no warm-up behind an elastic pool — the caps cost availability for no benefit. Automatic shift-back is correct where failover is itself automatic and cheap to reverse, because then no human was in the loop to confirm anything.
Why it holds up over time
That recovery concentrates demand is a permanent consequence of queueing, and any replacement for these mechanisms will need its own ceilings.
LessonDesign the recovery path before the failure path. Failure sheds load; recovery concentrates it, and concentrated load against a cold component is how one incident becomes two.
Shown on views06 14 18

Trust and heterogeneityWho is allowed to deny service to whom, and how runtimes that cannot answer the question are represented.

ADR-14

Regional registry authority, with a failover signal that does not depend on the registry

Accepted

Should there be one global registry, or one per region with federation between them?

Context
A single global registry gives one consistent answer, makes cross-region resolution trivial, and makes the global control plane a dependency of every region's ability to route — which contradicts ADR-01 at the regional scale instead of the request scale. Independent regional registries with asynchronous replication keep each region alive through a partition, and guarantee that the regions will sometimes disagree about which region is healthy. That disagreement arrives precisely when the answer matters most, during a partition, and a system that asks the registry whether a region is reachable is asking the wrong component: the registry's own reachability is the thing in question.
Decision
Each region operates an independent control plane with its own registry authority, able to serve all of its own resolution with no dependency on a global component. Cross-region resolution is served from an asynchronously replicated, read-only aggregate of the peers, and replication loss degrades cross-region resolution only. The signal that declares a region unhealthy is out of band and has no dependency on the registry it describes. Reconciliation after a partition is explicit and logged, never derived from timestamps.
How it is realised on AWS
Three independent regional stacks in us-east-1, eu-west-1 and ap-south-1, each with its own Aurora registry and its own control plane on EKS. Cross-region state replicates asynchronously into read-only aggregates. Route 53 Application Recovery Controller readiness checks and routing controls provide the out-of-band region health and failover signal, deliberately chosen because its data plane is separate from the platform's own.
Options weighed
  • ChosenRegional authority with read-only cross-region aggregates and an out-of-band failover signal: A region routes with every peer and every global component unreachable, which is the property that matters.
  • RejectedSingle global registry with regional read caches: One consistent answer, and the global plane becomes a dependency of every region's routing.
  • RejectedRegional authority with a global read-only aggregate as the single cross-region view: Nearly the chosen design, with one global component whose loss removes all cross-region resolution at once.
  • RejectedStrongly consistent multi-region registry: Removes disagreement entirely and makes every registry write pay a cross-region round trip, and a partition stop writes.
  • Right elsewhereSingle-region control plane for the whole estate: Right for an estate genuinely in one region; here it makes one region's bad day global.
Consequences
What it buys
  • A regional partition is a cross-region resolution degradation rather than a routing outage anywhere.
  • The failover decision does not depend on the system whose reachability is in question.
  • Each region's control plane can be deployed and drilled independently, which is how the static-serving claim gets tested.
What it costs
  • Regions will disagree about reality after a partition, and the reconciliation is explicit work rather than an automatic merge.
  • Three stacks to operate, upgrade and keep semantically identical.
  • Cross-region resolution is always slightly stale, so a cross-region failover routes on an older view than a local one.
Choose differently when
If the estate's traffic were genuinely global rather than regionally served — every request touching services in two regions — the staleness of a cross-region aggregate becomes a correctness problem and strong consistency starts to earn its cost. A single region of operation makes the whole federation question moot.
Why it holds up over time
Not asking a component about its own reachability is a permanent rule. The specific choice of regions and replication mechanism is not.
LessonNever let a system be the source of truth about whether it can be reached. Put the reachability signal somewhere with a different failure domain, even if it is less convenient.
Shown on views16 06 21
ADR-15

The power to deny service is a privilege, and every signal must be attributable

Accepted

Who is allowed to mark an instance unhealthy, and what stops that being an attack?

Context
The discovery plane's central capability is removing endpoints from rotation. Stated plainly, it is an authorised denial-of-service mechanism operating continuously across the entire estate, and most implementations treat its inputs as telemetry rather than as security-relevant assertions. If any workload can report health about any instance, then a single compromised pod can remove a service it merely calls. If an operator can always force an instance ineligible, the break-glass path is a denial-of-service path with a badge. And the one evidence class that cannot be forged about oneself — the outcomes callers observed — can still be forged about somebody else.
Decision
Every registration, health report and resolution request is authenticated with workload identity rather than network position. An instance may report health only about itself; a prober only about the instances assigned to it. A signal that cannot be attributed to a registered identity is discarded, not down-weighted. Marking an instance ineligible by hand is privileged, rate-limited per actor per service, requires a justification, and is audited whether or not it succeeds. A registered address must be proved to belong to the identity registering it. Resolution is authorised per caller-callee pair, so the registry is not a map of the estate for any compromised workload.
How it is realised on AWS
IRSA and instance profiles supply workload identity; all control-plane paths use mutual authentication. Prober assignment is held in the registry, so a prober report outside its assignment is rejected at ingest. Overrides go through the policy API, which checks authorisation and the disruption budget before anything is applied, and writes the attempt to the 5-year immutable audit store. Caller outcome reports are weighted by reporter population rather than by report volume, so one noisy reporter cannot dominate.
Options weighed
  • ChosenIdentity-attributed signals, self-only reporting, privileged and audited overrides: Treats the removal capability as what it is, and makes the common attack — report about a neighbour — structurally impossible.
  • RejectedTrust signals from any authenticated workload in the mesh: Simple, and gives every workload the ability to deny service to every other.
  • RejectedNetwork-position-based trust: Requires no identity plumbing and fails completely the first time anything inside the perimeter is compromised.
  • RejectedUnrestricted operator override: Fast under pressure and makes the break-glass credential the most powerful outage tool in the company.
  • DeferredSigned health assertions with per-report non-repudiation: Stronger than identity-scoped attribution, and the per-signal cost at 1.2 M results a second does not currently buy enough.
Consequences
What it buys
  • A compromised workload cannot remove a service it merely calls, which is the attack this capability otherwise invites.
  • Every denial is attributable after the fact, including the ones that were refused.
  • Per-pair resolution authorisation means a compromised workload cannot enumerate the estate.
What it costs
  • Identity plumbing on every path, including the high-volume passive ingest, where it is the per-sample cost.
  • Prober assignment becomes registry state that must be correct, or legitimate probe reports are rejected.
  • A refused override under incident pressure is a frustrating experience, and the refusal has to explain itself well to be accepted.
Choose differently when
In a single-tenant estate with no meaningful internal threat model, per-pair resolution authorisation is overhead and a flat mesh identity would do. If incident response ever showed that refused overrides were materially extending outages, the budget check should be made overridable with a second authoriser rather than removed.
Why it holds up over time
Reasoning about a capability by what it can destroy rather than by what it is called is a permanent discipline, and it survives every change of identity technology.
LessonName your system's capabilities by their effect. A health-check input is an authorisation to deny service, and it should be secured like one.
Shown on views19 20 10
ADR-16

One registry over heterogeneous runtimes, with capability declared per service

Accepted

How should workloads that cannot answer a readiness question be represented?

Context
The estate is not uniform and will not become uniform. Alongside twelve orchestrated clusters there are long-lived VM fleets, managed functions behind an invoke path, remaining data-centre services, and third-party endpoints that publish no health signal of any kind. Forcing all of them into one abstraction gives callers a single resolution surface and one mental model, at the price of an abstraction whose health contract is only as rich as its weakest member. Giving the heterogeneous tail a second mechanism keeps the orchestrated path clean and forces every caller to know how its callee is deployed, which is exactly the coupling a discovery plane exists to remove.
Decision
All workloads are services in one registry with one resolution surface. Each service declares a capability profile stating which health signals its runtime can actually supply and which are unavailable, and the platform never reduces a richer service's contract to the weakest profile in the estate. A service with no readiness signal is represented as eligible with health unknown, stated in the view rather than implied, and is governed by passive caller outcomes alone.
How it is realised on AWS
The service row carries a capability_profile; the health contract's required fields are validated against it, so a contract demanding readiness from a third-party endpoint is rejected at authoring time. Views carry the capability per endpoint so a client can tell an unprobed endpoint from a probed healthy one. Non-orchestrated fleets renew leases; orchestrated ones do not need to.
Options weighed
  • ChosenOne registry, capability declared per service: Callers get one surface and the weak cases stay honest rather than being hidden behind a uniform contract.
  • RejectedOne registry with a synthetic health adapter per runtime: Makes every endpoint look probed, which is a lie the caller cannot detect and will route on.
  • RejectedSeparate external-dependency registry with its own semantics: Keeps the orchestrated path pristine and makes every caller know its callee's deployment model.
  • RejectedLowest-common-denominator contract for the whole estate: Uniform and simple, and discards the readiness and dependency signals that most of the estate can actually supply.
  • Right elsewhereOrchestrator-native discovery plus a side channel for the rest: Right where the non-orchestrated tail is small and shrinking; here it is neither.
Consequences
What it buys
  • A caller resolves by name without knowing whether its callee is a pod, a fleet instance or somebody else's endpoint.
  • Capability is visible in the view, so 'health unknown' is a fact the client can act on rather than an assumption.
  • The orchestrated majority keeps its full contract instead of being levelled down.
What it costs
  • Clients must handle an endpoint whose health is unknown, and some will route to it as though it were verified.
  • Capability profiles are configuration that can be wrong, and a wrong profile silently weakens or over-promises a contract.
  • Four registration paths behind one surface is more platform code than a single-runtime design needs.
Choose differently when
If the non-orchestrated tail shrank to nothing, orchestrator-native discovery becomes sufficient and this registry is a layer with no job. If callers proved unable to handle unknown-health endpoints correctly, a separate registry for them — forcing an explicit decision at the call site — would be the safer design despite the coupling.
Why it holds up over time
Declaring capability rather than assuming it is how any abstraction over heterogeneous providers stays honest, and that outlives the particular runtimes.
LessonWhen unifying things that are not the same, make the differences declarable and visible. An abstraction that hides what it cannot deliver moves the failure to the caller, who has less information than you did.
Shown on views14 01 09

Every package used, in one table

These terms are used precisely in this package. Several are used loosely in the wider literature on service discovery and health checking, and the loose readings are what make two discovery designs disagree with each other while appearing to say the same thing.

PackageWhat it isWhat it does hereConsidered instead
Liveness Whether the process is running and its port is open. Decides whether to restart an instance; it is the orchestrator's question, not the router's. "Health", which collapses it with readiness and makes a restarting replica indistinguishable from a busy one.
Readiness The instance's own declaration that it is willing to receive traffic now. The instance's vote in its own eligibility; one of three evidence classes. "Up", which hides that the instance is asserting something it may be wrong about.
Dependency health The instance's view of the downstreams it needs in order to do useful work. Reduces weight by default; may set unreadiness only for a declared non-shared dependency. Folding it into readiness, which is how a shared database wobble becomes a total outage.
Eligibility The platform's single verdict per instance: eligible, degraded with a weight, ineligible, or unknown. The only thing published to clients; everything upstream is evidence for it. "Healthy/unhealthy", which has no way to say that the observer is the broken part.
Unknown The state of an instance whose evidence has gone silent because the observation path failed. Keeps the instance in rotation and requires corroboration before removal. Treating silence as unhealthy, which drains a healthy fleet the first time a prober network partitions.
View A per-service, versioned endpoint set with weights, zones, versions and topology priority tiers. The unit of publication, of monotonic application at the client, and of staleness. "The endpoint list", which hides the version and therefore hides how old it is.
Last-known-good cache The durable per-client copy of the most recent view it successfully applied. The authority for routing, and the reason a control-plane outage is a staleness event. "Cache", which implies an optimisation rather than the system of record for routing.
Fail static Continuing to act on the last known state when new information cannot be obtained. The platform's required behaviour on control-plane unavailability. "Fail open" and "fail closed", both of which describe a decision rather than the refusal to make a new one.
Withdrawal An instance becoming ineligible, and the propagation of that fact to callers. The urgent path, budgeted at 5 s to 99% of callers and never shed. "Deregistration", which is a different event: the instance ceasing to exist.
Minimum healthy fraction The proportion of a service's instances per zone below which health is disregarded. The backstop that turns correlated mass withdrawal into a declared degradation. "Panic threshold", which describes the mechanism's mood rather than its contract.
Flapping An instance oscillating between eligibility states faster than its damping absorbs. Triggers escalating quarantine and an owner-visible defect report. "Noise", which suggests the right response is more filtering rather than isolation.
Slow start A rising published weight that brings a newly admitted instance to full share over at least 60 s. The ramp that stops a cold replica being killed by its own admission. "Warm-up", which usually names an application behaviour rather than a routing weight.
Outlier ejection A client removing an endpoint that is failing for that client specifically. The fast path: it acts before the control plane has an opinion, under a per-client cap. "Circuit breaking", which is about protecting the caller rather than correcting the endpoint set.
Failure domain A region, zone or cluster, modelled explicitly so correlated failure is expressible. Carries topology priority tiers and makes evacuation a one-step preference shift. "Location", which carries no statement about what fails together.
Capability profile A per-service declaration of which health signals its runtime can actually supply. Keeps a third-party endpoint honest without levelling the whole estate to its contract. Assuming a uniform health contract, which makes an unprobed endpoint look verified.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.