pattern

Service Discovery

How a caller finds a healthy instance of a callee in an environment where instances appear and disappear continuously.

service-discoveryload-balancingnetflixdnsresilience

Definition

In a fleet where instances are created, destroyed, rescheduled and replaced continuously, hardcoded addresses are unusable. Service discovery is the mechanism that answers "where can I reach the payments service right now, and which of those addresses are healthy?"

Two shapes:

  • Client-side discovery. The caller queries a registry, receives the instance list, and chooses one itself. Netflix's approach — a registry plus a client-side load balancer — belongs here. The client can then make sophisticated choices: prefer the same availability zone, avoid instances with rising latency, weight by observed performance.
  • Server-side discovery. The caller sends to a stable address — a load balancer, a mesh sidecar, a Kubernetes Service — which resolves and balances. Simpler clients, one more hop, and the balancer is now on the critical path.

Why it matters

Discovery is where several distinct concerns quietly meet: naming, health checking, load balancing, and failure detection. Getting it wrong produces failures that look like application bugs — traffic sent to instances that are shutting down, or a thundering herd when a registry entry expires.

Industry example

Netflix's client-side registry plus client-side balancing was chosen for a specific reason: resilience through stale-but-usable data. Clients cache the instance list. If the registry itself becomes unavailable, callers keep using their last known list rather than failing. The registry is deliberately AP rather than CP — it prefers returning possibly-stale instance information over returning nothing.

That is the right trade for this problem. A slightly outdated list means some calls hit a dead instance and fail fast, which the caller's retry handles. An empty list means total outage. Many teams get this backwards and build a strongly consistent registry that becomes the single point of failure for everything.

The cost of the choice is real: clients are thicker, discovery logic must exist in every language the estate uses, and rolling out a fix to client behaviour requires redeploying every caller. That last point is what drove much of the industry toward sidecar-based meshes, which move the same logic out of the application and into a proxy that can be upgraded independently.

Implementation patterns

  • Registry with heartbeats and TTL. Instances register and renew; entries expire if renewal stops.
  • Platform-native discovery. Kubernetes Services and DNS. Simplest when you are already there; beware DNS caching in application runtimes, which is a recurring source of traffic sent to dead addresses.
  • Sidecar proxy / mesh. Discovery, balancing, retries, timeouts and mTLS handled outside the application process.
  • Graceful shutdown protocol. Deregister, wait for in-flight requests and for caches to expire, then stop accepting connections. Skipping the wait causes a burst of errors on every deploy.

Failure scenarios

  • A strongly consistent registry that goes down and takes the whole estate with it.
  • DNS TTLs ignored by the runtime, so instances that were removed keep receiving traffic for minutes or forever.
  • Health checks that test dependencies. An instance reports unhealthy because a downstream is slow; the platform removes it; capacity drops; the remaining instances get more load and also report unhealthy. The dependency's slowdown becomes your outage. Liveness should test the process; readiness may test dependencies, and the two must be different endpoints.
  • Registration without deregistration on crash, leaving ghost entries until TTL.

Trade-offs

Client-side gives better balancing decisions and resilience to registry failure, at the cost of client complexity and lockstep upgrades. Server-side and mesh give thin clients and central control, at the cost of an extra hop and an extra thing to operate. Neither is wrong; the estate's language diversity usually decides it.

Interview question

"Your service registry has been unavailable for ten minutes. What should be happening to traffic, and what would you check first?"