intermediate 3 min answer

A ride-hailing company moving off its monolith found engineers actively avoiding service calls because they could not see the latency or the failures those calls produced. Lyft's published answer was Envoy, open-sourced in September 2016: a separate process on every host that all service traffic flows through. Which concern was actually separated, what did the separation cost, and when is a shared library the better answer?

lyftenvoysidecarcross-cutting-concernobservabilityservice-mesh
Show the full answer Hide the answer

The situation they were in

The symptom Lyft described in the announcement was behavioural, not technical: engineers were afraid to make service calls, because a call that got slow or failed produced no comparable signal. Every caller had its own idea of a timeout, its own retry policy, and its own way of naming a latency metric. Two services reporting "p99 400 ms" for the same dependency could mean two different measurements.

The usual fix is a client library per language that does retries, timeouts, circuit breaking and metric emission correctly. That works until the estate is polyglot. Then the library exists once per runtime, each copy is maintained by whoever needed it, and the three copies drift. Worse, the drift is in a cross-cutting concern, so when it matters — during an incident — the numbers that would settle the argument are not comparable.

What was actually separated

Not "networking". The separated concern is the uniform observation and control of a call: how long it may take, how many times it may be retried, when the caller stops trying, and what the resulting numbers are called. Envoy's documented shape is a self-contained process next to every application server; the application talks to localhost and is unaware of the network topology, and the proxies together form the mesh that carries load balancing, service discovery, retries and circuit breaking.

The load-bearing property is that it is out of process, which is what makes it language-agnostic. A library is in-process, so its reach is exactly one runtime. Move the same logic to a process and its reach is every runtime on the host, including the ones nobody has written a library for yet.

What it costs

  • A second process per host to build, deploy, configure and page someone about. The mesh becomes infrastructure with its own release cycle.
  • Two extra localhost hops per request. Small per hop, and not free on a chain five services deep with a tight budget.
  • A new shared failure domain. The thing that now sees every call can also break every call: a bad configuration push reaches every host at once, which is a failure class the library version simply does not have.
  • A debugging indirection. "The call failed" now has one more place to look, and engineers must learn that place.

The decision rule

Count the runtimes that make outbound calls. One runtime and one team: put it in a library, because the library costs nothing to operate and you get compile-time enforcement. Three or more runtimes, or a number of teams large enough that library upgrades take quarters to land, and the out-of-process version starts winning — the deciding fact is not scale but how many divergent implementations the concern currently has.

When this is the wrong answer

A single-language estate with fewer than about twenty services should not run a mesh. The honest comparison is a mesh against a maintained library, not against the drift you have today, and a maintained library in one language is cheaper on every axis. The other wrong-answer case is adopting the mesh for observability alone: if nobody has agreed what to call the latency metric, a proxy will emit an unagreed name uniformly and the argument during the next incident will be exactly as long.