advanced 3 min answer

An internal fleet is moving service-to-service calls from HTTP/1.1 behind an L4 load balancer to gRPC over HTTP/2. Sequence the migration and say which load-balancing property the team has just deleted.

http2grpcload-balancingconnection-poolingmigration
Show the full answer Hide the answer

The property you just deleted

An L4 load balancer balances connections, and HTTP/2 makes one long-lived connection carry every request. The balancing decision is therefore made once, at connect time, and then holds for as long as that connection stays open. Under HTTP/1.1 with a pool of short-lived connections, the same load balancer reshuffled continuously and nobody had to think about it.

The symptom is not slowness; it is imbalance. After a scale-out event, new backends receive no traffic at all, because every caller already holds a connection to an old one.

The sequence

  1. Baseline per-backend request distribution before changing anything: requests per backend per minute, and the ratio of the busiest to the median backend. You cannot detect imbalance you never measured, and this is the one number the migration puts at risk.
  2. Choose the balancing layer before the protocol. Either an L7 proxy that balances per stream, or client-side balancing with a resolver that sees every endpoint and a round-robin or least-request policy. Both are reversible decisions; doing nothing is not.
  3. Migrate one caller at a time behind a flag, keeping the HTTP/1.1 listener live on a second port so rollback is a configuration change rather than a deploy.
  4. Set a maximum connection age with a grace period on the server. gRPC exposes this as MAX_CONNECTION_AGE: the server sends GOAWAY when a connection reaches the age, the client re-resolves and reconnects, and the fleet rebalances on a cycle you choose. Without it, a connection established at 09:00 can still be pinned to the same backend at 17:00.
  5. Jitter that age. A fixed 30-minute lifetime set identically everywhere produces a synchronised reconnect storm every 30 minutes, which is a self-inflicted thundering herd.
  6. Re-check the distribution against the baseline, then remove the HTTP/1.1 path.

Where it diverges, and how you would know

Per-backend CPU spread is the fastest signal: two backends at 90% while eighteen sit at 15% is this bug and nothing else. The confirming view is a histogram of active connections per backend, which is the quantity the L4 balancer is actually equalising. A new pod sitting at zero requests after a scale-out is the same bug seen from the other end.

Long-lived streams are the residual risk. Maximum connection age rebalances new streams only; streams already in flight move when the grace period closes them, so a service with hour-long streams rebalances on the order of hours.

The point of no return

Removing the HTTP/1.1 listener. Everything before it is a flag flip.

How long it really takes

The protocol switch is a week of work. The balancing-layer decision is the project, because it determines whether you now operate a mesh, a client-side library in every language you use, or a fleet of L7 proxies with their own capacity and failure modes.

When not to do this at all

With three backends and heavy traffic per caller, imbalance is cheap and HTTP/1.1 keep-alive is fine. The migration earns itself when per-request connection and header overhead is a measurable share of latency, when you need bidirectional streaming, or when a strict schema and generated clients are worth more than the balancing complexity. Adopting HTTP/2 between two services that exchange 5 requests per second is cost with no return.