Health Check & Service Discovery  ·  View 04 of 21  ·  People and journeys

Journey — Deploy Without Dropping Traffic

The trough is not the drain. It is the moment the new replica gets a full share of traffic before it is warm.

Editable source SVG draw.io All views
Service owner ships twice a day Goal — Replace every replica of my service in ten minutes with nobody noticing Trigger — A merged pull request; the pipeline asks to drain the first instance Done when — Old version gone, new version at full weight, zero caller-visible errors 1 · Ask to drain pipeline 2 · Drain ◆ moment of truth 3 · Terminate 4 · Admit new version ◆ moment of truth 5 · Ramp to full What happens Budget checked Marked draining Withdrawal pushed Grace period ends Readiness passes Weight 1 of 10 Weight rises 60 s Who acts Deploy pipeline Evaluation tier Orchestrator Evaluation tier Every client How it feels Confident Watching Burned Where it hurts Budget refuses, no reason Terminated before callers knew Cold replica at full share What answers it Typed refusal + floor Drain gated on measured delay Lease expiry backstop Control-plane slow start Ramp is uniform Journey — Deploy Without Dropping Traffic v 1.0 · owner Reliability Architecture

What answers the trough

  • Drain-to-termination is gated on the measured propagation delay, not a fixed sleep — the number moves when the platform does.
  • Slow start is published as a rising weight by the control plane, so every client honours the same curve.
  • A refused drain returns the floor it hit, so the pipeline can report a reason rather than retrying blindly.

Assumptions

  • Graceful drain visible to 99% of callers in 2 s, 99.9% in 5 s; new instance eligible 20 s after first readiness pass.
  • Ramp to full weight over at least 60 s; lease expiry removes an ungraceful exit within 30 s.

Residual risk

  • A service whose warm-up is longer than 60 s will still be hurt by the ramp. The ramp length is per service, and few teams tune it.