An edge network announces one anycast prefix from 40 locations. At one of them, the upstream transit begins dropping 8% of packets. BGP sessions stay established, the location's local health endpoint returns 200, and CPU is normal. What do users mapped there experience, and what stops it?
Show the full answer Hide the answer
Second by second
Nothing changes in routing, because BGP has no opinion about packet loss. It tracks whether a session is alive, using a 30-second keepalive and a 90-second hold timer in common configurations. A session dropping 8% of packets is alive.
For the users whose routes point at that location, every flow degrades at once. TCP treats loss as congestion and halves its window. A lost segment at the tail of a response has no following packet to trigger a duplicate acknowledgement, so recovery waits for a retransmission timeout, which Linux clamps at a 200 ms floor. A TLS handshake needs several packets to complete in both directions, so at 8% per-packet loss a meaningful share of handshakes stall or fail outright, and the user sees a page that hangs rather than an error.
The blast radius is the catchment of one location, not one fortieth of users. Catchments are wildly uneven: a large metro location can hold 10% of traffic while a small one holds 0.2%.
Why anycast does not help here
Anycast's failover story assumes a location either answers or stops announcing. Partial failure is outside the model. Worse, the local health endpoint returning 200 is honest: the server is healthy, the path to it is not. A health check that runs inside the failure domain cannot see the failure domain.
What actually stops it
- Health-triggered route withdrawal driven by outside observation. Probes from other networks, plus client-reported success rates, decide whether a location keeps announcing. The route is the actuator.
- Hysteresis and a quorum. Withdraw after several consecutive bad intervals, return after a longer clean period, and cap how many locations may withdraw at once. Without the cap, a correlated measurement fault withdraws half the fleet and the survivors fail under the load they inherit.
- Expect connection resets. Withdrawal moves the next packet of an established flow to a location holding no state for it, so in-flight connections break. Convergence is seconds to tens of seconds for neighbours and can stretch with a 30-second minimum route advertisement interval. Draining is not graceful; it is fast.
What would have to be true for it to self-heal
The transit provider's own loss detection would have to withdraw the path, which is exactly the assumption that failed. Self-healing requires the loss signal and the routing decision to live in the same system, which is why edge operators build their own traffic engineering on top of measurement rather than trusting BGP to notice.
Common weak answers
- "Anycast fails over automatically." Only for hard failure, and only when something withdraws the route.
- "Add more health checks." Not unless they run outside the location. More internal checks produce more confident wrong answers.
- "Lower the BGP timers." Faster session death detection does nothing for a session that is not dying, and aggressive timers add their own flapping risk.