ECMP Flow Pinning
also called Flow Hashing, Per-Flow Path Selection
Because a router picks one of several equal-cost paths by hashing a flow's five-tuple, one connection is confined to one link, so aggregate capacity is never per-connection capacity.
A nightly replication job moves 30 TB between two racks joined by four 25 Gbps links. It tops out around 23 Gbps and takes four times as long as planned. Nobody is congested: the bundle reports 25% utilisation, the switches show no discards, and both hosts have spare CPU.
The transfer is one TCP connection. A router or link-aggregation group spreads traffic by hashing fields of each packet, usually source and destination address, port and protocol, and sends every packet with the same hash down the same member link. That is deliberate, because packets of one flow must not be reordered: TCP reads reordering as loss and collapses its window. The consequence is arithmetic rather than opinion. A single flow's ceiling is one member link, whatever the bundle adds up to.
The same mechanism produces imbalance with a handful of large flows even when the hash is perfect. Eight elephant flows across four links do not land two per link; they land where the hash puts them, and one link can carry three while another carries one.
Why it matters
Capacity planning from aggregate utilisation misleads in both directions. A bundle at 25% can be saturating one member and dropping packets, and the service that suffers is the one with a single large connection: backup, replication, bulk ingest, or a tunnel carrying everything else.
It also decides where the fix lives. The remedy for a per-flow ceiling is almost always more connections in the application, which is cheap, rather than more bandwidth in the network, which is expensive and changes nothing.
Implementation patterns
- Parallel streams. N connections give N hash results and can use up to N links; object-storage transfer tools default to 8 to 16 parts in flight for exactly this reason. Use at least twice as many streams as paths so a collision leaves no link idle.
- Entropy inside tunnels. VXLAN varies the outer UDP source port per inner flow so the fabric can still spread. A tunnel with a fixed outer tuple pins an entire site's traffic onto one member link.
- Seed hashes per device, because successive tiers hashing the same fields with the same seed polarise traffic and later tiers see no spread at all.
- Instrument per-member-link counters, not the bundle: discards on one member with the aggregate at a third of capacity is the fingerprint.
- Flowlet or adaptive switching where the fabric supports it, rebalancing at gaps in a flow rather than per packet.
Industry example
The published grounding is data-centre networking research rather than company blogs. Hedera (2010) measured how collisions between long-lived flows waste a Clos fabric's bisection bandwidth, and CONGA (2014) answered it with flowlet-level rebalancing. Both start from the same observation: per-flow hashing is right for correctness and wrong for utilisation once flow sizes are heavily skewed.
A live-streaming platform in Twitch's mould shows the shape in production. Contribution feeds for one large event arrive as a few very large flows, so an ingest fabric can be a quarter idle while one member link drops packets and one broadcaster degrades.
Failure scenarios
- The single-flow ceiling, where a transfer plateaus at one link's rate and more bandwidth is bought with no effect.
- A tunnel with no entropy pinning a whole site onto one 10 Gbps member of a 40 Gbps bundle.
- Hash polarisation across tiers, which concentrates traffic although every device is correct in isolation.
- Rehashing on change. Adding or removing a member re-maps existing flows, so established connections see reordering or resets during a maintenance window.
- Mixed member speeds. A bundle of 10 and 25 Gbps links gives each flow the rate of whichever member it hashed onto, so throughput becomes a two-outcome lottery.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Per-flow hashing (the default) | No reordering; cheap and stateless in hardware | Per-flow ceiling of one link; imbalance with skewed flow sizes |
| Per-packet spraying | Near-perfect link utilisation | Reordering, which TCP treats as loss unless endpoints are built for it |
| Flowlet or adaptive switching | Most of the balance without reordering inside a burst | Needs fabric support and congestion feedback; more state to operate |
| Parallel application streams | Fixes the ceiling without touching the network | More connections and memory, and a per-stream failure to handle |
When not to use it
Do not engineer around flow pinning when no single flow comes close to a member link's rate: a service whose largest connection peaks at 200 Mbps on a 25 Gbps path gains nothing from parallel streams and pays in connection count. Reach for multipath transports only after the application-level split has been tried, because splitting a transfer across N connections is usually a configuration flag while MPTCP or multipath QUIC is a platform project.
Decision rule: if one connection must carry more than one member link's bandwidth, split it at the application layer; if the aggregate looks healthy while one flow is slow, go straight to per-member-link counters.
Interview question
Q: A cross-rack database restore runs at 9.4 Gbps on a path advertised as 40 Gbps. The network team reports no congestion and shows a bundle at 24% utilisation. What is happening, what do you measure to prove it, and what do you change?
What a strong answer covers: that per-flow hashing confines one connection to one member, so 9.4 Gbps is a 10 Gbps link rather than a shortage; that both parties are right because aggregate and per-member views answer different questions; that the proof is per-member throughput and discard counters plus the test that two parallel streams roughly double throughput; and that the fix is parallel streams in the restore tool, with tunnel entropy checked if the path is encapsulated.
Quick check
Quiz: Why can one TCP connection not use a four-link bundle's full capacity? Every packet of the flow hashes to the same member link, which is required to avoid reordering, so the flow's ceiling is one member's rate.
Flashcard: A transfer plateaus at a quarter of a bundle's capacity and nothing is congested. What is the cause, and the cheapest fix? — One flow pinned to one member by five-tuple hashing; run parallel streams, at least twice the number of paths.