Happy Eyeballs
also called Connection Racing, Dual-Stack Attempt Ordering
Starting a second connection attempt on the other address family a few hundred milliseconds after the first, so a broken IPv6 or IPv4 path costs the user a short delay instead of a connection timeout.
A client resolves a name and gets both an AAAA and an A record. Address selection prefers IPv6, the route exists, and the packets go nowhere: a tunnel is broken, a firewall drops the family, a carrier's NAT64 gateway is failing for that destination. Nothing returns an error. The client sits in the SYN retransmission ladder, which on Linux defaults runs roughly 127 seconds across six retries before the connect call gives up.
The user never waits that long, because the application's own timeout fires first, and the request fails on a network where the IPv4 path would have connected in 40 ms. On the server side there is no signal at all: a SYN that is never answered becomes no socket, no log line and no error rate.
Happy Eyeballs is the client-side answer. Start one family, and if it has not connected within a short delay, start the other in parallel and keep whichever completes first. RFC 6555 (2012) introduced it and RFC 8305 (2017) fixed the constants.
Why it matters
This is the failure class that makes enabling IPv6 risky. Publishing an AAAA record hands every client a path you cannot monitor from inside your estate, and if clients do not race, the first broken path between a user and your edge becomes an outage your dashboards describe as a quiet day.
The economics are lopsided. Racing costs one extra SYN on a small fraction of connections; not racing costs a subset of users their entire session.
Implementation patterns
- Resolution delay, 50 ms. If the A answer arrives first, wait up to 50 ms for the AAAA, so family preference is not decided by whichever DNS answer was quicker.
- Connection attempt delay, 250 ms recommended, with a 100 ms floor and a 2-second ceiling in RFC 8305. Below the floor you double SYN volume for nothing; above a second the user feels it.
- Interleave families, cap concurrent attempts, and cancel the losers as soon as one connection completes.
- Cache the winning family per destination for minutes rather than forever, so repeat connections skip the race and a repaired network is noticed.
- Know which of your clients actually race. Browsers and most modern runtimes do; a service dialling an explicit address, or a sidecar resolving on the application's behalf, does not.
Industry example
The grounding is a standard rather than a company: the 50 ms and 250 ms constants come from RFC 8305 (2017), and browsers shipped racing years before that. That gap is why the classic report exists, where a site loads in the browser and hangs from a command-line client on the same machine.
The problem is routine on mobile networks. A carrier running IPv6-only with NAT64 gives a handset no IPv4 path to race towards, so when translation fails for one destination, a client that cannot race has no second option.
Failure scenarios
- A custom client that resolves once and dials one address. No race, so a black-holed family produces a two-minute hang that every retry repeats.
- Racing without cancelling. The server sees two SYNs per connection, abandoned half-open entries fill connection-tracking tables, and per-family traffic metrics stop meaning what the capacity plan assumed.
- A family cache with no expiry. Clients quietly settle on IPv4 for months after IPv6 was fixed, so the migration never completes.
- Racing combined with 0-RTT early data, where both attempts deliver the early data and a non-idempotent request is duplicated.
- Resolution moved into a proxy or mesh, which removes the client's ability to race while everyone assumes it still happens.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Race both families | Worst case bounded near 250 ms; single-family breakage invisible to users | Up to double the SYNs; extra connection-tracking state; distorted per-family metrics |
| Serial with a 2-second connect timeout | One attempt; simple to reason about | Every user on a broken path pays 2 seconds on every new connection |
| One family only | Deterministic, no machinery | IPv6 cannot be rolled out to clients at all |
When not to use it
Inside a controlled single-family network, racing adds state and hides a misconfiguration you would rather see loudly, because a path that works only through fallback is a path nobody fixes. Skip it where a connection attempt is expensive in itself, such as a handshake with a hardware security module or a dependency that counts concurrent connections for licensing, since raced attempts double the count. And it is the wrong tool for a path that is slow rather than broken: the race picks a winner at setup and says nothing about steady-state throughput.
Decision rule: race anything initiated by a device over a network you do not control; do not race inside your own fabric.
Interview question
Q: Users on one mobile carrier report that your API hangs for about two minutes and then errors. Your dashboards show no errors, no elevated latency, and almost no traffic from that carrier. Take me from the symptom to the fix.
What a strong answer covers: that an unanswered SYN produces no server-side signal, so missing traffic is the evidence; that a dual-stack client preferring a black-holed AAAA explains both the duration and the silence, with the SYN ladder giving roughly 127 seconds; checking whether the client library races; the client fix of a 250 ms attempt delay plus a connect deadline shorter than the user's patience; and that client-reported success rate split by address family is the only instrument that catches this before the ticket.
Quick check
Quiz: Why does a black-holed IPv6 path produce no server-side error? The connection never completes, so there is no socket and no request to count; the evidence is traffic that is absent rather than traffic that failed.
Flashcard: What does a client lose by not racing address families, and what are the two RFC 8305 constants? — It loses the whole session on a broken path, up to roughly 127 seconds of SYN retries; the constants are a 50 ms resolution delay and a 250 ms connection attempt delay.