A 180 KB JSON response takes about 400 ms to reach a client 90 ms away on a connection whose handshake has just finished. The handler took 12 ms and the link has 200 Mbps spare. Measuring from the moment the request is written, what sets the floor?
Show the full answer Hide the answer
The arithmetic
A new TCP connection does not start at line rate. It starts at an initial congestion window of 10 segments, which RFC 6928 fixed in 2013 and Linux has shipped since 2.6.39 in 2011. At a 1460-byte segment that is about 14.6 KB in flight before the sender must wait for an acknowledgement.
The window roughly doubles per round trip, so the cumulative delivery looks like this at 90 ms RTT:
| Round trip | Window | Cumulative bytes |
|---|---|---|
| 1 (request out, first flight back) | 14.6 KB | 15 KB |
| 2 | 29 KB | 44 KB |
| 3 | 58 KB | 102 KB |
| 4 | 117 KB | 219 KB |
180 KB lands inside the fourth window. One round trip to carry the request plus four to ramp is five times 90 ms, roughly 450 ms, against 12 ms of server work. Bandwidth never becomes the constraint because the sender is never allowed to use it.
Why the other options fail
- The 64 KB receive window. A real limit in 1995. Window scaling has been on by default for two decades and receive buffers autotune into the megabytes. If this were the cause, throughput would plateau rather than ramp.
- The TLS handshake. It genuinely costs round trips, and that is why the stem excludes it: the clock starts after the handshake. Reaching for the handshake here means missing that data transfer has its own warm-up.
- Nagle's algorithm. It delays a small trailing segment by at most one round trip and only when a previous segment is unacknowledged. It cannot explain 400 ms, and disabling it is the most commonly applied fix for a problem nobody measured.
What to change, in order
- Reuse the connection. The window is per-connection state that grows as the connection is used, so the second response on a warm connection arrives in one or two round trips. Connection pooling and HTTP keep-alive beat every payload optimisation here.
- Get the first useful response under roughly 14 KB so it fits the first window. This is the real reason small first-paint payloads matter, and it is a hard threshold rather than a preference.
- Move the origin closer, because every one of these costs is a multiple of RTT.
When this is the wrong answer
On a warm connection, or inside one datacentre where RTT is under a millisecond, slow start is irrelevant and the time is in the application. The diagnostic is cheap: compare the transfer on a fresh connection against the second request on the same connection. If they differ by several RTTs, this is the cause; if they match, stop looking at the transport.