concept

Head-of-Line Blocking

One delayed item stalling everything queued behind it, even when the rest could have been processed - the failure mode that ordering guarantees create.

tcphttp2quiclatencystreamingordering

Head-of-line blocking occurs wherever a queue enforces order and one item is delayed. Everything behind it waits, regardless of whether it was ready. The property causing the problem is not slowness; it is the ordering guarantee, and it appears at several independent layers.

In TCP. The protocol delivers bytes in order, so a single lost packet stalls delivery of every byte behind it until the retransmission arrives — even bytes that already reached the receiver.

In HTTP/1.1. One request per connection at a time, so a slow response blocks the next request. This is why browsers opened multiple connections per origin.

In HTTP/2. Application-level blocking was solved by multiplexing streams over one connection — but all those streams still ride one TCP connection, so a lost packet stalls every stream, not just the one it belonged to. HTTP/2 moved the problem down a layer rather than removing it.

In message queues and stream partitions. A partition that guarantees ordering means a poison message or a slow consumer blocks everything behind it in that partition.

Why it matters

It is the reason ordering guarantees are expensive, and the reason "just make it ordered" is rarely a free choice. Every ordering guarantee creates a queue, and every queue can be blocked by its head.

Recognising it converts a mysterious latency symptom — "everything got slow at once, but only on this connection" — into a structural explanation.

Implementation patterns

  • Independent streams with independent delivery, which is what QUIC provides: streams are implemented in userspace over UDP, so loss affecting one does not block the others.
  • Multiple connections as the crude version, which is why HTTP/1.1 clients opened six per origin.
  • Partition by key in messaging, so ordering is guaranteed only within a key and one slow key does not block unrelated ones.
  • Dead-letter handling with a retry ceiling, so a poison message is removed from the head rather than blocking a partition indefinitely.
  • Selective reliability for real-time media: retransmit what still matters, discard what has expired.

Industry example

Video delivery is where the layered version bites hardest. A player fetches media segments, manifests, thumbnails and telemetry concurrently. Under HTTP/2, all of these share one TCP connection, so a packet lost while fetching a thumbnail delays the next media segment — a stall caused by something the user cannot see.

Moving to HTTP/3 over QUIC removes exactly this: streams are independent at the transport layer, so loss is contained to the stream it affected. Combined with QUIC's single-round-trip handshake and connection migration across network changes, this is why video platforms were among the earliest large-scale adopters.

Live ingest shows the opposite lesson. There, in-order delivery is not merely inconvenient but actively harmful: TCP will faithfully retransmit a frame from four seconds ago that nothing can use, consuming uplink the current frame needs. For real-time media, reliability and timeliness conflict, and timeliness wins — which is why the industry moved to UDP-based transports with application-level selective reliability.

Failure scenarios

  • Assuming HTTP/2 removed it, and being surprised by correlated stalls across unrelated streams.
  • A single ordered partition for a whole workload, where one slow or poisonous message halts everything.
  • Ordering requested by default in a message contract, when the business only needs per-customer ordering.
  • Retrying a stalled connection rather than opening a new one, prolonging the block.

Trade-offs

Removing head-of-line blocking means giving up an ordering guarantee, which some workloads genuinely need. The productive question is what is the smallest scope over which ordering is actually required — usually per user, per entity or per key, rather than globally. Narrowing the scope preserves the guarantee where it matters and removes the blocking everywhere else.

Interview question

"A client fetches ten resources concurrently over HTTP/2 and one packet is lost. What happens to the other nine? Now tell me what changes under HTTP/3, and what you give up."