pattern

Streaming Passthrough Gateway

also called Token Stream Relay, SSE Passthrough Proxy

Relaying a model's token stream through a shared proxy without buffering it, which makes the gateway's capacity unit concurrent long-lived connections rather than requests per second.

ai-gatewaysstreamingsseconcurrencydraining

A gateway in front of model providers starts as an ordinary reverse proxy. Then the product streams tokens and its job changes shape: it is no longer handling short requests, it is holding thousands of connections open for tens of seconds each.

The capacity arithmetic follows. At 50 requests per second with a mean response of 20 seconds, Little's law puts 1,000 streams open at any instant — a gateway sized at 50 of something must hold twenty times that. Thread-per-connection runtimes exhaust their workers in the low hundreds, which is why this component needs asynchronous I/O.

Why it matters

Three unspecified properties become load-bearing at once. Buffering destroys the product, because streaming exists so the first token arrives in a few hundred milliseconds and a proxy holding bytes until a buffer fills gives a pause then a wall of text. Deploys become user-visible, since a rolling restart cuts every in-flight stream mid-sentence. And cost accounting moves to the end of the stream, so an abandoned stream is spend with no record attached.

Implementation patterns

  • Disable response buffering explicitly on every hop. This is a documented default rather than a surprise: nginx has proxy_buffering on by default and holds the response in 4 KB or 8 KB buffers, and short answers never fill one. The per-response escape hatch is the X-Accel-Buffering: no header, which nginx honours unless it has been told to ignore it.
  • Set idle timeouts above the p99 stream duration on every hop, and send a heartbeat event so an intermediary does not treat a thinking model as a dead connection.
  • Drain for longer than a stream lives. A 30-second grace period in front of 90-second p99 streams truncates on every deploy.
  • Propagate cancellation upstream on client disconnect, or generation continues and you pay for tokens nobody reads.
  • Account on both paths: usage from the final event on success, tokens relayed on an abort, so abandonment is visible.
  • Cap concurrent streams per tenant, not just requests per minute: long-context traffic can exhaust the connection budget without exceeding any request-rate limit.
  • Run output guardrails on buffered chunks, a sentence at a time. A check needing the whole response cannot coexist with streaming, and learning that after launch forces a product change.

Industry example

The mechanism is easiest to see in reverse-proxy defaults documented since well before 2020: a proxy configured for ordinary web responses buffers them, which is correct for a 20 KB HTML page and wrong for a token stream, and the reported symptom is always "the model got slow" rather than "our proxy is batching". A platform running six products through one shared gateway then meets the second half: the gateway is a stateful component on every product's critical path, and its deploy schedule is a product-quality event.

Failure scenarios

  • Silent cost leak. Users navigate away, generation continues, and the bill includes undelivered output.
  • Truncation on every deploy, reported as the assistant "cutting off" and never reproducible in testing.
  • Worker exhaustion. A blocking runtime with 200 workers and 1,000 streams queues new requests behind open ones, so time-to-first-token collapses while CPU sits low.
  • A guardrail that cannot run, because the safety check needs the full response. That choice gets made under deadline pressure unless it was made in design.
  • Mid-stream provider failure with 300 tokens on screen, where a generic retry appends a second answer to the first.

Trade-offs

Passthrough buys the latency the product was built for and pays with a stateful, connection-bound component: more memory and file descriptors per instance, slower deploys, and harder autoscaling, since new instances cannot adopt existing streams. Buffering buys stateless scaling and whole-response checks, and loses the feature.

When not to use it

If the workload does not stream, do not build for streaming. Classification, extraction, embedding and batch summarisation are request-response, and a plain proxy with a short timeout is a far smaller thing to operate.

Prefer a library or sidecar over a network hop when the only requirement is key custody and token accounting. The gateway earns its place when something must be enforced that an application cannot be trusted to enforce — a quota, a routing policy, an audit record. When it is merely a convenient home for shared code, a client library gives the same benefit without a stateful component on every request path.

Interview question

Q: Our AI gateway holds streaming responses for six products. Size it for 50 requests per second with 20-second responses, then tell me what happens during a rolling deploy and what you would change.

What a strong answer covers: the concurrency calculation and that the capacity unit is connection-seconds; asynchronous I/O; buffering defaults on every hop and the latency symptom they cause; draining set against p99 stream duration; cancellation propagation and the cost of abandoned streams; and where output guardrails can run once the response is a stream.

Quick check

Quiz: 80 requests per second, 15-second mean stream. How many connections must the gateway hold? — About 1,200 concurrently, so capacity is planned in connections and a request-rate target alone under-provisions it.

Flashcard: Why does a new AI gateway often make time-to-first-token worse? — Reverse proxies buffer by default (nginx proxy_buffering, 4 to 8 KB buffers), so token events wait for a full buffer; disable buffering or send X-Accel-Buffering: no.