LLM Application Architecture advanced 8 min read 7 flashcards

Streaming Transport, Cancellation and Resumability

Why an HTTP 200 tells you nothing about a streamed response, what the infrastructure between your server and the user does to a long-lived stream, and how to make a stream survive the connection that carried it.

A user asks a question. The first token lands in 400 ms, the answer takes 40 seconds, and at second 22 the train goes into a tunnel. Three things are now true: the user has half an answer, your database has nothing, and the provider is still generating tokens you will be billed for. Streaming is introduced as a latency trick, and it is one, but what it really does is move the response out of the request/response model into a transport problem with its own failure modes.

A 200 is not an outcome

Server-sent events deliver a sequence of named events over one long-lived HTTP response. Anthropic's Messages API streams a fixed flow: message_start, then per content block a content_block_start, a run of content_block_delta events and a content_block_stop, then one or more message_delta events and a final message_stop (Anthropic, Streaming messages). The status line is written before any of that happens, which has a consequence most client code gets wrong: "when receiving a streaming response over server-sent events (SSE), an error can occur after the API returns a 200 response. In that case, error handling doesn't follow these standard mechanisms" (Anthropic, Claude API errors). A capacity error that would have been an HTTP 529 arrives as event: error with an overloaded_error payload, mid-answer.

So the terminal event, not the status code, is the result. Absence of message_stop is a failure even when nothing errored, and a client that treats "the iterator ended" as success persists truncated answers forever without an alarm firing.

The other reason to stream is that long non-streaming requests are structurally fragile. Anthropic advises streaming or the Batches API for anything over about 10 minutes, because "some networks may drop idle connections after a variable period of time", and the SDKs enforce it by validating that a non-streaming request is not expected to exceed a 10-minute timeout and by setting TCP keep-alive. The ping events sprinkled through a stream exist for the same reason: traffic keeps intermediaries from deciding a connection is idle.

The infrastructure in between

Streaming works on a laptop and dies in staging, and the reason is almost always a reverse proxy. nginx buffers upstream responses by default (proxy_buffering on), absorbing the whole response before forwarding it, so the client receives nothing until the upstream closes and then receives everything at once. The fix is proxy_buffering off on the streaming location, an X-Accel-Buffering: no response header for the layers that honour it, and a proxy_read_timeout raised well above the 60-second default that otherwise kills any answer longer than a minute.

Browsers add their own ceiling. Over HTTP/1.1 a browser holds roughly six connections per origin and each open stream consumes one; HTTP/2 multiplexes them over a single connection and the limit stops mattering. One chat stream per tab is fine either way, but a dashboard with ten live feeds on HTTP/1.1 starves itself, and the starvation looks like a backend problem.

Cancellation has to be propagated

A user who closes the tab sends nothing. The browser drops a socket; your server may not notice for a while, and even when it does, the provider is unaffected until your server closes its own socket upstream. Cancellation is an explicit chain: client abort signal, request context, provider call, and every tool call already in flight. Miss a link and you pay for output no one will read.

Input tokens are billed whatever happens, since they were processed before the first delta, and output accrues until generation actually stops. The accounting hook is that the usage field on message_delta events is cumulative, so a stream that ends early still tells you exactly what it consumed, provided you persisted the deltas rather than only the accumulated text.

Making the stream outlive its connection

SSE has a resumption mechanism: the server tags events with id:, a reconnecting client sends the last one it saw in a Last-Event-ID header, and the retry: field sets the reconnection delay. This only helps if something still holds the events, which the provider connection does not.

The shape that works is to decouple the model call from the user's connection. Give each generation a stream id, write every delta to a durable append-only log keyed by (stream_id, seq), and have the client reconnect with the last sequence it received. The generation then runs to completion whether or not anyone is attached, which is the right behaviour when the tokens are billed anyway, and the tunnel becomes a replay rather than a loss.

One cost to budget for: Anthropic's prompt cache lifetime runs from the start of the request that writes or reads the entry, not from the end of its response (Anthropic, Prompt caching). A four-minute stream leaves roughly one minute of the default five-minute window for the next turn to begin, so long answers quietly consume the cache window they depend on.

When it breaks

Partial JSON in tool calls. With fine-grained tool streaming the server emits input_json_delta fragments without validating them, so the accumulated string is not guaranteed to parse, and a response can stop at max_tokens midway through a parameter (Anthropic, Fine-grained tool streaming). Guard the parse and check the stop reason before acting.

Moderation re-buffers what you worked to stream. Scanning output before showing it reinstates the latency streaming removed. Either scan a sliding window and release text behind the frontier, or show text and retract it, which users notice.

The client becomes the source of truth. Persisting the assistant turn from text the browser accumulated lets a flaky network silently edit both your history and your billing record. Persist server-side from the deltas.

A reconnect that re-issues the call is not a resume. It generates and bills a second answer, and the two will differ. See latency, streaming and perceived speed for what the user makes of all this.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Anthropic, Streaming messages platform.claude.com
  2. Anthropic, Claude API errors platform.claude.com
  3. Anthropic, Prompt caching platform.claude.com
  4. Anthropic, Fine-grained tool streaming platform.claude.com
Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track