beginner 3 min answer Multiple choice

Your chat endpoint streams tokens to the browser. A request fails after 300 of an expected 500 tokens are already on the user's screen. The team wants to apply the same automatic retry policy the rest of their API uses. Why is this different?

streamingretriesidempotencylatencynvidia
Pick one
Show the full answer Hide the answer

The mechanism

A streamed response commits its status code and headers before the model has produced its second token. The server sent 200 OK at the moment the first chunk left, and 300 tokens have been painted on the screen. HTTP has no facility to withdraw bytes already delivered, so the framework's retry interceptor - which works by discarding a failed response and issuing a fresh request - has nothing to discard.

Error handling therefore has to move inside the stream. With server-sent events that means an explicit error event in the body, and a client that understands the difference between "the stream ended" and "the answer finished".

Why the model cannot resume

The appealing idea is to continue from token 300. Generation has no resume point. The only way to approximate one is to issue a new request containing the original prompt plus the 300 tokens as an assistant prefix, which means paying the full prefill again for a prompt that may be 3000 tokens, and accepting that the continuation may not match the style or the reasoning of what the user already read.

The cost asymmetry is worth knowing. Inference has two phases with different bottlenecks: NVIDIA's inference-optimisation guidance describes prefill, which processes the whole prompt at once, as compute-bound, and decode, which produces one token per forward pass, as memory-bandwidth-bound. A retry throws away decode work you have already paid for and repeats the expensive prefill, so a "cheap resume" is the most expensive option on the table.

Why the other options fail

"The retry is cheap because the model resumes from the last token" is the assumption most teams start with, and it is the one that produces duplicated or contradictory text in production. There is no server-side session holding the partial generation; a new request is a new request.

"HTTP forbids a second body" confuses the protocol with the response that is already in flight. Nothing forbids issuing another request - the problem is what the user has already seen, and what you now owe them.

"A stream that has started will not fail midway" is comfortably wrong: the upstream provider can drop mid-generation, the guardrail can trip on partial output, the pod can be evicted, and the connection can be lost. Mid-stream failure is the common case that streaming designs must handle.

The decision rule

Retry only before the first token reaches the client. Buffer until the first token arrives, and treat that moment as the commit point: before it, retries are ordinary and invisible; after it, the correct behaviour is to end the stream with an error event and offer the user a regenerate button, which is a decision they can make and you cannot.

When streaming is the wrong choice

For short outputs - under roughly 200 tokens - buffering costs a second or two of perceived latency and buys back retries, whole-response validation, output guardrails that need the complete text, and response caching. Streaming is bought with those capabilities, and for a classification result, a JSON payload or a short extraction it is a bad trade that teams make by default because chat interfaces do it.