Somebody else resolved it
A patch release to a distributed inference framework now refuses to fetch a caller's image when an HTTP proxy stands in the way, because the proxy resolves the destination where the server cannot check it. Turning the fetch back on means vouching, in an environment variable, for a control you do not own.
The argumentMultimodal serving made the inference tier an HTTP client dialling caller-supplied addresses from inside the cluster, and because every protection depends on the server resolving the name itself, an ordinary egress proxy reduces the control to an operator's environment variable vouching for somebody else's check.
On 7 October, NVIDIA's Dynamo — the distributed inference framework that fronts vLLM, SGLang and TensorRT-LLM workers behind a router — shipped v1.5.1. It is a patch release, and it reads like one: router overload recovery, min_tokens on tokenizer-free SGLang decode workers, a stop-sequence leak. It also carries two breaking changes. The first says this:
Workers that fetch request media through an ambient HTTP proxy must now set
DYN_MM_TRUST_EGRESS_PROXY=1. Without it, a policy-checked media fetch that would route through a proxy is refused, because the proxy resolves the destination outside Dynamo's address check.
Read that as the person who owns the deployment. A path that worked yesterday stops working, in a patch, because something the serving stack cannot see is doing the name resolution. The project is not apologising for it. It is the correct behaviour. It is also the most interesting sentence published about inference architecture this fortnight, and it is not about inference.
Nothing in a chat completion looks like a network operation. A multimodal request carries {"type": "image_url", "image_url": {"url": ...}} — a field, among fields, no different in shape from a temperature or a token limit. But the bytes have to arrive from somewhere, and in both of the major open serving stacks they arrive because the server goes and gets them. That one decision, taken for the obvious reason that it is convenient for the client, converts the inference tier into an HTTP client that opens connections to addresses chosen by whoever sent the prompt, from wherever the pods happen to sit. vLLM's own documentation says the quiet part out loud: the restriction matters most "if you run vLLM in a containerized environment where the vLLM pods may have unrestricted access to internal networks."
How that gets defended is worth following, because it is where the argument lives. On 19 September, Dynamo merged a change first titled "close DNS-rebinding gap in media fetch" and renamed, by the time it landed, to "pin validated DNS answers at connect time." The defect it closes is stated plainly in the pull request: validate_url resolves a hostname and checks the addresses, then the client resolves the name again to open the connection. "A server that answers differently the second time is checked on one address and dialed on another." The fix does not live in the validator. It is a filtering resolver wired into aiohttp's TCPConnector, so the client resolves once, keeps every address that passed, and dials only those — holding the hostname for TLS so certificate verification still works. The Rust frontend does the same through reqwest's dns_resolver.
That is careful work, and it establishes the thesis almost by accident. The protection is not a rule about URLs. It is a property of the server performing its own name resolution. Everything rests on the checker and the dialler being the same component, close enough together that no answer can change between them.
Now put a forward proxy in the path. The server resolves nothing; it hands a hostname to the proxy, and the proxy decides what that means. There is no address to validate and nothing to pin. So the fetch fails closed — and the only way to reopen it is an environment variable in which an operator asserts that the proxy performs an equivalent check. The trust boundary of the inference tier has become a boolean about another team's infrastructure.
Look at what else had to be built in the same cycle, and the shape becomes clearer. Inline data: URLs were accepted at any size; they are now capped at 16 MiB by DYN_MM_MAX_DATA_URL_MB, a limit that has to be set on both the Rust frontend and the Python workers, with a test whose job is to keep the two defaults from drifting apart. SGLang's image-diffusion and video handlers passed a client-supplied input_reference through to the generator's image_path "after only a non-empty check" — any path was accepted — and now require DYN_MM_LOCAL_PATH to name the one permitted directory. Image dimensions outside 1 to 4096 are rejected. Malformed base64 audio no longer takes a worker down. Remote references download through a policy-aware client that rechecks every redirect, into a temporary file capped at 64 MiB.
None of that is exotic. It is the standard inventory of any program that fetches untrusted URLs: address validation, redirect policy, size caps, path allowlists, one pooled client instead of several. What is notable is where it is being written — into a GPU inference server, in October 2026, as a patch — and what that says about how the tier was understood until now. We have been reasoning about model serving as compute. Tokens in, tokens out, a function of accelerators and a batch scheduler, with cost per million as the number that matters. Multimodality gave it an outbound interface, and nobody drew it on the diagram.
The boolean is also harder to hold than it looks. On the same day v1.5.1 published, a follow-up opened — still unmerged when I read it — because the first version of the resolver "exempted a configured proxy by host name on every connection, also on connections that did not go through the proxy." Since aiohttp caches DNS answers per connector, proxied and direct fetches need separate connectors; the client may now keep three connection pools, one per trust outcome, and the --http-max-connections help text had to be rewritten to say so. A review comment on that pull request observes that redirects taken through the trusted-proxy session could reach the proxy's own private endpoint directly. Meanwhile DYN_MM_ALLOW_INTERNAL=1 lets private addresses through and switches off the proxy refusal altogether — and the same pull request teaches the HTTP benchmark sweep to set it when it is unset, which is precisely how a flag becomes a default.
This is not one vendor's bad fortnight. vLLM gated per-request multimodal arguments on 26 September: mm_processor_kwargs and media_io_kwargs are "rejected by default," because they "can change image, video, or audio loading, sizing, sampling, and preprocessing behavior, causing excessive CPU, GPU, or memory use when controlled by an untrusted client." They come back only if the server is started with --trust-request-mm-kwargs, and the documentation is blunt about the condition: "Do not enable this option on an endpoint exposed to untrusted clients." On 6 October the same page was amended to state that vLLM "does not provide isolation between tenants that share the same server process." Its multimodal guide recommends --allowed-media-domains against server-side request forgery and VLLM_MEDIA_URL_ALLOW_REDIRECTS=0 so that redirects cannot be followed "to bypass domain restrictions." Two independent serving stacks, inside one fortnight, reaching the same conclusion: what a caller may ask the serving tier to do on its own network has to be narrowed, and the narrowing belongs to the operator.
The case for finding none of this troubling is strong, and it deserves to be put at full strength. Reimplementing network policy inside an application is a known mistake. Organisations spent two decades consolidating egress control into proxies exactly so that individual services would stop carrying their own allowlists and drifting out of step with each other. An inference server that tried to second-guess the proxy would be worse than one that declines and says why. Failing closed is the conservative choice. Shipping it as a breaking change in a patch release, rather than quietly logging a warning, is the strongest signal a project can send. And on that reading DYN_MM_TRUST_EGRESS_PROXY is not an abdication at all — it is an accurate declaration of where the control actually lives, which is more than most software manages.
All of that is right, and it is why the project's decision is not the thing to argue with. What I do not accept is that the declaration has anywhere to live. These are the projects' own records — merged pull requests, release notes, security pages — and they are unusually candid. Nothing downstream of them is. The architecture document says "inference service," with an arrow going in. The threat model, where one exists, is about the model: jailbreaks, injection, data leaving in tokens. Neither says that a chat completion is also an outbound GET, that the policy governing it is split between a Helm value and a proxy owned by a different team, or that whether those two agree is asserted once by whoever edited the values file and tested by nothing afterwards. There is no conformance test for "the proxy performs the equivalent check." There is no telemetry attribute saying this completion caused a fetch, to this address, which passed. The variable is a promise with no counterparty.
That gap widens as the tier acquires more clients of its own. Media fetching is the easy case: one hop, triggered by a documented field, with a size you can cap. Tool calls and retrieval make the same pods a client of databases, internal APIs and other models, with the destination chosen further and further away from anyone who reviewed the design. Each of those capabilities will arrive the way this one did — as a convenience in a request schema, implemented by the server because that is where it is easiest, and recognised as a trust boundary somewhere around the second security review. The architectural question stops being how many accelerators the tier needs and becomes what the tier can reach, and who said it could. In every other part of a serious estate, that second question has an owner, a change record and a test. In the serving tier it currently has an environment variable.
There is something almost generous about a fetch that fails closed, because it tells you what your serving tier was doing all along. A deployment that broke on 7 October broke because a proxy stood between its pods and the open internet, which means somebody had already decided that this traffic should be governed by something. The deployments that upgraded without noticing anything are the ones where the resolver answered, the socket opened, and nothing was ever in the way.
What this is argued from
Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.
- Dynamo v1.5.1 release notes
- fix(common-http) pin validated DNS answers at connect time (#14474)
- fix(common-http) exempt the egress proxy only on proxy connections (#15780)
- fix(sglang) validate diffusion input_reference and bound media fetches (#14435)
- fix(multimodal) cap the size of inline data URLs (#14437)
- Multimodal Inputs — vLLM documentation
- Security — vLLM documentation
Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.