Opaque, supplied by the caller
vLLM can now rewrite the weights of a running engine while requests are in flight, and its documentation says plainly that a paused request resumes under the new ones. The only version those weights carry on the serving path is a free string with a default value of "default".
The argumentvLLM has made a live model's weights mutable mid-request while leaving their only serving-path identity an optional string the trainer types, so the one instrument that could say what actually answered is a dev-mode endpoint you must stop serving to call.
There are two lines in vLLM's engine core that, between them, describe the state of model identity in production inference serving. They sit in the constructor, a few lines below where the config is logged:
# Opaque weight version supplied by the caller.
self._weight_version = "default"
That is the whole of it. Inside the process that is answering requests, the version of the model being served is a string whose default value is the word default, and the comment above it says exactly what the code thinks of it. Opaque. Supplied by the caller.
Read the two write paths and the picture sharpens. finish_weight_update accepts an optional version — in the control protocol, optional string weight_version = 1 — and if the caller does not supply one, the transfer completes and the previous value stands. There is also update_weight_version, whose docstring is a single sentence: "Set the weight version without updating weights." So the weights can change while the version does not, and the version can change while the weights do not. The field is a label, and nothing in the engine couples it to the bytes it names.
For most of the history of serving, that would be a curiosity rather than a problem, because the weights of a running model did not move. You built an image, you pinned a digest, you rolled out, and the question "which model produced this output" was answered by deployment state. What has changed in the last few months, and what shipped more of itself in the fortnight to 1 October, is that the parameters of a live vLLM engine are now a writable surface. There is a documented, pluggable weight transfer system with five backends — NCCL broadcast, CUDA IPC, sparse NCCL patches, sharding-aware NCCL M2N reshard, and a pull-based NIXL path over Ray Direct Transport — driven by a four-phase protocol of initialise, start, update, finish, where the update phase "may be called one or more times" for chunked transfers. It is not confined to a research harness. The documented way to switch it on in online serving is a flag: vllm serve my-model --weight-transfer-config '{"backend": "nccl"}'.
The reason this matters more than a housekeeping complaint is what the project's own documentation says about requests that are already running when a transfer lands. vLLM's pause API offers three modes — abort, wait, and keep — and keep is the one that makes asynchronous reinforcement learning possible, because it freezes in-flight requests instead of discarding them. The async RL page states the consequence without flinching: requests paused with mode="keep" "will produce tokens from the old weights before the pause and tokens from the new weights after resume." If the KV cache is retained rather than cleared, it adds, "some tokens in context may still reflect the old weights (stale KV cache)."
So a single completion can be the joint product of two parameter sets, with the boundary falling wherever the trainer's step happened to finish. The engine knows this happened. It has a field in which to say so. The field is a string the trainer may or may not have set, it is global rather than per-request, and the only ways to read it are a dev-mode GET /weight_info and a gRPC GetWeightVersion on a control service documented as being for "trusted sidecars". No completion carries it. Nothing in the metrics carries it. The question an architect would ask after the fact — which weights produced this answer — can be asked of the server, about this moment, and never of the answer.
The delta backend makes the bookkeeping problem concrete. Sparse NCCL sends weight patches as names, flat indices and values in checkpoint coordinates; it is explicitly "a delta backend: each call supplies replacement patches rather than a stable stream of the model's parameters," and the documentation assigns ownership of the consequences plainly — "the caller owns export/diff state and restart or reseed after a partial failure." What is resident in the engine, then, is a base checkpoint plus a sequence of patches, and the only record of which patches were applied lives in the process that sent them. The pull-based RDT engine is equally candid about how fragile a recorded plan is: it rejects expert-parallel load balancing outright, because EPLB "rearranges experts at runtime, which invalidates the recorded plan."
Now put that beside what landed on 30 September. vLLM merged a weight checker: an endpoint that hashes every model parameter with SHA-256 and returns a digest per shard, keyed dp{dp}:pp{pp}:pcp{pcp}:tp{tp}:{tensor_name}, so a caller can save a baseline, reset the weights, transfer them back and confirm they match. This is the real article — identity derived from the bytes rather than asserted about them. Read its limitations section and you find the price. It needs VLLM_SERVER_DEV_MODE=1. The engine must stay paused from checksum to compare, because "the weights are invalid between reset and the transfer, and EPLB moves experts while serving." The endpoint keeps no state, so the caller holds the baseline. Buffers, draft models and LoRA adapters are not checked. And hashing "copies every weight to CPU, so keep it off latency-sensitive paths."
In one fortnight, the project built both the honest answer to what is loaded and the reason nobody will ask it under traffic. The cheap identity is a label. The true identity is an audit that costs you the service.
The strongest objection is that this is training infrastructure being read as though it were a customer-facing endpoint. Rollout engines are not production. The HTTP control surface sits behind a development-mode environment variable. In an RL cluster the trainer is both the only writer of weights and the only consumer of outputs, so it already knows which step each sample belongs to, and nothing of consequence is lost by leaving the engine's version field at default. Evaluation, meanwhile, happens on exported checkpoints, which do have digests. On this reading the mutable engine is a scratch surface inside a loop, and demanding artifact discipline of it is a category error.
That objection is right about where the code is used today and wrong about where the ambiguity bites. The place it bites hardest is inside the loop, not outside it. In online RL the rollout is the training data, and a sequence that straddles a weight sync is a sample attributed to a policy that never generated it — the documentation says that is precisely what keep mode produces. The only thing that could mark the boundary is the trainer's own bookkeeping, which is the same bookkeeping that types the version string, and the protocol gives it no obligation to type anything. A system that cannot distinguish a sample from its own policy is not a governance problem; it is a measurement problem, and it sits upstream of every number the run reports. The second half of the objection weakens by the month. The same lifecycle is now exposed over the Rust frontend's gRPC control service, ServerInfo.rl_capabilities advertises weight_transfer_enabled and the configured backend so that a caller can discover the facility, and a dev-mode flag is a deployment default rather than an architecture. Capabilities that are advertised get consumed. Continuous post-training on a serving fleet is not a speculative destination; it is the direction every one of these merges points.
Three weeks ago, writing about packaging models as OCI artifacts, the argument here was that custody is not provenance — that a registry can prove a set of bytes has not changed while every property you would govern the model on remains an optional string the packager typed. This is the same screw, one turn further, and it is worth being clear about the difference. There, an artifact existed and its metadata was unverified. Here, at the moment of service, there is no artifact at all: there is a process whose parameters were last written over a fabric, at a time, by a peer, under a label nobody validated. Signing a checkpoint is a real control. It simply does not reach the thing that answered.
What follows for anyone running post-training against a live engine is unglamorous and cheap. Make the version mandatory in your own wrapper and derive it rather than type it — the trainer's commit and optimiser step, not a hand-set string. Record it per request at admission, not by racing a global GET against generation. Treat pause(mode="keep") as a change event that either gets forbidden on paths you must attribute, or gets a boundary marker written into the sample record. Run the weight checker as a release gate with the baseline stored alongside the eval results, and accept the pause it costs. None of that is new engineering; all of it is the kind of thing that only gets built after someone asks a question the system cannot answer.
Everything above is drawn from one project's own public record — its documentation, protocol definitions, source and commit history, read on main. That is a real limit on the piece: the network available for this run reached almost nothing else, and the competing serving stacks are not corroborated here. It is also the evidence a reviewer can check line by line, which is the trade worth making.
The distance between the two mechanisms is the whole argument. A SHA-256 over every parameter is what identity looks like when a system means it. A nullable string defaulting to default is what identity looks like when a system needs the throughput. Both now ship in the same server, and only one of them is on the fast path.
What this is argued from
Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.
- vllm/v1/engine/core.py, weight version accessors on main
- Async Reinforcement Learning, pause modes and the keep-mode contract
- Weight Transfer, the four-phase protocol and the HTTP control endpoints
- Weight Checker, SHA-256 verification of an RL weight update
- rust/proto/control.proto, the RL lifecycle control service
- Sparse NCCL, checkpoint-coordinate delta weight patches
- Sharded RDT Engine, pull-based NIXL weight transfer
- Native RL APIs in vLLM
Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.