Every key you ever served
vLLM now watermarks generated text in the sampler itself, which is the right place for it. But detection needs the exact profile that produced the text, the key is the one field the config refuses to print, and the documented remedy spends your statistical confidence on your own deployment history.
The argumentWatermarking makes generated text testable only against the exact serving profile that produced it, so provenance is now a records problem rather than a property of the text, and the remedy the documentation recommends makes the evidence weaker the longer you have been operating.
Someone forwards you a paragraph and asks whether your system wrote it. This is no longer a hypothetical conversation; it is a Tuesday afternoon with a lawyer on the call. And as of last month you have, in principle, the instrument to answer: vLLM's v0.30.0 release, tagged on 22 September, shipped watermarked generation with a detector that works on token IDs alone and never touches the model weights. You do not need the GPUs back. You do not need the checkpoint. You need eight configuration values and a tokenizer.
One of those values is a secret integer, and the line of code that defines it is written specifically to make sure it never gets written down.
key: int = Field(ge=0, repr=False, exclude=True)
That is WatermarkConfig in vllm/config/watermarking.py. repr=False keeps the key out of the object's printed form; exclude=True keeps it out of serialisation. This is unambiguously correct engineering — a secret that appears in a config dump is a secret in your log aggregator by Friday. It is also, read as an architect rather than as a reviewer, the whole problem in one line. The one value your future provenance claim depends on is the one value the system is careful to forget.
The feature itself deserves respect before the criticism starts. The RFC that proposed it, opened in late August, makes an argument that is hard to improve on: watermarking "is a modification of the sampling process, which cannot be modified by downstream consumers." If marking output is a thing a provider owes anyone, it cannot be a wrapper. It has to be in the sampler, below the API, where no client can route around it. And that is exactly where it landed. The documentation puts GPUWatermarkSampler at the point of "final stochastic token selection after temperature, min-p, top-k, and top-p are applied" — the last place a token choice is still free.
The mechanism is Gumbel-max: a keyed pseudorandom function turns the secret, the previous context_width tokens and each candidate token into the noise used for categorical sampling. The signal is not metadata bolted to the response. It is which words came out. Nothing can strip it the way you would strip an EXIF tag, because there is nothing to strip.
So the hard part is done, and the hard part was never the part that will fail.
Everything below is drawn from vLLM's own documentation, configuration code and review history. That is one project's account of itself, which cuts both ways: the candour is unusual and worth saying so, and the omissions are the ones a project is least likely to narrate about its own work.
Look at what the detector actually requires. The documentation is unusually direct: detection must match generation "including the tokenizer, PRF, watermarking algorithm, algorithm-specific watermarking configuration, and key." Unpack that against the config object and you get a profile with real dimensionality — algorithm, alpha, context_width, deduplicate_contexts, deduplicate_contexts_max_history, prf, allow_target_only_watermarking, the key, and whichever tokenizer the model shipped with that quarter. Change any of them and the test comes back negative on text you generated yourself.
Then comes the sentence that should be read twice. "In practice, this information is often unavailable when checking a piece of text. Deployments should therefore retain the set of candidate configurations they have served, test the text against each candidate, and correct for multiple testing, for example with a Bonferroni correction to the resulting p-values."
That is the real architecture of this feature, stated plainly by the people who built it, and it is not a watermark. It is an archive with a statistical penalty attached.
Bonferroni is simple arithmetic: to hold a family-wise error rate, you divide your significance threshold by the number of comparisons. Rotate your watermark key quarterly — as you should, because it is a secret, and because the PRF that ships today, Philox4x32-10, is documented as "not a cryptographic PRF" with no "key-recovery or forgery resistance" — and after three years you hold twelve candidate profiles. A passage that would clear a threshold of 0.01 against a single known profile now has to clear roughly 0.00083. Change context_width once for robustness, try dual_key_gumbel for a quarter to get speculative decoding back, move to a model with a different tokenizer, and the denominator grows again.
This is the part worth arguing about, because it inverts an assumption nearly everyone is carrying. We have been told that watermarking makes text self-evidencing. It does the opposite. It makes text evidencing of a configuration, and the strength of the evidence is inversely proportional to how much configuration history you are carrying. Good key hygiene degrades your provenance. Routine model upgrades degrade your provenance. There is no component in this system, or in any system next to it, whose job is to hold that tension — and nobody has been asked to decide which of the two they would rather have.
It gets more specific than that, and the specifics fall on exactly the traffic that matters. The watermark is not uniformly present in your output; its density is set by decisions made for throughput. allow_target_only_watermarking lets you keep speculative decoding with an algorithm that does not natively support it, and the cost is named: the signal "is diluted in proportion to the share of output tokens supplied by accepted drafts." Your draft acceptance rate is a performance number that moves with load, quantisation and draft-model version. It is now also the dilution factor on your legal evidence. Requests at temperature=0 bypass watermarking entirely and "emit a warning once per worker" — once, per worker, for an unbounded number of unmarked responses. Beam search does not apply it at all. And if you set deduplicate_contexts to "all" to get non-distortion across a conversation, the documentation warns that "a prompt that already contains the structure of the answer such as a tool result the model extends, may therefore leave little of the answer watermarked." That is a description of agentic serving. The workload most likely to produce text somebody later disputes is the workload where the mark thins out most.
Then there is the boundary. In sampling_params.py the field reads watermarking: bool = True, and the per-request watermarking: false opt-out is accepted on the OpenAI-compatible APIs, the Rust chat and completions APIs, the Rust token API and the gRPC GenerateRequest. The documentation's instruction is correct and quietly damning: deployments that require watermarking "must restrict this field to trusted callers, or strip and validate it at the ingress boundary, so untrusted clients cannot opt out." Your compliance posture is now a request-field allowlist in a gateway. Those are maintained by whoever last touched the gateway.
And the detector cannot be shared. "Repeated queries can use this information to construct text that imitates watermarked output or to modify watermarked text so it is no longer detected." The reference server in examples/basic/online_serving/watermark_detection_server.py is forty lines and returns a p-value, which makes it both the instrument of proof and the instrument of forgery. So verification cannot be delegated outward, to a regulator, a journalist, a court-appointed expert, or the person whose work was allegedly imitated. Provenance here is structurally an in-house function, answered by the party with the most interest in the answer.
The strongest case against all of this is that none of it is a defect. It is an honest implementation of what the research actually supports, built by people who documented every limitation rather than burying it, and it is enormously better than the nothing that preceded it. Nobody promised a universal public detector, and the obligations driving this work — the RFC names the EU AI Act's Article 50(2) and California's SB 942, a citation I report as the proposer's rather than as my own reading of either text — are obligations on providers to mark their output, not to hand the world an oracle. A keyed, in-house, statistically honest detector is precisely the right tool for an abuse investigation or a discovery request. That is a serious position and I hold most of it.
Where it stops being sufficient is the distance between marking output and being able to prove which output was yours. The first is now a flag in a serving config. The second is a register: every watermark profile ever deployed, with the dates it was live, the key under custody that survives the cluster it ran on, the tokenizer version, the decoding settings that were in force. Nothing creates that register. Nothing validates it. No retention policy covers it, because no policy author has encountered the idea that a cryptographic secret from three years ago is load-bearing for a claim you have not yet needed to make. The sampler shipped. The filing cabinet did not.
This is becoming the characteristic failure of the whole evidence layer we are assembling. Attestation, signed model artefacts, SBOMs, and now watermarks: each presents as a property of a thing, and each turns out, on inspection, to be a pointer into records whose upkeep was assigned to nobody. The artefact is the part we build, because building is what we are good at. The archive is the part that decides whether any of it counts.
The signal rides in the tokens, where it is nearly impossible to remove. The proof rides in a list somebody has to remember to maintain. Two years of discussion have gone into how to mark the output. Almost none has gone into the question that will actually be asked in the room with the lawyer — not whether the text is watermarked, but whether anyone can still say with which key.
What this is argued from
Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.
- Text watermarking
- WatermarkConfig
- SamplingParams, watermarking field
- RFC, Native Text Watermarking Support
- Watermarked generation and detection (Gumbel-max algorithm)
- Dual-key gumbel-max watermarking for speculative decoding support
- vLLM v0.30.0 release notes
- Minimal reference server for watermark detection
Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.