A confident model cannot sign its work
vLLM now ships a text watermark that provably does not change what a model says. The guarantee is real. The accounting behind it is not being read: the signal is made out of randomness the sampler was already spending, so it is loudest on prose and nearly silent on code.
The argumentA distortion-free watermark is built from the sampler's own randomness, so its detection strength measures how much freedom the model had rather than whether a machine wrote the text, and the output we most want attributed is the output with the least freedom to spare.
Two numbers from the same set of measurements, taken with the same model, the same key and the same detector. On creative writing, the text watermark that shipped in vLLM 0.30.0 is caught almost every time by about a hundred distinct scored tokens, at a one per cent false positive rate. On MBPP, a Python programming benchmark, it reaches roughly 69% at four hundred scored tokens, and about 43% if the detector has to try a hundred candidate configurations before it finds the right one. Nothing about the watermarking scheme differs between those two lines. What differs is how much the model had left to decide.
That gap is the whole subject. A distortion-free text watermark is not something added to the text. It is a reallocation of randomness the sampler was going to spend anyway, which means its strength is a measurement of the model's freedom rather than of the text's origin. The corollary is uncomfortable and almost never stated: the output an organisation most wants to be able to attribute later, structured answers, code, short replies, anything the model was nearly certain about, is the output with the least room to carry a mark.
What the trick actually does
Start with ordinary sampling. The model produces a logit for every token in the vocabulary; softmax turns those into probabilities \( p_v \); you draw one token according to those probabilities. The Gumbel-max trick reaches the same place by a different route. Draw an independent uniform \( U_v \sim \mathcal{U}(0,1) \) for every token in the vocabulary, convert each into Gumbel noise,
add it to the log-probabilities, and take the largest:
The result is exactly categorical sampling: \( \mathbb{P}(v^\star = v) = p_v \). No approximation. vLLM's write-up notes that its Model Runner v2 sampling path already worked this way before watermarking existed, because adding noise to logits and reducing to an argmax parallelises better on a GPU than a softmax followed by a draw.
The watermark changes one thing. Instead of fresh randomness, each \( U_v \) comes from a pseudorandom function of a secret key \( K \), the last few generated tokens \( \mathbf{w} \), and the candidate token itself:
Four tokens of context by default. The key makes the pattern unguessable; including the candidate token gives every candidate its own value; including the recent context makes the values differ at almost every step. Because the context and the emitted token both survive into the output, anyone holding the key can recompute the exact numbers that were used.
Detection follows from that. For each token \( x_t \) in a suspect passage, recompute \( U_{x_t} \) and score it
If the text was not watermarked, the emitted token is independent of the keyed value, so \( U_{x_t} \) behaves like a uniform draw and \( s_t \) is exponential with mean one. Sum \( n \) of them and you get a Gamma distribution with shape \( n \) and scale one, which gives a calibrated one-sided p-value. Watermarked text drifts high, because Gumbel noise increases with \( U_v \), so the argmax systematically prefers candidates that drew large values. The implementation is exactly this, twelve lines of it: noise = -torch.log(-torch.log(uniforms)), then argmax(logits + noise), and on the detection side -torch.log1p(-uniforms) summed and pushed through a Gamma survival function.
Now look at where the signal comes from. The argmax only notices the noise when the logits leave it a choice. If one token dominates, \( \log p_v + G_v \) is maximised by that token whatever the noise did, the emitted token stops depending on \( U \), and the score collapses back to mean one. No excess, no evidence. This is not a bug or an implementation shortcut. It is the price of the guarantee: a scheme that provably never changes the output distribution can only express itself through choices that were already free. vLLM makes the boundary case explicit, and greedy requests at temperature=0 bypass watermarking entirely and log a warning.
So entropy is the carrier. The mark rides on the model's uncertainty, and where there is no uncertainty there is nothing to ride on.
The dial that was removed
It is worth seeing what this replaced. The influential earlier construction, the green-list scheme, hashes the preceding context to split the vocabulary into a favoured green set and a red set, then adds a bias \( \delta \) to the green logits. That scheme has a dial. Turn \( \delta \) up and the watermark gets easier to detect while the output visibly degrades; turn it down and quality returns while detection needs more text. Distortion against detectability, in one parameter an operator can set.
The Gumbel-max scheme refuses the trade. It fixes distortion at zero and, in doing so, removes the dial. What is left to spend is length and entropy, and an operator controls neither. Users decide how long the answers are and what kind of question gets asked. The quality evidence bears the guarantee out: on Qwen3.5-27B, GSM8K, MBPP and IFEval move by roughly a point in both directions between watermarked and unwatermarked runs, with overlapping error bars. The performance evidence is equally clean. Across eight keys and batch sizes from 1 to 256, mean matched throughput changed between −1.1% and +2.0%, with no consistent slowdown, on one H100 with three speculative tokens.
Then the compounding starts, and each layer of it removes marked tokens rather than weakening the mark.
Single-token non-distortion does not give sequence-level non-distortion. If a four-token context repeats, so do its keyed values, and choices that ordinary sampling would have made independently become correlated. vLLM's own worked example is a model that writes 1 + 1 + 1 +, returns to the context (1, +, 1, +), draws the same favourable value for 1, and loops. The fix is context deduplication: when a context repeats, sample normally and do not watermark that position. It costs almost nothing in throughput, at most 0.19% end to end, because the cost is not compute. The cost is coverage. Set the scope to all and the engine also skips any context that already appears in the prompt, which the documentation says plainly can leave little of a long-context answer marked when the prompt already contains the shape of the reply, for instance a tool result the model is extending.
Speculative decoding takes another share. Applying one watermark to both draft and target distributions shrinks their overlap and lowers the acceptance rate, so vLLM splits the key in two: accepted draft tokens use one, target recovery and bonus tokens the other. Acceptance is preserved and the detector must now score against both keys and blend the results, which dilutes both. The calibration is the interesting part. The generator routes 10% of tokens to the second key by default; the detector weights that key at 20%, because, as the write-up explains, drafts are accepted more often when they are easy, and easy means low entropy, and low entropy means little signal. The stack has already priced the thing this piece is about.
Finally, the detector needs to know what it is looking for: the key, the tokenizer, the algorithm and its parameters. The documentation concedes that in practice this information is often unavailable when someone checks a piece of text, and advises deployments to keep every configuration they have served, test against each and correct for multiple testing. That is where the 69% becomes 43%. Detection of text is not a public capability; it is a private membership test that a provider can run against its own logs.
What a learner should take from this
Keep two guarantees apart. Distortion-free is a claim about the output distribution. Detectable is a claim about a test's power. They are independent, and the paper-thin word "watermarked" covers both, so ask which one a system is selling you.
Then ask what the carrier is. For images and audio the carrier is perceptual redundancy, and there is a lot of it. For text the carrier is entropy, and there is only as much as the question left over. If you want to know how well a passage can be marked, look at the token log-probabilities and see how much the sampler was actually free to vary. That is a measurable quantity, not a guess.
And stop accepting detection accuracy as a single number. It is a surface over scored tokens, output domain and the number of candidate configurations tested, and the vLLM documentation goes further still: because a deployment serves one fixed key, repeated structures across documents reuse the same values, so the realised false positive rate is key-dependent even after deduplication. Measure it on your own unwatermarked traffic before you believe is_watermarked.
The strongest counterargument is that none of this is concealed. The engineers who built it wrote the limitations down, published the detection-power curves by domain, and noted that a detector exposing scores is also an oracle for scrubbing and forgery, and that the default pseudorandom function is not cryptographic and resists neither. Watermarking was never meant to stop a determined adversary, who in an open-weights stack simply omits the flag. It is provenance for cooperative traffic at scale, and prose is where the volume is.
That defence holds, and it is narrower than the claim being built on top of it. What shipped on 22 September, in a stable release of the serving engine much of the open ecosystem runs, is the first evidence that a provably distortion-free watermark costs nothing in quality or throughput. That evidence is real, and the argument I am making rests almost entirely on one project's own record, which is the limit of it. But the mark it produces is not a record that a machine wrote a passage. It is a record of how freely the machine chose, which means it is thinnest on exactly the output that now travels between machines with no human reading it, and thickest on the prose a person was going to judge anyway.
What this is argued from
Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.
- Watermarking in vLLM
- Text watermarking, feature documentation at tag v0.30.0
- Gumbel-max watermark generation and detection primitives
- WatermarkDetector, scoring and context deduplication
- WatermarkConfig defaults
- SamplingParams, the per-request watermarking field
- vllm release metadata, upload times for 0.29.0 and 0.30.0
Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.