A text watermark is a statistical signal hidden in which tokens a model chooses. It is not a visible label and not metadata: it lives in the words themselves, so it survives copy-paste into a plain text box where a C2PA manifest or an HTTP header would be lost. The provider holds a secret key; during sampling the key nudges the model towards a pseudo-random subset of the vocabulary that changes at every position; later, anyone with the key can count how often a text landed in those subsets and compute how unlikely that count would be for text written without the key.
This article is about the mechanism and the statistics: how the green-list scheme biases logits, how detection turns into a z-test, how distortion-free schemes such as Gumbel-max sampling and SynthID Text differ, how watermarks are attacked, and how to run a detector without harming people with false accusations. The wider system a watermark plugs into, with citations, audit logs and a verifier service, is covered in LLM output provenance architecture.
Why watermark, and what it cannot prove
Three tools answer the question 'did our model write this?' and they fail differently. Post-hoc classifiers guess from style; they need no cooperation from the generator but have high error rates, drift with every new model and are biased against non-native writers. Metadata and signed manifests, described in C2PA for AI-generated content, are exact but are stripped by any copy of the raw text. Watermarks sit between them: they require the generator to cooperate at sampling time, they give a measurable false positive rate rather than a guess, and they degrade gradually under editing instead of vanishing.
The cost is that a watermark only marks text from models whose sampling you control. Anyone running open weights can sample without it, so a missing watermark proves nothing. Treat it as evidence that a specific text came from your service, never as a general AI detector.
Architecture
The green-list scheme in code
The scheme from Kirchenbauer and colleagues (ICML 2023) is the easiest to understand and still the reference point. At each step it hashes the previous token with the secret key, uses that hash to seed a random generator, and splits the vocabulary into a green list holding a fraction gamma of tokens and a red list with the rest. It then adds a constant delta to the logits of every green token before softmax. When the model is uncertain, that small push makes green tokens much more likely; when the model is confident, the push changes nothing, which protects quality on facts and syntax. The paper's typical setting is gamma 0.25 and delta 2.0.
import hashlib, numpy as np
def green_mask(key: bytes, prev_token: int, vocab: int, gamma: float) -> np.ndarray:
seed = int.from_bytes(hashlib.sha256(key + int(prev_token).to_bytes(4, "big")).digest()[:8], "big")
rng = np.random.default_rng(seed)
mask = np.zeros(vocab, dtype=bool)
mask[rng.permutation(vocab)[: int(gamma * vocab)]] = True
return mask
def watermark_logits(logits, prev_token, key, gamma=0.25, delta=2.0):
return logits + delta * green_mask(key, prev_token, logits.shape[-1], gamma)
def detect(tokens, key, vocab, gamma=0.25):
seen, hits = set(), 0
for prev, tok in zip(tokens, tokens[1:]):
if (prev, tok) in seen: # repeated pairs would inflate the count
continue
seen.add((prev, tok))
hits += green_mask(key, prev, vocab, gamma)[tok]
t = len(seen)
z = (hits - gamma * t) / np.sqrt(t * gamma * (1 - gamma))
return t, hits, zDetection needs only the tokenizer and the key, not the model. Under the null hypothesis (text written without the key) each scored token is green with probability gamma, so the green count is binomial and the z-score is approximately standard normal. The deduplication line matters: repeated phrases, boilerplate and code produce the same (context, token) pair many times, and counting each repetition as independent evidence inflates z on perfectly human text.
Worked example: reading a z-score
Suppose a moderator submits a 200-token essay. With gamma 0.25 an unwatermarked text is expected to have 50 green tokens, with a standard deviation of the square root of 200 x 0.25 x 0.75, which is 6.12. The detector counts 95 green tokens after deduplication. The z-score is (95 - 50) / 6.12, which is 7.35. The one-sided normal tail beyond 4 is about 3 in 100,000, and beyond 7.35 it is far smaller, so this text almost certainly passed through the watermarked sampler.
Now suppose a student pasted 60 watermarked tokens into a 340-token essay they wrote themselves. If the watermarked part scores about 0.6 green and the rest about 0.25, the total is roughly 36 + 85 = 121 green of 400 expected 100, a standard deviation of 8.66 and a z of 2.4. A whole-document test misses it. A sliding-window test over 100-token spans finds the hot region, at the price of many more tests, so the threshold must be corrected for the number of windows or the false positive rate silently multiplies.
Entropy and the seeding window
The signal per token depends on how much freedom the model had. If the model puts 0.99 probability on one token, a delta of 2 barely moves it, and that position carries almost no evidence. This is why watermarks are weak on code, on lists of facts, on short answers and on low-temperature output, and strong on open-ended prose. The practical consequence is a minimum length: publish the number of scored tokens your detector needs for a target power, measured on your own traffic, and refuse to give a result below it.
The context used to seed the partition is the second knob. Seeding on one previous token is robust, because an edit only disturbs one or two positions, but the green lists are easy to learn from enough samples, which enables spoofing. Seeding on a longer window of previous tokens hides the key better but every edit now destroys the signal over the whole window. There is no setting that is both maximally robust and maximally secret.
Distortion-free schemes and SynthID Text
The green-list scheme changes the output distribution, which slightly lowers quality and can be measured. Distortion-free schemes keep each token's marginal distribution exactly the model's, and instead correlate the random choices with the key. Scott Aaronson's Gumbel-max proposal draws a keyed pseudo-random number r for every vocabulary entry and picks the token that maximises r to the power 1/p; averaged over keys this is ordinary sampling, but the chosen tokens have suspiciously large r values, and the detector sums minus log(1 - r) over the text. Kuditipudi and colleagues extended the idea with a fixed key sequence and an edit-distance alignment so the test survives insertions and deletions.
SynthID Text, published by Google DeepMind in Nature in 2024, uses tournament sampling: several candidate tokens drawn from the model compete in rounds scored by keyed g-functions, and the winner is emitted. It is available in Hugging Face Transformers, where you pass a SynthIDTextWatermarkingConfig to generate() through watermarking_config.
from transformers import AutoModelForCausalLM, AutoTokenizer, SynthIDTextWatermarkingConfig
cfg = SynthIDTextWatermarkingConfig(
keys=load_keys_from_kms(), # 20-30 random integers; this IS the secret
ngram_len=5, # context window for seeding; larger = more detectable, more brittle
)
out = model.generate(**tok(prompts, return_tensors="pt", padding=True),
watermarking_config=cfg, do_sample=True, max_new_tokens=512)Its detector is a trained classifier per model and configuration rather than a closed-form z-test, so calibrate its threshold on your own human and watermarked samples before using its scores.
Attacks
| Attack | What it does | Effect | Mitigation |
|---|---|---|---|
| Paraphrase | Rewrite with another model (for example DIPPER) | Removes most of the signal on short texts | Longer texts, windowed tests, accept limits |
| Emoji attack | Ask for a token after every word, then delete them | Breaks every seeding context | None in-scheme; it is why no claim of robustness is absolute |
| Dilution | Mix watermarked spans into human text | Lowers whole-document z | Sliding windows with multiple-test correction |
| Spoofing | Learn green lists from many outputs, then write marked text | Frames a human or your service | Longer seeding context, key rotation, rate-limited detection |
| Detector oracle | Query the detector repeatedly while editing | Finds minimal edits that clear it | Authenticate and rate-limit the detection API |
| Open weights | Generate without the watermark | Nothing to detect | None; absence is not evidence |
Spoofing deserves attention: work on watermark stealing (Jovanovic and colleagues, 2024) showed that a modest number of queries can let an attacker estimate green lists and produce text that tests as watermarked. A watermark that can be forged cannot be used to accuse anyone, which is a strong argument for treating it as one signal among several.
Operating a watermark
The key is the whole security of the scheme. Store it in a KMS, give the sampler and the detector separate service identities, and version it: every generated response should log the key version so a detector can test against the right key after rotation. Rotating keys limits the damage of a leak and the value of stealing, but every old key must remain available for detection for as long as you want to be able to answer questions about old text.
Expose detection as an internal, authenticated API that returns the number of scored tokens, the z-score or p-value, the key version and the window that triggered, never a bare 'AI-written' verdict. Rate-limit it per caller, because an open detector is an oracle for evasion. Measure false positives on a held-out corpus of human text from the domains you actually serve, including non-native writing and code, and put that measured rate in every report. Where outputs feed downstream filters, document where the watermark is applied relative to output guardrails, since a guardrail that rewrites text also weakens the mark.
Batching needs one detail: the partition depends only on the key and the context, so the processor can compute masks for every sequence in a batch at once, and caching masks for frequent contexts keeps the per-token overhead small even with large vocabularies. Test that streaming and non-streaming paths both apply the processor; a fast path that skips it ships unmarked text that no dashboard will notice.
Finally, decide what the score is used for. Platform moderation of coordinated campaigns, the use case in LLM-powered disinformation, aggregates many texts and tolerates per-text uncertainty. Academic or employment decisions about one person do not, and a watermark score alone should never be the basis for one.
Trade-offs
Delta and gamma trade detectability against quality: larger delta gives higher z at a given length but measurably shifts word choice. The seeding window trades robustness against spoofing. Distortion-free schemes remove the quality cost but make detection depend on more careful key handling, and some need an alignment search at detection time. Watermarking every response costs almost nothing at serving time, a hash and a vector add per token, but it commits you to keeping keys and a detector alive for years.
Failure modes
- Counting repeated n-grams as independent evidence, which flags human boilerplate.
- Testing a whole document when only a span is generated, which misses dilution.
- Running many windows or many keys at a single-test threshold, which multiplies false positives.
- Tokenising suspect text with a different tokenizer from the sampler, which scrambles every context.
- Losing the mapping from response to key version after rotation.
- Applying the watermark before a rewrite step, so the shipped text carries little signal.
What to do next
- Implement the green-list sampler and detector above on a small open model, and plot z against length for watermarked and human text from your own domain.
- Measure the minimum scored-token count that gives your target power at z 4, and enforce it.
- Add deduplication and a sliding-window mode with a corrected threshold.
- Try SynthID Text in Transformers and calibrate its detector on held-out samples.
- Run a paraphrase and an emoji attack against your own scheme and record the surviving z.
- Put the key in a KMS, log the key version with every response and rate-limit detection.
- Write down which decisions may use the score, and require a second signal for any decision about a person.