A text watermark is a statistical signal hidden in which tokens a model chooses. It is not a visible label and not metadata: it lives in the words themselves, so it survives copy-paste into a plain text box where a C2PA manifest or an HTTP header would be lost. The provider holds a secret key; during sampling the key nudges the model towards a pseudo-random subset of the vocabulary that changes at every position; later, anyone with the key can count how often a text landed in those subsets and compute how unlikely that count would be for text written without the key.

This article is about the mechanism and the statistics: how the green-list scheme biases logits, how detection turns into a z-test, how distortion-free schemes such as Gumbel-max sampling and SynthID Text differ, how watermarks are attacked, and how to run a detector without harming people with false accusations. The wider system a watermark plugs into, with citations, audit logs and a verifier service, is covered in LLM output provenance architecture.

Why watermark, and what it cannot prove

Three tools answer the question 'did our model write this?' and they fail differently. Post-hoc classifiers guess from style; they need no cooperation from the generator but have high error rates, drift with every new model and are biased against non-native writers. Metadata and signed manifests, described in C2PA for AI-generated content, are exact but are stripped by any copy of the raw text. Watermarks sit between them: they require the generator to cooperate at sampling time, they give a measurable false positive rate rather than a guess, and they degrade gradually under editing instead of vanishing.

The cost is that a watermark only marks text from models whose sampling you control. Anyone running open weights can sample without it, so a missing watermark proves nothing. Treat it as evidence that a specific text came from your service, never as a general AI detector.

Architecture

Sampling-time watermark: the key biases generation, the same key tests the textModel logitsone score per tokenSeeded partitionPRF(key, prev tokens)Biased sampling+delta on green listOutput textto userRecompute listssame key, same PRFCount green hitsdedupe repeated n-gramsz-score / p-valuethreshold, never verdictsuspect text (possibly paraphrased, truncated, mixed with human text)Key serviceKMS, rotation, auditDetection needs no model call: only the tokenizer, the key and the hash scheme.
The same key and hash scheme drive sampling and detection; the detector never calls the model.

The green-list scheme in code

The scheme from Kirchenbauer and colleagues (ICML 2023) is the easiest to understand and still the reference point. At each step it hashes the previous token with the secret key, uses that hash to seed a random generator, and splits the vocabulary into a green list holding a fraction gamma of tokens and a red list with the rest. It then adds a constant delta to the logits of every green token before softmax. When the model is uncertain, that small push makes green tokens much more likely; when the model is confident, the push changes nothing, which protects quality on facts and syntax. The paper's typical setting is gamma 0.25 and delta 2.0.

import hashlib, numpy as np

def green_mask(key: bytes, prev_token: int, vocab: int, gamma: float) -> np.ndarray:
    seed = int.from_bytes(hashlib.sha256(key + int(prev_token).to_bytes(4, "big")).digest()[:8], "big")
    rng = np.random.default_rng(seed)
    mask = np.zeros(vocab, dtype=bool)
    mask[rng.permutation(vocab)[: int(gamma * vocab)]] = True
    return mask

def watermark_logits(logits, prev_token, key, gamma=0.25, delta=2.0):
    return logits + delta * green_mask(key, prev_token, logits.shape[-1], gamma)

def detect(tokens, key, vocab, gamma=0.25):
    seen, hits = set(), 0
    for prev, tok in zip(tokens, tokens[1:]):
        if (prev, tok) in seen:          # repeated pairs would inflate the count
            continue
        seen.add((prev, tok))
        hits += green_mask(key, prev, vocab, gamma)[tok]
    t = len(seen)
    z = (hits - gamma * t) / np.sqrt(t * gamma * (1 - gamma))
    return t, hits, z

Detection needs only the tokenizer and the key, not the model. Under the null hypothesis (text written without the key) each scored token is green with probability gamma, so the green count is binomial and the z-score is approximately standard normal. The deduplication line matters: repeated phrases, boilerplate and code produce the same (context, token) pair many times, and counting each repetition as independent evidence inflates z on perfectly human text.

Worked example: reading a z-score

Suppose a moderator submits a 200-token essay. With gamma 0.25 an unwatermarked text is expected to have 50 green tokens, with a standard deviation of the square root of 200 x 0.25 x 0.75, which is 6.12. The detector counts 95 green tokens after deduplication. The z-score is (95 - 50) / 6.12, which is 7.35. The one-sided normal tail beyond 4 is about 3 in 100,000, and beyond 7.35 it is far smaller, so this text almost certainly passed through the watermarked sampler.

Now suppose a student pasted 60 watermarked tokens into a 340-token essay they wrote themselves. If the watermarked part scores about 0.6 green and the rest about 0.25, the total is roughly 36 + 85 = 121 green of 400 expected 100, a standard deviation of 8.66 and a z of 2.4. A whole-document test misses it. A sliding-window test over 100-token spans finds the hot region, at the price of many more tests, so the threshold must be corrected for the number of windows or the false positive rate silently multiplies.

Entropy and the seeding window

The signal per token depends on how much freedom the model had. If the model puts 0.99 probability on one token, a delta of 2 barely moves it, and that position carries almost no evidence. This is why watermarks are weak on code, on lists of facts, on short answers and on low-temperature output, and strong on open-ended prose. The practical consequence is a minimum length: publish the number of scored tokens your detector needs for a target power, measured on your own traffic, and refuse to give a result below it.

The context used to seed the partition is the second knob. Seeding on one previous token is robust, because an edit only disturbs one or two positions, but the green lists are easy to learn from enough samples, which enables spoofing. Seeding on a longer window of previous tokens hides the key better but every edit now destroys the signal over the whole window. There is no setting that is both maximally robust and maximally secret.

Distortion-free schemes and SynthID Text

The green-list scheme changes the output distribution, which slightly lowers quality and can be measured. Distortion-free schemes keep each token's marginal distribution exactly the model's, and instead correlate the random choices with the key. Scott Aaronson's Gumbel-max proposal draws a keyed pseudo-random number r for every vocabulary entry and picks the token that maximises r to the power 1/p; averaged over keys this is ordinary sampling, but the chosen tokens have suspiciously large r values, and the detector sums minus log(1 - r) over the text. Kuditipudi and colleagues extended the idea with a fixed key sequence and an edit-distance alignment so the test survives insertions and deletions.

SynthID Text, published by Google DeepMind in Nature in 2024, uses tournament sampling: several candidate tokens drawn from the model compete in rounds scored by keyed g-functions, and the winner is emitted. It is available in Hugging Face Transformers, where you pass a SynthIDTextWatermarkingConfig to generate() through watermarking_config.

from transformers import AutoModelForCausalLM, AutoTokenizer, SynthIDTextWatermarkingConfig

cfg = SynthIDTextWatermarkingConfig(
    keys=load_keys_from_kms(),     # 20-30 random integers; this IS the secret
    ngram_len=5,                   # context window for seeding; larger = more detectable, more brittle
)
out = model.generate(**tok(prompts, return_tensors="pt", padding=True),
                     watermarking_config=cfg, do_sample=True, max_new_tokens=512)

Its detector is a trained classifier per model and configuration rather than a closed-form z-test, so calibrate its threshold on your own human and watermarked samples before using its scores.

Attacks

AttackWhat it doesEffectMitigation
ParaphraseRewrite with another model (for example DIPPER)Removes most of the signal on short textsLonger texts, windowed tests, accept limits
Emoji attackAsk for a token after every word, then delete themBreaks every seeding contextNone in-scheme; it is why no claim of robustness is absolute
DilutionMix watermarked spans into human textLowers whole-document zSliding windows with multiple-test correction
SpoofingLearn green lists from many outputs, then write marked textFrames a human or your serviceLonger seeding context, key rotation, rate-limited detection
Detector oracleQuery the detector repeatedly while editingFinds minimal edits that clear itAuthenticate and rate-limit the detection API
Open weightsGenerate without the watermarkNothing to detectNone; absence is not evidence

Spoofing deserves attention: work on watermark stealing (Jovanovic and colleagues, 2024) showed that a modest number of queries can let an attacker estimate green lists and produce text that tests as watermarked. A watermark that can be forged cannot be used to accuse anyone, which is a strong argument for treating it as one signal among several.

Operating a watermark

The key is the whole security of the scheme. Store it in a KMS, give the sampler and the detector separate service identities, and version it: every generated response should log the key version so a detector can test against the right key after rotation. Rotating keys limits the damage of a leak and the value of stealing, but every old key must remain available for detection for as long as you want to be able to answer questions about old text.

Expose detection as an internal, authenticated API that returns the number of scored tokens, the z-score or p-value, the key version and the window that triggered, never a bare 'AI-written' verdict. Rate-limit it per caller, because an open detector is an oracle for evasion. Measure false positives on a held-out corpus of human text from the domains you actually serve, including non-native writing and code, and put that measured rate in every report. Where outputs feed downstream filters, document where the watermark is applied relative to output guardrails, since a guardrail that rewrites text also weakens the mark.

Batching needs one detail: the partition depends only on the key and the context, so the processor can compute masks for every sequence in a batch at once, and caching masks for frequent contexts keeps the per-token overhead small even with large vocabularies. Test that streaming and non-streaming paths both apply the processor; a fast path that skips it ships unmarked text that no dashboard will notice.

Finally, decide what the score is used for. Platform moderation of coordinated campaigns, the use case in LLM-powered disinformation, aggregates many texts and tolerates per-text uncertainty. Academic or employment decisions about one person do not, and a watermark score alone should never be the basis for one.

Trade-offs

Delta and gamma trade detectability against quality: larger delta gives higher z at a given length but measurably shifts word choice. The seeding window trades robustness against spoofing. Distortion-free schemes remove the quality cost but make detection depend on more careful key handling, and some need an alignment search at detection time. Watermarking every response costs almost nothing at serving time, a hash and a vector add per token, but it commits you to keeping keys and a detector alive for years.

Failure modes

  • Counting repeated n-grams as independent evidence, which flags human boilerplate.
  • Testing a whole document when only a span is generated, which misses dilution.
  • Running many windows or many keys at a single-test threshold, which multiplies false positives.
  • Tokenising suspect text with a different tokenizer from the sampler, which scrambles every context.
  • Losing the mapping from response to key version after rotation.
  • Applying the watermark before a rewrite step, so the shipped text carries little signal.

What to do next

  1. Implement the green-list sampler and detector above on a small open model, and plot z against length for watermarked and human text from your own domain.
  2. Measure the minimum scored-token count that gives your target power at z 4, and enforce it.
  3. Add deduplication and a sliding-window mode with a corrected threshold.
  4. Try SynthID Text in Transformers and calibrate its detector on held-out samples.
  5. Run a paraphrase and an emoji attack against your own scheme and record the surviving z.
  6. Put the key in a KMS, log the key version with every response and rate-limit detection.
  7. Write down which decisions may use the score, and require a second signal for any decision about a person.
Key takeaway: A text watermark biases sampling with a secret key so that the key holder can later run a statistical test on the words alone. It gives a measurable false positive rate on text from your own service, weakens on short, low-entropy or paraphrased text, can be stolen and spoofed, and proves nothing when absent. Run it as a keyed, logged, rate-limited signal with published thresholds, not as a verdict on a person.