A glitch token is an entry in a model's vocabulary that the model barely learned. The tokenizer can produce it, so text containing the right string turns into that token ID, but during training the ID appeared so rarely that its embedding never became meaningful. Feed one in and the model may refuse to repeat it, substitute a different word, produce unrelated text, or fall into a loop. In 2023 the token "SolidGoldMagikarp" made the phenomenon famous; in 2024 Land and Bartolo showed that under-trained tokens are common across many open model families and can be found automatically.

For security teams, glitch tokens matter less as a party trick than as one concrete instance of a broader problem this page's slug points at: token smuggling, where what a filter sees and what the model sees diverge at the token level. This article explains the mechanism from first principles, shows how to audit your own model's vocabulary with code, lays out what an attacker can and cannot get from these tokens, and gives defences and a checklist. It is written for people running or evaluating models, and gives no exploit strings for production systems.

How a token ends up untrained

A language model sees integer IDs, not characters. A tokenizer, usually byte-pair encoding or a unigram model, is trained first on some corpus and fixes the vocabulary; see how BPE builds a vocabulary. The model is then trained, often on a differently filtered or differently mixed corpus. Each ID owns one row of the input embedding matrix, and, in models that do not tie weights, one row of the output (unembedding) matrix.

A row only learns when its token appears. If a string was frequent in the tokenizer's corpus, perhaps a username repeated thousands of times in one scraped forum, but that forum was filtered out of the training data, the ID exists while its row receives almost no gradient. The row stays close to its initial values, shaped mainly by side effects such as weight decay and by updates that push all output rows together. At inference time the model receives a vector that does not mean anything it has learned. What happens next is unpredictable: the model may treat it like a nearby common token, ignore it, or derail.

The original 2023 investigation by Jessica Rumbelow and Matthew Watkins found a cluster of such tokens in the GPT-2/GPT-3 vocabulary, including several Reddit usernames, and observed that their embeddings sat unusually close to the centroid of the embedding space. Asked to repeat them, models of that era produced evasions or different words. Later tokenizers changed vocabularies, but the mechanism did not go away. Land and Bartolo's EMNLP 2024 paper, "Fishing for Magikarp", combined tokenizer analysis, weight-based indicators and prompting to detect under-trained tokens and found them to be prevalent across a diverse set of models, with counts that varied by model.

Where glitch tokens come from, and where defenders can catch themTokenizer corpusbuilds the vocabularyModel training datadifferent mix, filteredVocabularyevery entry gets an IDFrequent tokensmany gradient updatesRare or absent tokensembedding rows barely trainedUser inputtextTarget tokenizertext to IDsModelOutputrare ID: evasive, wrong or looping outputGuard classifierits own tokenizerMismatch: the guard and the target may not see the same tokensDefender checkpoints1. scan the vocabulary offline (round trip, weights, probes)2. flag confirmed IDs on the target's own token stream3. cap output tokens and detect loops
A vocabulary built on one corpus and a model trained on another leave some embedding rows untrained. At inference such IDs cause erratic output, and a guard model with a different tokenizer may not even see the same tokens. Defenders can catch them offline, at input, and at output.

Token smuggling: what glitch tokens give an attacker

"Token smuggling" is used loosely in the security literature for any trick where the content a defence inspects differs from the content the model acts on, at the level of tokens. Glitch tokens are one mechanism. Related ones have their own pages here: invisible Unicode characters and confusables, covered in Unicode smuggling defence, and optimised adversarial suffixes, covered in GCG universal adversarial attacks. The shared lesson is that a filter is only as good as its agreement with the model about what the input is.

Concretely, glitch tokens give an adversary four things, none of which is a reliable universal jailbreak:

  • Unpredictable behaviour. Inputs that steer the model into nonsense, loops or maximum-length outputs. In a metered API that is a cost and availability problem.
  • Filter disagreement. A guard classifier with a different tokenizer, or a keyword filter working on normalised text, may treat a string as harmless while the target model receives something it handles badly, or the reverse.
  • Fingerprinting. Which strings break a model reveals which tokenizer and often which model family sits behind an API, which helps an attacker choose other techniques.
  • Degraded safety behaviour. Because an untrained embedding pushes the model off its usual distribution, refusals and other trained behaviours may be less consistent around it. Treat this as a reason to test, not as an established bypass.

Auditing a vocabulary in three steps

You can audit any open-weights model you deploy. The approach follows the structure of Land and Bartolo's method: cheap tokenizer checks, then weight-based indicators to rank candidates, then prompting to verify. The code below is a starting point written against Hugging Face Transformers; adapt the details, especially the reference set, to your model.

Step 1, tokenizer analysis. Decode every ID and re-encode the string. IDs that never come back as themselves are unreachable from normal text. Many are harmless (partial UTF-8 byte sequences, special tokens), so this step classifies rather than condemns.

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL = "your-org/your-model"          # an open-weights model you are allowed to inspect
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype=torch.float32).eval()

special = set(tok.all_special_ids)
unreachable, partial_bytes = [], []
for i in range(len(tok)):
    if i in special:
        continue
    s = tok.decode([i])
    if "\ufffd" in s:                   # replacement char: an incomplete UTF-8 piece
        partial_bytes.append(i)
    elif tok.encode(s, add_special_tokens=False) != [i]:
        unreachable.append(i)
print(len(unreachable), "unreachable;", len(partial_bytes), "partial-byte tokens")

Step 2, weight indicators. Under-trained rows look alike, because they all received the same shared updates and little else. Take a reference set of IDs you are confident were never trained, such as reserved or unused special-token slots many tokenizers ship with, average their unembedding rows, remove the first principal component that all rows share, and score every token by cosine similarity to that average. High scores are candidates. Embedding-norm outliers are a cruder alternative when you have no reference set.

U = model.get_output_embeddings().weight.detach().float()     # [vocab, d_model]
ref_ids = torch.tensor(REFERENCE_UNUSED_IDS)                  # chosen by you, per model

centred = U - U.mean(dim=0)
_, _, V = torch.pca_lowrank(centred, q=1)
pc1 = V[:, 0]

def strip_pc1(x):
    return x - (x @ pc1).unsqueeze(-1) * pc1

ref_mean = strip_pc1(U[ref_ids].mean(dim=0))
score = torch.nn.functional.cosine_similarity(strip_pc1(U), ref_mean.unsqueeze(0), dim=1)
candidates = [i for i in score.topk(1000).indices.tolist()
              if i not in special and i not in partial_bytes]

Step 3, verify by prompting. A well-trained token can be repeated. Ask the model to copy the string and measure the probability it assigns to the token at the point where the copy should start. Very low probability on a repetition task is strong evidence the token is under-trained.

@torch.no_grad()
def repeat_prob(i):
    s = tok.decode([i])
    prompt = f'Copy the text between the quotes exactly.\nText: "{s}"\nCopy: "'
    ids = tok(prompt, return_tensors="pt").input_ids
    if i not in ids[0].tolist():
        return None                      # string did not tokenize back to i in context
    probs = torch.softmax(model(ids).logits[0, -1], dim=-1)
    return probs[i].item()

confirmed = [i for i in candidates if (p_ := repeat_prob(i)) is not None and p_ < 0.01]

Two cautions. The repetition prompt depends on the model: instruction-tuned models need their chat template, and the threshold should be calibrated on tokens you know are healthy. And the check above only works when the token survives in context; strings with leading spaces or unusual boundaries may tokenize differently inside the prompt, which is why the function returns None instead of guessing.

Worked example: a pre-deployment audit

An illustrative audit, with made-up but realistic proportions, of a 50,000-token vocabulary before a deployment:

StageResultWhat the team did
Round trip1,100 unreachable or partial-byte IDsLabelled byte fragments as expected; kept the rest as weak candidates
Weight indicatorTop 1,000 by similarity to the reserved-slot meanExcluded specials and byte fragments, leaving 870
Repetition probe60 IDs with copy probability below 0.01Recorded as the confirmed list, with decoded strings and scores
Manual reviewMostly scraped usernames, markup fragments and long code identifiersChecked a sample by hand to confirm the probe was not misfiring

The confirmed list then fed three controls: an input check that flags requests containing those IDs, a red-team suite that tries each one against the guard and the model, and a release gate that re-runs the audit whenever the tokenizer or checkpoint changes. The red-team results are what decide whether flagged inputs are blocked, stripped or merely logged.

Defences

  • Check IDs, not strings. Run the target model's own tokenizer on the final prompt, after templates and retrieval have been added, and compare the ID stream against the confirmed list. String blocklists miss variants and context-dependent tokenization.
  • Make guard and target agree. Where possible, run input classifiers on the same token stream as the model, or normalise text once and pass the normalised form to both. Where they must differ, include glitch candidates for both tokenizers in your red-team suite; see LLM red teaming.
  • Bound the damage at output. Set a maximum output length per route, detect repeated n-grams and stop generation early, and meter tokens per user so a looping response costs little.
  • Choose the response deliberately. Blocking every confirmed ID will reject legitimate text such as usernames, code identifiers or rare words in some languages. Stripping changes meaning. Logging and routing to a stricter policy is often the right middle ground; decide per product.
  • Fix it at the source if you train. Build the tokenizer on the same data mix the model will be trained on, check that every token occurs in training often enough, and when adding tokens during fine-tuning, initialise their embeddings sensibly and make sure they are trained. Pruning unused entries is another option; see tokenizer vocabulary trimming.
  • Re-audit on every change. A new checkpoint, a new tokenizer or a merged adapter changes the list. Make the audit a pipeline step with a stored report.

Failure modes

  • Trusting the indicator alone. Weight scores rank; they do not prove. Without the prompting step you will flag healthy tokens and block real users.
  • Bad reference set. If the reference IDs were actually trained, the mean points somewhere meaningful and the ranking is noise. Inspect the reference rows first.
  • Auditing the wrong artefact. Auditing the base model and deploying a fine-tune, or auditing one tokenizer revision and serving another, gives a list that no longer applies.
  • Closed models. Without weights you cannot run step 2. Probe behaviourally with the repetition test through the API, rate-limit yourself, follow the provider's terms, and report findings through their disclosure channel.
  • Over-reading old examples. The 2023 tokens are tied to a specific tokenizer generation. A string that broke one model says little about another until you test it.

Trade-offs

The audit is cheap compared with training: one pass over the vocabulary for round trips, one matrix computation, and a forward pass per candidate. The real costs are false positives and maintenance. Hard blocking gives the strongest guarantee and the worst user experience for legitimate rare strings; logging gives the best experience and relies on output limits to contain damage. For most deployments the balance is: always enforce output caps and loop detection, always log confirmed IDs, and block only on routes where erratic output is itself a harm, such as agents with tools, where an off-distribution step could trigger an action.

What to do next

  1. List every model and tokenizer revision you serve, including fine-tunes and adapters.
  2. For each open-weights model, run the three-step audit above and store the confirmed list with scores and decoded strings.
  3. Add an ID-level input check on the final prompt and decide per route whether matches are blocked, stripped or logged.
  4. Set output token caps and repetition detection on every route, and confirm they fire with a synthetic looping response.
  5. Add the confirmed tokens to your red-team suite against both the guard classifier and the target model.
  6. Make the audit a release gate that re-runs whenever a checkpoint or tokenizer changes, and read Land and Bartolo's paper for the full method.
Key takeaway: Glitch tokens are vocabulary entries whose embeddings were barely trained, usually because the tokenizer and the model saw different data. Audit every open-weights model you serve with round-trip checks, weight-based ranking and repetition probes, check the final prompt's token IDs rather than strings, keep guard and target tokenization aligned, cap and loop-check outputs, and re-run the audit whenever the model or tokenizer changes.