Every language model sees text through its tokenizer, and the tokenizer is chosen once, before pre-training, then lives forever in the embedding matrix. Get it wrong and you pay on every training step and every served token: longer sequences, wasted parameters, a model that cannot count digits, or a product that costs three times as much in some languages as in English.

The algorithms themselves are covered elsewhere on this site, starting from byte-pair encoding in depth. This article is the decision guide: which choices exist, what each one costs in parameters, compute and behaviour, how to measure candidates on your own data, and which failure modes to test before you commit.

Advertisement

The five choices

A tokenizer is a pipeline; each stage is a separate choiceNormalizeNFKC? lowercase?Pre-tokenizeregex, digits, spacesSubword modelBPE / Unigram / WPPost-processBOS, EOS, templateIDsto embeddingChoices made here are frozen into the embedding matrix and every checkpoint after it.Vocabulary size V drives:embedding params V x doutput logits 2 x V x d FLOPs/tokenshorter sequencesfewer decode stepscost grows with Vbenefit shrinks with VMeasure both on your own corpus; the right V is where the curves cross for your model size.
The tokenizer pipeline and the two sides of the vocabulary-size trade-off.

A tokenizer is not one decision but five, made at different pipeline stages.

AxisOptionsWhat it changes
Base unitUnicode characters, bytes, or bytes with fallbackWhether any input can be encoded without an unknown token
Subword algorithmBPE, WordPiece, Unigram LM, none (byte or character level)Which segmentations are learned and whether alternatives can be sampled
Vocabulary sizeRoughly 8k to 256kParameters, softmax cost, sequence length
Pre-tokenizationWhitespace, regex, digit splitting, noneWhich merges are allowed to cross word, number or punctuation boundaries
NormalizationNone, NFC, NFKC, lowercasing, accent strippingWhether the original text can be reproduced exactly

Published models show the range: GPT-2 used byte-level BPE with 50,257 entries; BERT base used WordPiece with 30,522 and lowercasing in its uncased variant; T5 used a SentencePiece Unigram model with 32,000; Llama 2 used SentencePiece BPE with 32,000 and byte fallback; Llama 3 moved to a tiktoken-style byte-level BPE with 128,256 entries. OpenAI's tiktoken encodings grew from about 100k entries in cl100k_base to about 200k in o200k_base. The direction of travel for large multilingual models is bigger vocabularies over bytes.

Algorithm and base unit

BPE starts from characters or bytes and repeatedly merges the most frequent adjacent pair. Encoding replays the merges, so it is deterministic and fast. WordPiece is similar but scores merges by likelihood gain, and marks word-internal pieces with ##. Unigram LM goes the other way: it starts from a large candidate vocabulary and prunes pieces that least reduce corpus likelihood; at encode time it picks the most probable segmentation, and can sample alternatives, which enables subword regularization during training. The detail lives in the Unigram LM tokenizer article.

In practice the algorithm matters less than the base unit and the pre-tokenization rules. Byte-level vocabularies (or character vocabularies with byte fallback) guarantee that every input round-trips with no unknown token, which matters for code, logs, emoji and any language underrepresented in training. Pure byte models with no subwords, such as ByT5, remove the tokenizer entirely but make sequences several times longer, which attention pays for quadratically.

Advertisement

Vocabulary size: the arithmetic

Vocabulary size is the one number everyone argues about, so do the arithmetic. With model width d and vocabulary V, the input embedding has V x d parameters, and an untied output projection has as many again. Computing logits costs about 2 x V x d floating-point operations per token.

Worked example, d = 4096:

VEmbedding paramsUntied in + outLogit FLOPs per token
32,000131M262M0.26 GFLOP
128,256525M1.05B1.05 GFLOP
256,0001.05B2.10B2.10 GFLOP

For an 8B-parameter model the 128k vocabulary costs about 1B parameters, roughly an eighth of the model, and the logit layer adds about 1 GFLOP to roughly 16 GFLOP of forward compute per token. For a 1B model with d = 2048, the same vocabulary is about 260M parameters per matrix, a quarter of the budget, which is why small models often keep smaller vocabularies or trim them, as described in vocabulary trimming.

The benefit side is compression. Suppose measurement shows the 128k tokenizer produces 15 percent fewer tokens than the 32k one on your corpus. Then a fixed training budget in tokens covers 15 percent more text, a fixed context window holds 15 percent more content, and generation needs 15 percent fewer decode steps. Compression gains are largest for languages and domains the small vocabulary served badly, so measure per language, not on the aggregate. Also note that a rare token gets few gradient updates: past the point where new entries are seen often in training, extra vocabulary is mostly undertrained rows.

Pre-tokenization and normalization

Pre-tokenization splits text into chunks before the subword model runs, and merges never cross a chunk boundary. GPT-2's pattern is a good reference:

import regex  # the third-party module, needed for \p{L} and \p{N}

GPT2_PAT = regex.compile(
    r"""'s|'t|'re|'ve|'m|'ll|'d| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)|\s+"""
)
print(GPT2_PAT.findall("Total: 12345 tokens, isn't it?"))
# ['Total', ':', ' 12345', ' tokens', ',', ' isn', "'t", ' it', '?']

Three rules here have big downstream effects. Leading spaces attach to the following word, so " tokens" and "tokens" are different tokens, which is why prompts that end with a trailing space behave oddly. Letters and digits are separated, so no token mixes them. And digit runs are left whole, so the BPE learns arbitrary chunks of numbers like 123 and 45, which makes digit-level arithmetic harder. Later patterns such as the one behind cl100k_base cap digit runs at three, and SentencePiece offers split_digits for one digit per token. For code-heavy models, also check how runs of spaces and tabs are handled, since indentation is a large share of code tokens.

Normalization should usually be minimal. NFKC folds compatibility characters, which helps search but makes the tokenizer lossy: full-width letters, ligatures and some math symbols no longer round-trip. Lowercasing is fine for a classifier and wrong for a generator.

Measuring candidates on your own corpus

Do not choose from published tables. Take a sample of your actual training or serving text, split by language and domain, and measure each candidate.

from transformers import AutoTokenizer

def profile(name, docs_by_slice):
    tok = AutoTokenizer.from_pretrained(name)
    out = {}
    for slice_name, docs in docs_by_slice.items():
        chars = words = tokens = roundtrip_fail = 0
        for d in docs:
            ids = tok.encode(d, add_special_tokens=False)
            chars += len(d); words += len(d.split()); tokens += len(ids)
            if tok.decode(ids) != d:
                roundtrip_fail += 1
        out[slice_name] = dict(
            chars_per_token=round(chars / tokens, 2),
            tokens_per_word=round(tokens / max(words, 1), 2),   # fertility
            roundtrip_fail=roundtrip_fail,
        )
    return out

for name in ["gpt2", "bert-base-uncased", "t5-small", "./my_tokenizer"]:
    print(name, profile(name, slices))

Read three things. Characters per token is compression; compare slices against your English or majority slice to see the cost multiplier other users pay. Fertility, tokens per whitespace word, is meaningless for languages without spaces, so use characters per token there. Round-trip failures reveal normalization and unknown-token loss; a generator should have none. Multilingual tokenization explains why the per-language gap appears and how sampling the training corpus closes it.

Training your own tokenizer

If no existing tokenizer fits, train one on a sample that mirrors your intended training mix, not the raw crawl. Two common routes:

# Hugging Face tokenizers: byte-level BPE
from tokenizers import Tokenizer, models, trainers, pre_tokenizers, decoders

tok = Tokenizer(models.BPE())
tok.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=False)
tok.decoder = decoders.ByteLevel()
trainer = trainers.BpeTrainer(
    vocab_size=64000,
    special_tokens=["<|endoftext|>", "<|pad|>"],
    initial_alphabet=pre_tokenizers.ByteLevel.alphabet(),
)
tok.train(["mix_sample.txt"], trainer)
tok.save("my_tokenizer.json")

# SentencePiece: Unigram with byte fallback and single-digit numbers
import sentencepiece as spm
spm.SentencePieceTrainer.train(
    input="mix_sample.txt", model_prefix="sp64k", vocab_size=64000,
    model_type="unigram", byte_fallback=True, split_digits=True,
    character_coverage=0.9995,
)

Reserve special tokens now, including spare unused ones, because adding rows later means resizing the embedding of every checkpoint. Weight the sample so low-resource languages and code are present at the share you want them to compress at, then re-run the profiler on held-out text.

Worked example: a trilingual assistant with code

Worked example. A team is pre-training a 3B model for an assistant that answers in English, Hindi and Spanish and writes Python. Candidates: a 32k SentencePiece BPE, a 64k byte-level BPE trained on their mix, and a 128k byte-level BPE. Their profile on held-out data, in characters per token, comes out roughly 4.1, 4.4 and 4.6 for English, but 2.0, 3.1 and 3.5 for Hindi.

At d = 3072, the 128k vocabulary costs about 394M parameters per matrix against 197M for 64k, so tied embeddings are mandatory and still spend 13 percent of the model on the table. The 64k tokenizer captures most of the Hindi gain, which is where the product's cost multiplier was worst, at half the embedding cost. They choose 64k, byte-level, three-digit number chunks, no normalization, and reserve 256 special-token slots.

What the tokenizer costs at serving time

The tokenizer keeps costing money after training. Serving cost, latency and context capacity are all denominated in tokens, so compression differences turn directly into product differences.

Price and quota. APIs bill per token. If Hindi text compresses at 2.0 characters per token and English at 4.1, a Hindi user pays about twice as much per character of text and hits token rate limits about twice as fast. Publish per-language cost estimates from your profiler, not from an English rule of thumb such as four characters per token.

Latency. Generation is one decode step per output token. Fewer tokens for the same text means proportionally less time to finish a response, and a larger vocabulary adds only a small logit-layer cost per step at typical model sizes. Prefill also shrinks with compression, which matters for long documents and retrieval-heavy prompts.

Context. A 32k-token window holds very different amounts of text depending on the language and the tokenizer. Size retrieval chunks, truncation limits and memory budgets in tokens measured with the production tokenizer, never in characters or words.

Speculative decoding and model families. A draft model must share the target model's vocabulary, or at least have a mapping, for its proposals to be verified token by token. Choosing one tokenizer across a family of model sizes keeps that option open and lets you reuse data pipelines, token-count budgets and evaluation harnesses unchanged.

Failure modes

  • Undertrained tokens. Entries present in the tokenizer's training text but rare in the model's training data get almost no updates and can produce erratic output. Train the tokenizer on the same mixture as the model, and scan for tokens with near-initial embedding norms.
  • Train/serve mismatch. A different tokenizer version, a missing BOS token or a changed chat template shifts every position. Pin the tokenizer file by hash next to the checkpoint; the serving side is covered in LLM tokenizer architecture.
  • Special-token injection. If user text containing a literal <|endoftext|> is encoded as the special token, users can end turns or forge roles. Encode user text with special tokens disabled.
  • Boundary artifacts. Trailing spaces, prompts that end mid-word and digits split unpredictably all change what the model sees. Test them explicitly.
  • Changing tokenizers mid-project. A new tokenizer invalidates embeddings and every token-count-based number: context lengths, cost estimates, data budgets. Decide once.

What to do next

  1. Build a held-out sample per language and domain that mirrors real traffic, a few megabytes each.
  2. Run the profiler on three to five candidate tokenizers and tabulate characters per token, fertility and round-trip failures per slice.
  3. Compute embedding and logit cost for each vocabulary at your model width, tied and untied.
  4. Pick the smallest vocabulary that closes the worst slice's gap to within your cost tolerance.
  5. Decide digit, whitespace and normalization rules deliberately, and write them down.
  6. If you train one, reserve spare special tokens and pin the tokenizer file by hash with every checkpoint.
  7. Add tests for special-token injection, trailing spaces and round-trip losslessness to CI.
  8. Read tokenizers and their consequences next for the conceptual background.
Key takeaway: A tokenizer is five frozen choices: base unit, algorithm, vocabulary size, pre-tokenization and normalization. Byte-level coverage and sensible digit and whitespace rules matter more than BPE versus Unigram. Vocabulary size trades embedding parameters and logit compute against shorter sequences, so compute both at your model width and measure compression per language on your own data. Pick the smallest vocabulary that fixes your worst slice, reserve special tokens, and pin the tokenizer with every checkpoint.