A language model never reads text. It reads a sequence of integers, each one an index into a vocabulary of byte strings that a tokenizer learned from a corpus. Most write-ups explain how that vocabulary is trained. This article is about the other half: what the integer view does to the behaviour of a model you prompt, fine-tune and evaluate, and the bugs that appear where your strings meet those integers.
Every split and id below was produced by running OpenAI's open-source tiktoken library (version 0.13) with the cl100k_base and o200k_base encodings, so the splits are real rather than illustrative. The algorithms themselves are covered in how tiktoken works and the design decisions in tokenizer design choices; this page assumes only that you know a tokenizer turns text into ids and back.
What the model actually sees
Modern LLM tokenizers are mostly byte-level BPE. Text is first cut by a regular expression into pre-tokens (roughly: a word with its leading space, a run of punctuation, a short run of digits), each pre-token is turned into UTF-8 bytes, and learned merges are replayed to glue bytes into the longest known units. Because the base alphabet is all 256 byte values, there is no unknown token: anything can be encoded, rare strings simply cost more ids.
Three consequences matter more than the algorithm. First, the space belongs to the following word: hello is id 15339 in cl100k while hello is id 24748, and Hello and Hello are two more unrelated ids. The model learns four separate embeddings and has to discover that they mean nearly the same thing. Second, the same string can tokenize differently in context, because pre-token boundaries depend on neighbours. Third, ids are only meaningful for one vocabulary: o200k assigns hello the id 24912, so an id list is useless without the name and version of the tokenizer that made it.
| Input | cl100k_base pieces | o200k_base pieces |
|---|---|---|
strawberry | str | aw | berry | st | raw | berry |
strawberry | one token | one token |
12345678 | 123 | 456 | 78 | 123 | 456 | 78 |
The answer is 42 | The | answer | is | space | 42 | same shape |
你好 | 2 tokens | 1 token |
The prompt boundary
Generation continues from the last id of your prompt, so the model's next choice is conditioned on exactly where you cut. The classic failure is a trailing space. In cl100k, Q: capital of France?\nA: Paris ends with the pieces : and Paris. If your prompt template ends with A: including the space, the prompt ends with a lone space token (id 220), and the natural continuation Paris now starts with a space the model cannot repeat. It must emit a bare Paris after a lone space, a sequence it rarely saw in training. Quality drops in ways that look like the model being dim: odd capitalisation, a stray extra space, or a different answer.
The fix is not a universal rule like never end with a space, because pre-tokenizers differ. Both cl100k and o200k split digits away from the preceding space: The answer is 42 encodes as The, answer, is, a lone space, then 42. For numbers the trailing space is the canonical form; for words it is not. The robust practice is to end prompts on a natural boundary such as a newline or a colon, let the model produce the space, and test templates by checking that encoding the prompt and a typical answer separately gives the same ids as encoding them together.
Inference libraries can also repair the cut automatically with token healing: back off the last token or two of the prompt and constrain the first generated token to start with the removed characters. The mechanics belong to constrained decoding, explained in the math of guided generation. Healing is a safety net, not a substitute for templates that cut in the right place.
The training boundary: prompt, completion and the loss mask
The same boundary bites harder in supervised fine-tuning, because it is silent. A typical SFT example is a prompt and a completion, and you want loss only on the completion. The tempting implementation tokenizes the two strings separately, concatenates the id lists and masks the first list. That produces a token sequence the tokenizer would never produce for the joined text, so the model is trained on non-canonical sequences it will never see at inference, and the mask may cover the wrong characters.
Measured on the France example: with the prompt ending in A: and the completion Paris, separate encoding ends with :, , Paris while encoding the joined text ends with A, :, Paris, one token shorter. Without the trailing space the two methods agree. The safe pattern is to encode the full text once and derive the mask from character offsets, refusing any example whose boundary falls inside a token:
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
def build_example(prompt: str, completion: str, max_len: int = 4096):
full = prompt + completion
ids = enc.encode(full)
_, starts = enc.decode_with_offsets(ids) # char offset where each token starts
boundary = len(prompt)
if boundary not in starts and boundary != len(full):
raise ValueError(f"prompt/completion boundary splits a token at char {boundary}")
labels = [tok if start >= boundary else -100 # -100 = ignored by cross-entropy
for tok, start in zip(ids, starts)]
return ids[:max_len], labels[:max_len]Hugging Face fast tokenizers give the same information through return_offsets_mapping=True. For chat models the template adds role markers and end-of-turn tokens; render it with the model's own chat template, mask everything except assistant turns, and make sure the end-of-turn token is inside the loss, or the fine-tuned model forgets how to stop. Template and special-token handling are covered in the LLM tokenizer deep-dive.
The span boundary: tokens are not characters
Anything that maps model positions back to characters, such as highlighting cited evidence, token-level confidence, NER spans or attribution heatmaps, has to cope with tokens that do not align with characters. Byte-level BPE can cut a single character in two. In cl100k, 東京で会議 encodes to seven tokens and the first two are the two halves of the UTF-8 bytes for the first character; decode_with_offsets reports both starting at character 0, and the last character is split the same way. Decoding either half alone gives invalid UTF-8.
Two rules follow. When you stream, buffer bytes and emit only complete UTF-8 sequences; never decode token by token into strings. When you draw spans, convert token positions to byte ranges, widen each range to the enclosing character boundaries, then merge overlaps. Latin text hides the problem: Refund naïve café order 98765 splits as Ref, und, naï, ve, café, order, a space, 987, 65, all on character boundaries, so tests written in English pass and production in Japanese breaks.
Digits, spelling and why models miscount letters
Tokenization explains several behaviours people blame on reasoning. Spelling and letter counting: strawberry is a single id in both encodings, so the model never sees its letters; it can only have memorised facts about that id. Without the leading space the word becomes three pieces, and capitalised or misspelled forms differ again. Questions about characters are answered from association, not inspection.
Arithmetic: cl100k and o200k chunk digit runs left to right in groups of up to three, so 12345678 becomes 123, 456, 78 and 1234567 becomes 123, 456, 7. Place value is not aligned between numbers of different lengths: the last chunk holds the units digit in one number and the hundreds in another. Models learn around this, but carries and long multiplications remain error-prone.
Practical mitigations: route exact arithmetic, counting and string manipulation to a tool or code interpreter; when the model must operate on characters, present them separated by spaces or one per line so each becomes its own token; format numbers consistently in prompts and training data; and evaluate on inputs whose surface form matches production, because a benchmark of clean integers says little about invoice totals with thousands separators.
Counting, cost and special tokens
Context limits, rate limits and prices are all counted in tokens of the model's own tokenizer, and the ratio of tokens to characters varies by language, domain and vocabulary. Across our samples, 你好 costs two tokens in cl100k but one in o200k, and the o200k vocabulary is roughly twice as large (n_vocab reports 200,019 against 100,277, both including special tokens). Moving a product to a model with a different tokenizer changes every budget, even if the prompts stay identical. The disparity across languages is quantified in multilingual tokenization.
Count with the exact tokenizer for the exact model, and add the overhead the chat template inserts per message. Where a provider exposes a token-counting endpoint, prefer it for anything near the limit, because a local estimate of the wrong encoding can be off by tens of percent on code or non-Latin text. Keep a margin for the reply: the context window is shared by input and output.
Special tokens deserve a separate check. tiktoken refuses by default to encode text that contains a special-token string such as the end-of-text marker: the call raises ValueError unless you pass allowed_special or disallowed_special. With disallowed_special=() the same text becomes seven ordinary tokens; with allowed_special="all" it becomes the single control id 100257. User-supplied text must always take the first path, so that a pasted control string cannot end a turn or inject a role.
Worked example: a classifier that loses accuracy in production
Consider an illustrative case. A team fine-tunes a model to answer support questions with a one-word category, then serves it. Offline accuracy is 94 percent; in production it falls to 81. Working back through the boundaries finds three tokenization bugs, each of which you can check in an afternoon.
- Training boundary: the data builder encoded
Category:and the label separately. Every training example ended with a lone space token followed by a bare word, a sequence the serving path never produced. Rebuilding with the offset-based builder above, and dropping the trailing space from the template, recovered most of the gap. - Prompt boundary: the serving template still ended with the space, so even after retraining the first generated token was non-canonical. Ending the template at the colon fixed it.
- Evaluation: the scorer compared strings exactly, so outputs with a leading space were marked wrong in one harness and right in another. Normalise whitespace in the scorer, and log the generated ids as well as the text so you can see which form the model chose.
The general lesson: write a test that encodes each template with a representative answer both ways and asserts the id sequences match, run it in CI for every tokenizer you ship, and log ids alongside text for a sample of production traffic.
Failure modes
| Symptom | Tokenization cause | Fix |
|---|---|---|
| Odd first word, extra or missing space | Prompt ends mid pre-token (trailing space) | End templates on a newline or colon; token healing |
| Fine-tuned model worse than offline eval suggests | Prompt and completion encoded separately | Encode once, mask by character offsets |
| Model never stops after fine-tuning | End-of-turn token outside the loss mask | Include the end-of-turn token in labels |
| Mojibake when streaming | Decoding partial UTF-8 per token | Buffer bytes, emit complete characters |
| Highlight spans off by one in CJK text | Token cuts a character | Map through byte offsets, widen to characters |
| Context overflow after a model switch | Different vocabulary, different counts | Count with the target tokenizer, keep a margin |
| User text ends the turn | Special-token string encoded as a control id | Encode user text with special tokens disallowed |
| Cached dataset gives garbage | Ids built with another tokenizer version | Key caches by tokenizer name and file hash |
Trade-offs
Larger vocabularies shorten sequences, which cuts attention cost and price per character, but they spend parameters on embedding rows and leave rarer tokens undertrained; smaller vocabularies generalise across surface forms but make every sequence longer. Digit chunking of up to three keeps numbers compact but misaligns place value, while single-digit tokenization helps arithmetic and lengthens numeric text. Exposing raw ids in your application (logit bias, stop ids, cached tensors) is fast but couples you to one tokenizer; working in strings and re-encoding at the boundary costs a little CPU and survives model changes. For most application teams the right default is strings at every interface, ids only inside a component that owns the tokenizer.
What to do next
- Write down the tokenizer name and version for every model you call, and key token caches and datasets by it.
- Add a CI test that encodes each prompt template with a typical answer separately and jointly and fails if the ids differ.
- Rebuild SFT data by encoding full text once and masking by character offsets; reject examples whose boundary splits a token.
- Check that end-of-turn tokens are inside the loss and that user text is encoded with special tokens disallowed.
- Make streaming and span highlighting byte-aware and test them with Japanese, Chinese and emoji inputs.
- Route exact counting, spelling and arithmetic to tools, and evaluate on production-shaped strings.
- Recompute token budgets with the target tokenizer before any model switch, including template overhead.