CHAT TEMPLATES · THE STRING THE MODEL ACTUALLY SEES

Chat Template Renderer

Render the same conversation with the chat templates shipped by Qwen2.5, Llama 3.2, Gemma 2 and Phi-3 and see the exact text, special tokens and failures.

Direct tool · updates as you edit

Your experiment

Start with the Llama 3.2 template and a system, user, assistant, user conversation. Read the rendered prompt, then switch templates, try the conversation without a system message, two user messages in a row and a tool result, turn off the generation prompt and encode the text again.

Every input recomputes the result immediately; there is no animation because nothing here unfolds over time. An input outside its allowed range is rejected with a message and the previous valid result stays on screen.

Rendered prompt, with each newline shown as \n

Computed data

Metrics

Templates are ported from the chat_template field of each tokenizer_config.json (Qwen2.5-0.5B-Instruct and Phi-3-mini-4k-instruct from their publishers, Llama-3.2-1B-Instruct and gemma-2-2b-it from public mirrors), fetched on 2026-10-10, for conversations without tool definitions. Llama 3.2's template writes today's date when the renderer provides strftime_now, as recent transformers releases do; this page uses the template's fallback, 26 Jul 2024. The double-BOS check uses each tokenizer_config's add_bos_token value.

One conversation, four strings

A chat model never sees a list of messages; it sees one string built by the chat template stored in its tokenizer_config.json. The opening Llama 3.2 render starts with <|begin_of_text|>, writes a system header with Cutting Knowledge Date: December 2023 and a date before the system text, wraps every turn in <|start_header_id|> and <|eot_id|>, and ends with an open assistant header: 395 characters and 15 special tokens. Qwen2.5 renders the same messages in ChatML with <|im_start|> and <|im_end|>, 194 characters. Pick a Sample conversation or edit the box Messages, one per line as role: text, where the role is system, user, assistant or tool, and every template re-renders it at once.

Templates make choices

Without a system message Qwen2.5 inserts its own: You are Qwen, created by Alibaba Cloud. You are a helpful assistant. Llama 3.2 still writes the system header with only the date lines. Llama and Gemma trim spaces around each message, while Qwen2.5 and Phi-3 keep them exactly, so the spaces sample produces different strings. Phi-3 ends with <|endoftext|> when the generation prompt is off, and Gemma calls the assistant role model.

When a template refuses

Gemma 2's template raises System role not supported for any conversation that starts with a system message, which is why the opening conversation gives TEMPLATE ERROR there. It also raises Conversation roles must alternate when two user messages follow each other or a tool message sits where a user turn should be. Phi-3's template has no branch for a tool role, so the tool result is silently left out: DROPPED 1 MESSAGE. Qwen2.5 wraps it in <tool_response> inside a user turn and Llama 3.2 in an ipython header.

The double BOS trap

Llama 3.2 and Gemma 2 write their BOS token into the template and also set add_bos_token true. If the rendered text is tokenized again with special tokens enabled, which is what Encode the rendered text again with special tokens simulates, the tokenizer adds a second BOS: DUPLICATE BOS. The fix is to tokenize directly with apply_chat_template or to pass add_special_tokens=False. Qwen2.5 and Phi-3 set add_bos_token false, so encoding again does not duplicate anything. Add the generation prompt ends the string with an open assistant turn for the model to complete; turning it off gives the training format, which ends after the last turn.

Reading the tool and its limits

The lanes list the parsed messages and the template settings, the table shows each message with the text its template produced, and the text panel prints the whole prompt with every newline shown as \n. Templates were ported for conversations without tool definitions; when tools are passed, Qwen2.5 and Llama 3.2 add tool instructions to the prompt. Lines that do not start with a role are joined to the previous message.