A chat request often starts as a list of messages. Each item has a role and content. The role names the speaker. The content holds the message. Common roles are system, user, and assistant. Before a causal language model generates a reply, a chat template turns that list into the sequence the model expects. Hugging Face’s chat template guide describes this conversion.
The template marks each turn
The formatted sequence contains message text and markers for its structure. A marker can identify a speaker or end a message. The model then continues that sequence of tokens. The visible message list and the model input have different shapes, even though they carry the same conversation. The guide shows message dictionaries and their formatted output side by side.
The markers depend on the model. In the guide, Mistral-7B-Instruct surrounds a user message with [INST] and [/INST]. Zephyr-7B uses markers such as <|user|> and <|assistant|>. These models share a base model, yet their chat formats differ. A template supplies the format each model expects. The guide warns that using the wrong control tokens can sharply hurt performance.
The ending sets up the reply
The final tokens matter. With add_generation_prompt=True, apply_chat_template can append the start of an assistant message. The model can then continue from that point as the assistant. Some models do not use a separate generation prompt, so this option has no effect for them. The examples show the added assistant prefix and explain the exception.
A different option, continue_final_message=True, removes the ending that would close the last message. It lets generation continue that message, for example when an assistant reply has already been started. The two options serve different continuations and cannot be used together.
What to do
- Use the chat template supplied with the model’s tokenizer. Pass messages as
roleandcontentdictionaries. - Inspect the formatted output with
apply_chat_template(..., tokenize=False). Check how it marks speakers and where it leaves the assistant turn. - When tokenizing with
apply_chat_template, usetokenize=True. If you format first and tokenize later, setadd_special_tokens=Falseto avoid duplicate special tokens. Hugging Face explains this tokenization detail.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.