What a chat template is
A model does not read your messages array. It reads one
flat string. The chat template is the Jinja program that turns the
first into the second: it decides where a system turn goes, what marks
the start and end of each turn, and - the part that matters most here -
which placeholder token stands in for an image, an audio clip or a
video.
The template ships inside the checkpoint, written by whoever published the weights. We do not pass a template of our own, and we do not patch the one we are given: every model is served with the file its own checkpoint publishes. This page publishes those files so you can read them, and links each one to its upstream source so you can check that claim rather than take it.
Three reasons to want the file:
- Reproducing a vector or a completion locally
- Running the same checkpoint yourself and getting a different answer usually means a different prompt string, not different weights. The template is where that difference lives.
- Understanding what you are billed for
- The scaffolding the template adds - the turn markers, an injected system message, the vision delimiters around each image - is part of the prompt, so it is counted as input tokens like everything else. A short row with an image is mostly scaffolding.
- Phrasing embedding inputs
- On the embedding models a text row and a media row do not go through the same path, and the template is the difference. See chat templates and embedding rows below.
Download a template
One file per model, byte for byte as the checkpoint publishes it. The source column links the exact upstream file, pinned to the commit the copy was taken from, so the two can be compared directly.
| model | template | taken from |
|---|---|---|
| qwen3.8-27b | qwen3.8-27b.jinja | chat_template.jinja |
| glm-5.2 | glm-5.2.jinja | chat_template.jinja |
| qwen3.6-27b | qwen3.6-27b.jinja | chat_template.jinja |
| gemma4-31b | gemma4-31b.jinja | chat_template.jinja |
| gemma4-12b | gemma4-12b.jinja | chat_template.jinja |
| qwen3-14b | qwen3-14b.jinja | tokenizer_config.json |
| nemotron-3-nano-omni-30b | nemotron-3-nano-omni-30b.jinja | chat_template.jinja |
| qwen3-embedding-8b | qwen3-embedding-8b.jinja | tokenizer_config.json |
| qwen3-vl-embedding-8b | qwen3-vl-embedding-8b.jinja | chat_template.jinja |
| qwen3-vl-embedding-2b | qwen3-vl-embedding-2b.jinja | chat_template.jinja |
index.json carries the same list in machine-readable form: repository, commit, source file name and byte length per model.
One model renders no template at all. The decision
model laya is an encoder that reads a row's text and its
questions as they are (see
decision jobs), so there is no
template to publish for it and it has no row in the table above.
Where the template lives differs by checkpoint. Newer
ones ship a standalone chat_template.jinja file, which is
copied here unchanged. Two do not:
qwen3-14b and qwen3-embedding-8b keep the
template as a JSON string under the chat_template key of
tokenizer_config.json. For those two the published
.jinja is that string with its JSON escaping
undone, which is the only edit made to any file on this page. The
source link in the table opens the JSON it came out of.
Checkpoints track their upstream repository. The commits above are what was served when this page was last built. A publisher who changes a template changes what the model reads, and we pick that change up. If you depend on byte-exact behaviour, keep the copy you tested against and diff it against this one before you assume a run will reproduce.
qwen3-vl-embedding-8b and
qwen3-vl-embedding-2b publish the same template: the two
files here are identical. Their vectors are still not comparable to
each other - a shared prompt format is not a shared vector space.
gemma4-31b and gemma4-12b publish the same
template too, down to the audio and video branches the 31B has no
encoder for.
How an image, audio clip or video reaches the model
Media never travels through the template. A content part carrying an image, an audio clip or a video is replaced by a short placeholder string, and the media itself goes to the model's own encoder, which turns it into the run of tokens that placeholder expands to. So the template tells you where your media sits in the prompt and what wraps it; the encoder decides how many tokens it costs.
These are the placeholders each model's template emits, one per media part, in the order your parts appear:
| model | image | audio | video |
|---|---|---|---|
| qwen3.8-27b | <|vision_start|><|image_pad|><|vision_end|> | not accepted | <|vision_start|><|video_pad|><|vision_end|> |
| qwen3.6-27b | <|vision_start|><|image_pad|><|vision_end|> | not accepted | not accepted |
| gemma4-31b | <|image|> | not accepted | not accepted |
| gemma4-12b | <|image|> | <|audio|> | <|video|> |
| nemotron-3-nano-omni-30b | <image> | <so_embedding> | <video> |
| qwen3-vl-embedding-8b | <|vision_start|><|image_pad|><|vision_end|> | not accepted | <|vision_start|><|video_pad|><|vision_end|> |
| qwen3-vl-embedding-2b | <|vision_start|><|image_pad|><|vision_end|> | not accepted | <|vision_start|><|video_pad|><|vision_end|> |
| glm-5.2, qwen3-14b, qwen3-embedding-8b | not accepted | not accepted | not accepted |
"Not accepted" means the model carries no encoder for that modality.
Such a part is not rejected at submit: the job fails while its input is
being checked and priced, with
failure_code: invalid_input. Pick a model whose input
column in the models table covers what
your rows carry.
The part types accepted for each modality, all of them the usual OpenAI-compatible spellings:
- image
image_url,image,input_image- audio
input_audio,audio_url- video
video_url,video
Nothing is inserted between your parts. A text part followed by an image part renders as the text immediately followed by the placeholder, with no space and no newline of their own. If you want a separator between your caption and your image, end the text part with it.
nemotron-3-nano-omni-30b reorders your
parts. Its template is the exception to the rule above: it
collects every media placeholder to the front of the message, one per
line, in the order images, then video, then audio, and puts all of
your text after them, whatever order you sent the parts in. Two or
more images are also numbered, rendering as
<image 1><image> <image 2><image>
rather than a bare placeholder each. A caption written to sit after
its picture ends up in front of it here, so write the prompt to read
correctly with the media first.
What a row renders to
A text-and-image row for qwen3.8-27b, as it appears in
your JSONL file:
{"custom_id": "1", "method": "POST", "url": "/v1/chat/completions", "body": {
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": [
{"type": "text", "text": "What is in this picture?\n"},
{"type": "image_url", "image_url": {"url": "https://example.com/bike.jpg"}}
]}]
}}
and the string its template produces. Newlines are shown as
\n, and the panel wraps for width - the real string is one
line:
<|im_start|>system\nReasoning effort is set to xhigh. Please think carefully
through the task, validate key assumptions, consider plausible alternatives,
and prioritize correctness, consistency, and clarity in the final
answer.<|im_end|>\n<|im_start|>user\nWhat is in this
picture?\n<|vision_start|><|image_pad|><|vision_end|><|im_end|>\n
Two things there are the template's doing rather than yours. It opens a system turn your row never asked for, carrying a reasoning-effort instruction. And your image part became a pad token wrapped in vision delimiters, with the encoder's output for the picture itself expanded in place of the pad. Both are counted as input tokens.
An audio row for gemma4-12b, which injects no system
turn. Its template also trims the whitespace off each text part, so a
trailing newline of yours does not survive into the prompt:
{"role": "user", "content": [
{"type": "text", "text": "Transcribe this clip."},
{"type": "input_audio", "input_audio": {"url": "https://example.com/call.mp3"}}
]}
<|turn>user\nTranscribe this clip.<|audio|><turn|>\n
What happens when your row carries a system message of its own is
decided by the template, and the models disagree about it.
qwen3-vl-embedding-8b and
qwen3-vl-embedding-2b use your text in place of their
default system turn. qwen3.8-27b keeps its reasoning-effort
instruction and appends your text after a blank line, so the prompt
carries both. Read the template for the model you are actually using
before assuming either.
Chat templates and embedding rows
The embedding models are the case where the template matters most, because only some of your rows go through it.
A row whose input is a plain string takes the completion
path: the string is embedded exactly as you wrote it, with no template,
no turn markers and no injected system message. A row whose
input is a list of content parts carrying an image or a
video has no single string to send, so it goes as a chat message and
does get rendered through the template above. Two rows of the same
model can therefore reach it in two different formats.
On qwen3-vl-embedding-8b and
qwen3-vl-embedding-2b that rendering is:
<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|><|im_end|>\n
The system message there is the template's own default, used because a row sends a single user message and has no system turn of its own. The prompt ends after the user turn: no assistant header is appended, because the pooling endpoint embeds the conversation rather than continuing it.
That gives you the string to match if you want a text query phrased identically to your image rows, which is worth doing when you search images with text. How embedding inputs are processed has the copy-paste recipe for it, and for putting your own instruction in front of either side.
qwen3-embedding-8b takes text only, so every row of its
goes the plain-string path and its template is never applied to
anything you send. It is published above for completeness, and because
it is the template Qwen's own example code applies before embedding -
which is what makes a local run's vectors differ from ours unless you
match one side to the other.