Chat templates

The exact file each model renders your rows with, and what your media turns into inside it.

What a chat template is

A model does not read your messages array. It reads one flat string. The chat template is the Jinja program that turns the first into the second: it decides where a system turn goes, what marks the start and end of each turn, and - the part that matters most here - which placeholder token stands in for an image, an audio clip or a video.

The template ships inside the checkpoint, written by whoever published the weights. We do not pass a template of our own, and we do not patch the one we are given: every model is served with the file its own checkpoint publishes. This page publishes those files so you can read them, and links each one to its upstream source so you can check that claim rather than take it.

Three reasons to want the file:

Reproducing a vector or a completion locally
Running the same checkpoint yourself and getting a different answer usually means a different prompt string, not different weights. The template is where that difference lives.
Understanding what you are billed for
The scaffolding the template adds - the turn markers, an injected system message, the vision delimiters around each image - is part of the prompt, so it is counted as input tokens like everything else. A short row with an image is mostly scaffolding.
Phrasing embedding inputs
On the embedding models a text row and a media row do not go through the same path, and the template is the difference. See chat templates and embedding rows below.

Download a template

One file per model, byte for byte as the checkpoint publishes it. The source column links the exact upstream file, pinned to the commit the copy was taken from, so the two can be compared directly.

index.json carries the same list in machine-readable form: repository, commit, source file name and byte length per model.

One model renders no template at all. The decision model laya is an encoder that reads a row's text and its questions as they are (see decision jobs), so there is no template to publish for it and it has no row in the table above.

Where the template lives differs by checkpoint. Newer ones ship a standalone chat_template.jinja file, which is copied here unchanged. Two do not: qwen3-14b and qwen3-embedding-8b keep the template as a JSON string under the chat_template key of tokenizer_config.json. For those two the published .jinja is that string with its JSON escaping undone, which is the only edit made to any file on this page. The source link in the table opens the JSON it came out of.

Checkpoints track their upstream repository. The commits above are what was served when this page was last built. A publisher who changes a template changes what the model reads, and we pick that change up. If you depend on byte-exact behaviour, keep the copy you tested against and diff it against this one before you assume a run will reproduce.

qwen3-vl-embedding-8b and qwen3-vl-embedding-2b publish the same template: the two files here are identical. Their vectors are still not comparable to each other - a shared prompt format is not a shared vector space. gemma4-31b and gemma4-12b publish the same template too, down to the audio and video branches the 31B has no encoder for.

How an image, audio clip or video reaches the model

Media never travels through the template. A content part carrying an image, an audio clip or a video is replaced by a short placeholder string, and the media itself goes to the model's own encoder, which turns it into the run of tokens that placeholder expands to. So the template tells you where your media sits in the prompt and what wraps it; the encoder decides how many tokens it costs.

These are the placeholders each model's template emits, one per media part, in the order your parts appear:

modelimageaudiovideo
qwen3.8-27b<|vision_start|><|image_pad|><|vision_end|>not accepted<|vision_start|><|video_pad|><|vision_end|>
qwen3.6-27b<|vision_start|><|image_pad|><|vision_end|>not acceptednot accepted
gemma4-31b<|image|>not acceptednot accepted
gemma4-12b<|image|><|audio|><|video|>
nemotron-3-nano-omni-30b<image><so_embedding><video>
qwen3-vl-embedding-8b<|vision_start|><|image_pad|><|vision_end|>not accepted<|vision_start|><|video_pad|><|vision_end|>
qwen3-vl-embedding-2b<|vision_start|><|image_pad|><|vision_end|>not accepted<|vision_start|><|video_pad|><|vision_end|>
glm-5.2, qwen3-14b, qwen3-embedding-8bnot acceptednot acceptednot accepted

"Not accepted" means the model carries no encoder for that modality. Such a part is not rejected at submit: the job fails while its input is being checked and priced, with failure_code: invalid_input. Pick a model whose input column in the models table covers what your rows carry.

The part types accepted for each modality, all of them the usual OpenAI-compatible spellings:

image
image_url, image, input_image
audio
input_audio, audio_url
video
video_url, video

Nothing is inserted between your parts. A text part followed by an image part renders as the text immediately followed by the placeholder, with no space and no newline of their own. If you want a separator between your caption and your image, end the text part with it.

nemotron-3-nano-omni-30b reorders your parts. Its template is the exception to the rule above: it collects every media placeholder to the front of the message, one per line, in the order images, then video, then audio, and puts all of your text after them, whatever order you sent the parts in. Two or more images are also numbered, rendering as <image 1><image> <image 2><image> rather than a bare placeholder each. A caption written to sit after its picture ends up in front of it here, so write the prompt to read correctly with the media first.

What a row renders to

A text-and-image row for qwen3.8-27b, as it appears in your JSONL file:

{"custom_id": "1", "method": "POST", "url": "/v1/chat/completions", "body": {
  "model": "qwen3.8-27b",
  "messages": [{"role": "user", "content": [
    {"type": "text", "text": "What is in this picture?\n"},
    {"type": "image_url", "image_url": {"url": "https://example.com/bike.jpg"}}
  ]}]
}}

and the string its template produces. Newlines are shown as \n, and the panel wraps for width - the real string is one line:

<|im_start|>system\nReasoning effort is set to xhigh. Please think carefully
through the task, validate key assumptions, consider plausible alternatives,
and prioritize correctness, consistency, and clarity in the final
answer.<|im_end|>\n<|im_start|>user\nWhat is in this
picture?\n<|vision_start|><|image_pad|><|vision_end|><|im_end|>\n

Two things there are the template's doing rather than yours. It opens a system turn your row never asked for, carrying a reasoning-effort instruction. And your image part became a pad token wrapped in vision delimiters, with the encoder's output for the picture itself expanded in place of the pad. Both are counted as input tokens.

An audio row for gemma4-12b, which injects no system turn. Its template also trims the whitespace off each text part, so a trailing newline of yours does not survive into the prompt:

{"role": "user", "content": [
  {"type": "text", "text": "Transcribe this clip."},
  {"type": "input_audio", "input_audio": {"url": "https://example.com/call.mp3"}}
]}

<|turn>user\nTranscribe this clip.<|audio|><turn|>\n

What happens when your row carries a system message of its own is decided by the template, and the models disagree about it. qwen3-vl-embedding-8b and qwen3-vl-embedding-2b use your text in place of their default system turn. qwen3.8-27b keeps its reasoning-effort instruction and appends your text after a blank line, so the prompt carries both. Read the template for the model you are actually using before assuming either.

Chat templates and embedding rows

The embedding models are the case where the template matters most, because only some of your rows go through it.

A row whose input is a plain string takes the completion path: the string is embedded exactly as you wrote it, with no template, no turn markers and no injected system message. A row whose input is a list of content parts carrying an image or a video has no single string to send, so it goes as a chat message and does get rendered through the template above. Two rows of the same model can therefore reach it in two different formats.

On qwen3-vl-embedding-8b and qwen3-vl-embedding-2b that rendering is:

<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|><|im_end|>\n

The system message there is the template's own default, used because a row sends a single user message and has no system turn of its own. The prompt ends after the user turn: no assistant header is appended, because the pooling endpoint embeds the conversation rather than continuing it.

That gives you the string to match if you want a text query phrased identically to your image rows, which is worth doing when you search images with text. How embedding inputs are processed has the copy-paste recipe for it, and for putting your own instruction in front of either side.

qwen3-embedding-8b takes text only, so every row of its goes the plain-string path and its template is never applied to anything you send. It is published above for completeness, and because it is the template Qwen's own example code applies before embedding - which is what makes a local run's vectors differ from ours unless you match one side to the other.