The family
Nemotron is NVIDIA's own open-weight line, trained and published by the company whose cards the rest of this catalogue runs on. The models are released under the NVIDIA Open Model License, which permits commercial use, and NVIDIA publishes the checkpoint in several precisions rather than leaving the quantization to whoever serves it.
One Nemotron model is on offer here: the omni checkpoint, which reads audio, images and video alongside text and is the longest-form listener in our catalogue.
What we serve
| model | size | input | quantization | weights served |
|---|---|---|---|---|
| nemotron-3-nano-omni-30b | 30B MoE, 3B active | text + image + audio + video | fp8 | nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8 |
- Size and shape
- 30B total parameters with about 3B active per token, a MoE (Mixture of Experts) built on a hybrid of Mamba-2 state-space layers and transformer layers. Only the active parameters do work on each token, which is why it serves at a text rate near the cheapest here despite its size. It carries a vision encoder and a separate audio encoder, and answers in English.
- Audio, by the forty minutes
- Its audio encoder emits 12.5 tokens per second of clip, the lowest rate of any model here, and takes a clip up to 40 minutes long. The encoder itself would read an hour; the 32,768-token context a row is served in is what sets the ceiling, since 40 minutes of speech is already 30,000 of those tokens. That is still a whole meeting or a whole support call in one row, rather than a file cut into pieces and stitched back together afterwards. The rest of the numbers are under limits below.
- Video and documents
- Video up to 60 seconds, sampled at 2 frames per second and charged at 256 tokens a frame, which fills most of a row on its own. Images go through a tiling encoder: a small picture is one 256-token tile, and a large or long one is split into as many as 13, which is what lets it read a dense page scan at full resolution. That tiling is also why it carries the highest per-image ceiling on the pricing page.
- Output
-
Text, with reasoning on by default. Submit with
reasoning: falseto turn the trace off when you only want the answer, or per row withchat_template_kwargs.enable_thinking. A short prompt the model answers outright, with no thinking, still lands in the answer field (content, theoutputcolumn), never in the trace.
Limits
A row is served in a 32,768-token context: at most 32,768 input tokens with the media counted in, at most 32,768 output tokens, and on the cards this model runs on today the input and the output together have to fit inside 32,768 as well. Every second of media spends that budget at a fixed rate, which is what sets each ceiling below.
- Audio
- 40 minutes (2,400 seconds) per clip, at 12.5 tokens per second. A full clip is 30,000 tokens, which leaves about 2,700 for your prompt and a short answer - enough for a summary or a list of findings, not for a full transcript of the whole recording. The audio encoder would accept an hour; the served context will not, so a longer clip is rejected. Split a long recording, or ask for less output.
- Video
- 60 seconds per clip, sampled at 2 frames per second at 256 tokens a frame: a full minute is about 31,000 tokens, nearly the whole row.
- Images
- 256 tokens per 512 px tile, up to 13 tiles, so at most 3,328 tokens for one picture, and up to 2 images in a row.
- A row that does not fit
-
It is rejected, and nothing is billed for it. If it is one of the
rows sampled when the job is submitted, submit answers
400with the reason; otherwise the job fails before it can be approved, withfailure_code: invalid_inputand a reason naming the row and the number it went over. A hostedaudio_urlwith noduration_secondsdeclared is taken as a full 40 minutes, so it is admitted and priced as one - declare the real length. The per-model table is under limits in the docs.
Benchmark scores
Published figures, each from the source named beside it. Different benchmarks and different harnesses, so read down the column rather than across it.
| benchmark | score | source |
|---|---|---|
| AA Intelligence Index | 15 (estimated) | Artificial Analysis |
| VoiceBench, DailyOmni, WorldSense | best reported open-weight scores | model card |
Those two rows measure different things, and the second is the one to
buy it for. The Intelligence Index is a general-reasoning score: at 15
this model sits below gemma4-12b (22), the other model
here that hears. The voice and audio-video boards are where it leads,
which is the work a job sends 40 minutes of audio for:
hearing what is in a recording and reporting it accurately. Ask it to
do hard reasoning about what it heard and the index is the number that
predicts the result.
When to pick it
- Long recordings
- Calls, meetings, interviews, podcasts. Nothing else here takes a clip of that length in one row, and at 12.5 tokens per second 40 minutes of speech is 30,000 tokens of input, which comes to about a cent on this model's rates.
- Dense page scans and documents
- The tiling image encoder reads a full page rather than a downscaled version of it. Pay attention to the per-image ceiling when you estimate: it is the highest on the pricing page, and a 13-tile page carries 3,328 input tokens where a fixed-cost encoder charges 280.
- Cheap text at volume
-
Its text rate is among the cheapest here, behind
qwen3-14band level withgemma4-12b, because only 3B of its 30B parameters work on each token. - Not for
- Languages other than English, and reasoning-heavy work where the index gap matters. For those, send the job to Gemma or Qwen, or to GLM for the hardest reasoning.
Pass nemotron-3-nano-omni-30b exactly as written when you
submit a job. Its rates and its minimum
job price are on the pricing page.