Credits and estimates
You buy prepaid credits (1 credit = 1 US cent, USD) and jobs consume them. Rates are per model, not one blended list price: a smaller model that serves faster on cheaper hardware costs you less.
Estimates are free, and the credit amount shown at approval is a hard cap: an approved job never consumes more. Unused credits stay on your balance. Your job gets its own estimate based on current prices; the tables below are just representative examples.
Text output models
| model | input modality | input / 1M tokens | cached input / 1M | output / 1M tokens | price per image | price per clip |
|---|---|---|---|---|---|---|
| qwen3.8-27b | text | $0.18 | $0.045 | $0.90 | - | - |
| image + video | $0.30 | $0.075 | $1.20 | $0.0001 | $0.0004 | |
| gemma4-31b | text | $0.20 | $0.05 | $1.00 | - | - |
| image | $0.24 | $0.06 | $1.20 | $0.0002 | - | |
| gemma4-12b | text | $0.12 | $0.03 | $0.60 | - | - |
| image + audio + video | $0.30 | $0.075 | $1.50 | $0.0002 | $0.0002 | |
| qwen3-14b | text | $0.025 | $0.006 | $0.125 | - | - |
| nemotron-3-nano-omni-30b | text | $0.12 | $0.03 | $0.60 | - | - |
| image + audio + video | $0.30 | $0.075 | $1.50 | $0.0011 | $0.0002 |
Cached input is any part of a prompt the server has already processed in the same job - typically a shared instruction prefix repeated across rows. It bills at a quarter of the input rate, automatically: batch jobs with a common prompt prefix routinely see more than half their input tokens at the cached rate, which is why output tokens now carry more of the price than input tokens. There is nothing to configure and no way to lose money by it: a cache miss simply bills at the normal input rate.
The price per image is a ceiling on everything an image adds to your bill. Inside the model an image becomes input tokens through the encoder, and those tokens bill at the normal input rate, plus a small flat per-image charge for the fetch and encode work. The listed price per image covers the sum of both parts at the largest image the model's vision band accepts, so a run can come in under it but never over. The models differ here because their encoders differ: nemotron-3-nano-omni-30b tiles a large picture into as many as 13 tiles of 256 tokens, which is why it carries the highest ceiling. The two Gemma models encode every image to a fixed 280 tokens, and qwen3.8-27b caps an image at 200, which is why it carries the lowest ceiling.
Audio and video ride the token rates: a clip becomes input tokens through the model's encoder - on nemotron-3-nano-omni-30b 750 tokens per minute of audio, on qwen3.8-27b 105 tokens per sampled video frame at one frame per second, or at the higher rate a job sets with video_fps, which multiplies the frames charged. Each clip then adds a small flat charge, the price per clip in the table (the fetch and unpack work a clip forces whether it is three seconds or thirty, measured on production machines). One model also carries a per-minute surcharge for decode work that grows with the clip's length: qwen3.8-27b, at 0.03¢ per video minute, which puts a minute of video through it at about a third of a cent, answer included. The two models that hear, gemma4-12b and nemotron-3-nano-omni-30b, carry no per-minute surcharge at all: a clip costs the tokens its encoder makes of it plus the flat per-clip charge, which is why 40 minutes of speech through nemotron-3-nano-omni-30b comes to about two cents with the transcript in it. A job that carries any media is priced on its model's media rates, which sit above its text rates: media rows do less work per second of GPU time, so they cost us more per token. Every model that takes media shows a text row and a media row above.
Embedding output models
| model | embeds | text rows / 1M input tokens | media rows / 1M | per 1M images | per 1M video clips | realtime queries / 1M |
|---|---|---|---|---|---|---|
| qwen3-vl-embedding-8b | text + image + video | $0.03 | $0.08 | $4 | $60 | $0.121 |
| qwen3-vl-embedding-2b | text + image + video | $0.007 | $0.02 | $2 | $20 | $0.0077 |
| qwen3-embedding-8b | text | $0.03 | - | - | - | $0.033 |
Realtime embedding requests (see realtime embeddings) bill per input token at the realtime rate in the last column, from the same credit balance as batch jobs. Text queries only: image and video embedding stays on the batch API.
Embedding jobs bill input tokens only - no output-token or
cached-token charges. A row is priced by what it carries: text-only
rows at the text rate, rows with any media (images, video, or both)
at the media rate, plus the flat per-item charge from the table for
each image or clip - the fetch and decode work a media item forces
regardless of how few tokens it becomes. The two qwen3-vl-embedding
models put text, images, and video (up to 4 images or 1 video per
row) in one vector space, and every embedding model truncates to a
narrower vector via the per-request dimensions
parameter. See embedding jobs for
the job shape.
Decision output models
| model | answers | batch jobs / 1M input tokens | realtime queries / 1M | minimum job price |
|---|---|---|---|---|
| laya | choice, score, yes/no | $0.005 | $0.033 | $0.05 |
A decision model answers typed questions about a text instead of writing a reply: one answer per question, with a probability per option, in a single pass. It bills input tokens only, and it reads the text once per question, so a request's billed tokens are the text plus each question's own words, summed over the questions. There are no output tokens. Batch jobs bill at the first rate; realtime requests (see realtime decisions) bill at the second, from the same credit balance. See decision jobs for the request shape and the question types.
Minimum job price
Every job also has a minimum price: qwen3.8-27b $0.12, gemma4-31b $0.29, gemma4-12b $0.05, qwen3-14b $0.07, nemotron-3-nano-omni-30b $0.06, laya $0.05, and $0.02 for each embedding model - the at-cost price of starting the dedicated GPU machine a job runs on. It is a floor, not a fee: a job whose token price clears it pays nothing extra.
The flip side is that batch pricing rewards volume: once a job's token price clears the minimum, the machine start is not billed anymore, so the overhead share of what you pay falls toward zero as the job grows. Batching more rows into one larger job is the cheapest way to run them.
What jobs actually cost
| example job | model | assumed per row | estimate |
|---|---|---|---|
| Classify 1M support tickets | qwen3-14b | ≈300 tokens in, 10 out | ≈ $9 |
| Caption 100,000 photos | gemma4-31b | 1 image, ≈200 tokens in, 60 out | ≈ $32 |
| Transcribe 1,000 hours of audio | nemotron-3-nano-omni-30b | audio at 12.5 tokens/s, ≈12k transcript tokens out per hour | ≈ $32 |
| Describe 10,000 one-minute videos | nemotron-3-nano-omni-30b | video at 512 tokens/s, ≈500 tokens out | ≈ $101 |
Rows sharing a prompt prefix bill its repeat occurrences at the cached input rate, so cache-heavy jobs come in noticeably under these figures.