Base URL and authentication
Every endpoint below is relative to the batch API base URL
https://api.anex.sh/batch/v1 and needs an API key, created
on the console's Keys page:
Authorization: Bearer <your-api-key>
Keys belong to the organization, not to a person: any key of an organization can see and manage that organization's jobs, and each action records which key performed it. An id that belongs to another organization answers 404, the same as an id that does not exist.
An organization with EU data routing, which we turn on per organization
on request, keeps its data in the EEA (European Economic Area) and
uses a second base URL, https://api-eu.anex.sh/batch/v1,
with the same endpoints and the same keys. That API stores and reads
the data kept in the EEA; the address above answers a request for such
data with 421, naming the API to use. See
data residency.
The examples use $BASE for the base URL and
$TOKEN for the key.
The API is these eight endpoints, in the order a job meets them:
POST /inputs- Ingest an input file from a URL, or get a presigned target to upload one directly.
GET /inputs/{id}- Check an input's status until it is
ready. POST /jobs- Submit a job: an input plus a model.
GET /jobs/{id}- Poll a job for its status, estimate, and row counts.
POST /jobs/{id}/approve- Approve the estimate - required before any paid work starts.
POST /jobs/{id}/cancel- Cancel a job.
GET /jobs/{id}/result- Download the result file.
DELETE /jobs/{id}/data- Delete a job's data now rather than waiting for it to expire (a job still in flight is cancelled first).
There is no list-jobs and no list-models endpoint: keep the job ids you submit (the console's Jobs page shows every job of the organization), and see models below for the model ids.
The same API is described machine-readably by an OpenAPI 3.1
document at /openapi.json
(YAML): every operation has an id,
typed parameters and response schemas, so it drops straight into an
SDK generator or an agent's tool list. Every page of these docs also
answers Accept: text/markdown with a Markdown rendering
of itself.
Models
The model column below is the exact string to pass as
the model field when submitting a
job - copy it verbatim; there are no aliases or version suffixes.
The console's model dropdown selects from the same set.
Every model takes text; the input column lists what a model accepts on
top of that. A content part your chosen model has no encoder for is
not rejected at submit - the job fails while its full input is being
checked and priced (the estimating stage, when the job
has one), with failure_code: invalid_input - so pick a
model that actually covers what your rows carry.
The context column is the budget one row is served in, counted in tokens: at most that many input tokens including everything the media encodes to, at most that many output tokens, and on the cards these models run on today the input and the output together have to fit inside it as well. Media costs a fixed number of those tokens per model - a second of audio, a sampled video frame, an image - and the per-model rates and per-clip ceilings are in limits. A row that does not fit is rejected, and nothing is billed for it.
| model | input | output | context | quantization | exact weights served |
|---|---|---|---|---|---|
| qwen3.8-27b | text + image + video | text (reasoning) | 32,768 | fp8 | Qwen/Qwen3.8-27B-FP8 |
| gemma4-31b | text + image | text (reasoning) | 32,768 | fp8 | RedHatAI/gemma-4-31B-it-FP8-dynamic |
| gemma4-12b | text + image + audio + video | text (reasoning) | 32,768 | fp8 | RedHatAI/gemma-4-12B-it-FP8-Dynamic |
| qwen3-14b | text | text (reasoning) | 32,768 | fp8 | Qwen/Qwen3-14B-FP8 |
| nemotron-3-nano-omni-30b | text + image + audio + video | text (reasoning) | 32,768 | fp8 | nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8 |
| qwen3-embedding-8b | text | embedding vector | 32,768 | unquantized | Qwen/Qwen3-Embedding-8B |
| qwen3-vl-embedding-8b | text + image + video | embedding vector | 32,768 | unquantized | Qwen/Qwen3-VL-Embedding-8B |
| qwen3-vl-embedding-2b | text + image + video | embedding vector | 32,768 | unquantized | Qwen/Qwen3-VL-Embedding-2B |
| laya | text | typed decisions | 8,192 | unquantized | convaiinnovations/laya |
qwen3.8-27b takes a thinking effort level: low,
medium or xhigh (the default). Set it for a
job with reasoning_effort on submit, or for one row with
chat_template_kwargs.reasoning_effort.
The weights column names the published build each model is served from, so the version and precision are never a guess: an fp8 model runs that fp8 checkpoint, and an unquantized model runs the original published weights. Which models are actually on offer can vary by deployment - treat this table as the shape, not a guaranteed allowlist.
The context column is the general per-row budget. Rows carrying images
or video on qwen3.8-27b may take up to 131,072 input
tokens, so its long or densely sampled clips fit. Two models take a
job-level video sampling rate (video_fps):
qwen3.8-27b and nemotron-3-nano-omni-30b,
each up to 20 frames per second; gemma4-12b does not take
one yet.
The three embedding models return vectors instead of a text completion, and take a different job shape (see embedding jobs below):
- qwen3-embedding-8b
- 4096-dimension output, native; request any narrower width from 32
up to 4096 with
dimensions- 32,768-token context - qwen3-vl-embedding-8b
- text, image, video - one vector space across all three;
4096-dimension output, native; request any narrower width from 64
up to 4096 with
dimensions- 32,768-token context; per 1M input tokens: $0.02 text-only rows, $0.03 rows with images, $0.10 rows with video - qwen3-vl-embedding-2b
- budget sibling of the 8B; text, image,
video - one vector space across all three; 2048-dimension output,
native; request any narrower width from 64 up to 2048 with
dimensions- 32,768-token context; per 1M input tokens: $0.005 text-only rows, $0.01 rows with images, $0.05 rows with video
The two qwen3-vl-embedding-* models take image and video
content parts in input alongside text, all mapped into
the one vector space above - see embedding
jobs below for the parts contract. Like
qwen3-embedding-8b they bill input tokens only, with no
separate per-image or per-video charge - a media part becomes input
tokens through the model's own encoder, the same as text - and a
row's rate follows what it carries: text-only, images, or video.
laya is a decision model. It answers questions about a
text. A row gives it a text and a set of typed questions
(a choice between options, a score on a scale, a yes or no) and gets
one answer per question with a probability per option, in a single
pass and with no generated text. It reads text only, in more than 100
languages, and bills input tokens only. Its context is the longest
text one question is read over; the model picks one of two internal
checkpoints by the text's script and language, and the one it uses
for Latin-script text reads 512 tokens, so a longer text is cut to
that. See decision jobs for the row shape and
realtime decisions for the
single-request surface.
The input file
The input is a .jsonl file: one OpenAI-batch request
object per line. Each
line must be a JSON object with custom_id,
method, url and body, and
url must be /v1/chat/completions; a file that
targets any other endpoint is rejected.
{"custom_id":"row-1","method":"POST","url":"/v1/chat/completions","body":{"messages":[{"role":"user","content":"Say hello in French."}],"chat_template_kwargs":{"enable_thinking":false}}}
From a line's body, three things are used:
messages (multi-turn is preserved),
chat_template_kwargs (the per-row thinking setting and thinking effort,
chat_template_kwargs.reasoning_effort), and
response_format. The model and the output-token cap are
always the job-wide values - a model or
max_completion_tokens on the line does not override them -
and any other body field is not carried through.
A model whose modalities (see Models above) include image, audio, or video accepts those as extra content parts, the same shapes as the OpenAI API:
{"type":"image_url","image_url":{"url":"https://example.com/photo.jpg"}}
{"type":"input_audio","input_audio":{"data":"<base64>","format":"wav"}}
{"type":"audio_url","audio_url":{"url":"https://example.com/clip.mp3","duration_seconds":42}}
{"type":"video_url","video_url":{"url":"https://example.com/clip.mp4","duration_seconds":30}}
image_url, audio_url, and
video_url each accept a public URL or a data:
URI (use https:// for a remote host - it's what our SSRF
protections and pricing probe are built to assume);
input_audio carries its bytes as bare base64 (no
data: prefix) plus a format. Inline media is
size-capped per modality (see limits). A
type we don't recognize at all, and a real type your model
simply has no encoder for, both pass submit and fail the job once it
is checked (see errors).
Supported formats - audio: wav, mp3, m4a, aac, ogg,
opus, flac, aiff (wma is not supported). Video: mp4, webm, mov, mkv,
avi. Images: png, jpeg, webp, gif, bmp, tiff. These are container
formats: inside an allowlisted video container the codec still has to
be one the fleet decodes - H.264, H.265, VP8 and VP9 are the safe
choices, and an .mp4 carrying AV1 can still fail at
serving time. A part naming an unsupported container fails the job
while it is still free, with failure_code: invalid_input
naming the offending line.
We work out a part's format from, in order: the format
field (input_audio only), the media type of a
data: URI, the file extension on a URL, and the first
bytes of anything sent inline. Inline bytes that contradict the
declared format are rejected. For an audio or video URL with no file
extension we read the Content-Type from a
HEAD request made while the job is being estimated; if
your host answers, and that answer is
application/octet-stream or names something off the
allowlist, the job is rejected. That probe is best effort and covers
the first few dozen URLs in a job, so a clip we could not ask about is
admitted and can still fail at serving time. Three ways to fix
that: put a file extension on the URL, serve the object with its real
content type, or send the bytes inline. Images are never probed this
way, so an extensionless image URL with no data: media
type is passed through to the model as-is.
duration_seconds - audio and video are
priced and bounded by duration, which we otherwise have to work out
without downloading your media. Add an optional
duration_seconds inside the media object
(input_audio, audio_url, or
video_url) to state the real length yourself.
Undeclared, we work it out in the order below, and how we got there changes what you pay:
- A clip that states its own length - MP4, MOV,
M4A, WebM and MKV all do - is read exactly, whether you send it inline
as a
data:URI or host it. The clip is then priced and sampled at its real length. - Anything else we can size - an AVI, or a clip whose container we could not read - is estimated from its length in bytes at a conservative bitrate for the format. This is only an estimate: a clip encoded well below that bitrate reads short, and one encoded well above it reads long and can be refused for exceeding the model's clip ceiling.
- A clip we cannot size at all is assumed to run
the model's full duration ceiling, the worst case that still fits the
row's context: 40 minutes on
nemotron-3-nano-omni-30b, 30 s ongemma4-12b. Such a row is admitted, and it is checked and billed at that ceiling, so a 20-second clip sent that way tonemotron-3-nano-omni-30bis priced as 40 minutes of audio.
Declaring the length skips all of that and is the whole fix:
{"type":"audio_url","audio_url":{"url":"https://example.com/clip.mp3","duration_seconds":42}}
Declared, it's validated against the model's duration cap as stated,
and the per-minute surcharge (see
credits and billing) is billed on
exactly that length. Undeclared, an inline clip is estimated from the
byte size (a conservative floor, so the estimate only ever runs long)
and a hosted clip is taken at the model's ceiling. Admission always
resolves a duration offline, without waiting on your host. A
follow-up HEAD request only refines the byte-size
estimate for pricing on rows admission already accepted; it
never affects whether a row is admitted. Either way, declaring the
real duration gets you an accurate check and an accurate charge.
Embedding jobs
Point a job at one of the embedding models and the file carries a vector job's input instead of a chat prompt.
JSONL - url must be
/v1/embeddings, and body carries
input and an optional dimensions. On every
embedding model, input may be a single non-empty string -
one row is one embedding:
{"custom_id":"row-1","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":"a plain text string"}}
On qwen3-vl-embedding-8b and
qwen3-vl-embedding-2b, input can instead be
a list of chat-style content parts - text,
image_url, and video_url - so a row embeds
an image or a video clip, alone or alongside text, into the same
vector space as a plain text row:
{"custom_id":"row-2","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"image_url","image_url":{"url":"https://example.com/photo.jpg"}}]}}
{"custom_id":"row-3","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"video_url","video_url":{"url":"https://example.com/clip.mp4"}}],"dimensions":64}}
{"custom_id":"row-4","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"text","text":"caption for the image"},{"type":"image_url","image_url":{"url":"data:image/png;base64,iVBORw0..."}}]}}
image_url and video_url each accept a public
URL or a data: URI, the same shapes as
chat content parts. A part's modality must
be one the chosen model actually supports - a media part sent to
qwen3-embedding-8b (text only) is rejected naming the
modality, and no embedding model
accepts audio. A row's parts must include at least one with media
content or non-empty text. A list that isn't shaped as content parts -
not every element an object - is rejected with the same message as a
bare list of strings
(embeddings input must be a single string per row, not a list
(submit one row per input)); submit one row per string
instead. A messages key is rejected too
(an embeddings request carries 'input', not 'messages'),
and so is any encoding_format other than
"float". A file's rows must all target the same
endpoint - a job can't mix /v1/embeddings rows with
/v1/chat/completions rows.
Inline image and video parts are size-capped the same as inline chat
media, a video clip's duration is capped the same way too, and a clip
has to hold at least 2 frames (see
limits). Video on the embedding models is
always sampled at 1 frame per second; the job-level
video_fps is for generation models only. A row can carry at most 4 images
and 1 video clip on a model that takes them; a row over that isn't
rejected up front - the surplus is refused at serving time, the same
as chat media.
dimensions (optional)
truncates the output vector to a narrower width, Matryoshka style - an
integer from the model's own floor (32 for
qwen3-embedding-8b, 64 for
qwen3-vl-embedding-8b and
qwen3-vl-embedding-2b) up to the model's native width.
Set it job-level at submit to apply to every row, or per-row in a
JSONL line's body.dimensions; when both are set, the
row's own value wins for that row. Leave it out entirely for the
model's full native width.
For what reaches the model - whether any instruction or chat template is applied to your text, how the vector is pooled and normalized, and why two embeddings of the same text are not bit-identical - see how embedding inputs are processed.
Generation-only fields don't apply here: max_output_tokens,
reasoning, and response_format are all
rejected with 422 on an embedding model. The reverse also
holds - dimensions is rejected with 422 on a
chat-completions model.
Pricing and estimation work differently too: an embedding job bills
input tokens only, and its estimate comes from an exact token count
rather than a sample - see
embedding output and
poll a job below. qwen3-embedding-8b
bills $0.03 per 1M input tokens (0.003 cents per
1k). The two qwen3-vl-embedding-* models price a row by
what it carries: qwen3-vl-embedding-8b bills $0.03 per
1M input tokens on text-only rows and $0.08 on rows carrying any
media (images, video, or both);
qwen3-vl-embedding-2b bills $0.007 and $0.02 for the
same two. An image or video part becomes input tokens through the
model's own encoder, the same as text, and each item also adds a
small flat charge for the fetch and decode work it forces regardless
of how few tokens it becomes: on
qwen3-vl-embedding-8b $4 per million images and $60 per
million video clips, on qwen3-vl-embedding-2b $2 per
million images and $20 per million video clips.
A job's vectors are only comparable to another job's when both used
the same model, at the same dimensions width if either
job set one. Every embedding model produces its own vector space;
qwen3-vl-embedding-8b and qwen3-embedding-8b,
the two production-facing families, are not compatible with each
other.
How embedding inputs are processed
This section is for readers comparing our vectors against a local run of the same checkpoint, or deciding how to phrase their inputs. It applies to both embedding jobs and realtime embeddings.
We embed exactly the text you send. The string in
input reaches the model verbatim. Nothing is prepended,
appended, or rewritten: no instruction prefix, no task description,
and no system message. A batch row sends
{"model": ..., "input": "<your text>"} and a
realtime request sends
{"model": ..., "input": ["<your text>"]}, and that
is the whole transformation.
This is deliberate. Several embedding models are served under one OpenAI-compatible API, and their instruction conventions differ by model and by task, so the convention stays yours to choose. Instructions for embedding inputs below has each model's convention with a worked example of writing it into your own rows.
Text and media do not take the same path. A row whose
input is a string is embedded as written. A row whose
input is a list of content parts carrying an image or a
video has no single string to send, so it goes to the model as a chat
message, and a chat message is rendered through the model's own chat
template first. On qwen3-vl-embedding-8b and
qwen3-vl-embedding-2b a single-image row becomes:
<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|><|im_end|>\n
The system message there is the template's own default, not ours. The prompt ends after the user turn: no assistant header is appended, because the endpoint embeds the conversation rather than continuing it. The full template for every model is published, with its upstream source, on chat templates.
Both kinds of row land in the same vector space and retrieval across them works. But a text row and an image row of the same subject reach the model in different formats, and if you are searching images with text queries that difference is worth closing. See instructions for embedding inputs for the string to send.
Pooling and normalization. A vector is the model's last-token hidden state, L2 normalized to unit length, which is the pooling these checkpoints are built for. A string row has one end-of-text token appended before it is embedded, so that token is the position pooled; a media row's chat rendering does not get one, and pools its final template token instead. This is the one difference between the two paths you cannot close from your side, and it is a single token at the end of an otherwise identical prompt.
What differs from a local run of the checkpoint's example code is the prompt, not the weights: our vectors will not match a reference run element for element unless the reference is fed the same string. That is why phrasing your inputs the same way on both sides of a comparison matters more than matching us exactly.
dimensions truncates before normalizing, Matryoshka style,
so a narrowed vector comes back at unit length rather than as a raw
slice of the full-width one. Because every returned vector is unit
length, cosine similarity and dot product give the same ranking, and no
further normalization on your side is needed.
Weights. Every embedding model runs the unmodified upstream checkpoint at its published precision, with no quantization, distillation, or fine tuning of ours. The exact Hugging Face repository for each model is in the models table, and its chat template, byte for byte as that repository publishes it, is on chat templates.
Vectors are not bit-for-bit reproducible. Embedding
the same text twice can return numbers that differ in their last
digits: floating point results depend on the GPU, the batch the request
landed in, and the order of operations, so identical output is never
guaranteed across two machines or two runs. Cosine similarity between
two such vectors should still come out very close to
1.0. If you compare a repeated embedding of the same text,
at the same model and the same dimensions width, and get
meaningfully less than that, send us the examples through
support and we will look into it.
Instructions for embedding inputs
Embedding models are trained to be steered by a short instruction in front of the text: what the text is for, what you intend to match it against. Qwen reports that dropping the query-side instruction costs roughly 1 to 5 percent on their retrieval benchmarks, so this is worth doing if you are tuning for retrieval quality.
There is no instruction field to set. We add nothing to
your text, which means the instruction is simply the first part of the
string you send, and you get to choose its wording and its format. The
three recipes below are the ones worth copying.
1. A task instruction on qwen3-embedding-8b. This model's own convention is a two-line prefix on the query side, with documents left bare:
{"custom_id":"q1","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-embedding-8b","input":"Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: how do I cancel a running job"}}
{"custom_id":"d1","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-embedding-8b","input":"Cancelling a job stops new chunks from being dispatched. Work already running is drained and billed."}}
Replace the task sentence with your own: it describes the retrieval task, not the query. Keep it identical across every query in a corpus, because a query embedded under one instruction and a query embedded under another are not scored on the same footing.
2. A text query phrased like your image rows. On
qwen3-vl-embedding-8b and
qwen3-vl-embedding-2b, media rows are rendered through the
model's chat template and text rows are not. To put a text query in the
same format your images are in, write the template's own output into
the string yourself:
<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\nYOUR QUERY<|im_end|>\n
which in a JSONL row, next to the image row it is meant to match, is:
{"custom_id":"img1","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"image_url","image_url":{"url":"https://example.com/bike.jpg"}}]}}
{"custom_id":"q2","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":"<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\na red bicycle leaning on a fence<|im_end|>\n"}}
Those marker strings are the model's own special tokens and are recognized as single tokens, so the query row's prompt matches the image row's exactly, other than the one end-of-text token described under pooling. Send this format for every text row in the corpus or none of them.
3. Your own instruction, on both sides. A media row is
sent as a single user message, so its system turn always carries the
template's default and there is no way to replace it. Put your
instruction in the user turn instead, where a text row and a media row
can both carry it: on the media row as a leading text
part, on the text row as the same words in the same position.
{"custom_id":"img2","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"text","text":"Describe this image for searching purposes.\n"},{"type":"image_url","image_url":{"url":"https://example.com/bike.jpg"}}]}}
{"custom_id":"q3","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":"<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\nDescribe this image for searching purposes.\na red bicycle leaning on a fence<|im_end|>\n"}}
The trailing \n on the instruction is doing real work.
Nothing is inserted between your content parts, so an instruction that
does not end in a separator runs straight into the image placeholder
that follows it.
Instructions that suit this model are short and say what the vector is
for: Describe this image for searching purposes.,
Represent this product photo for retrieval by shoppers.,
Summarize this frame for matching against text captions.
Which one wins is an empirical question on your own data, and the cost
of trying three is three embedding runs.
You are billed for the instruction. It is part of the prompt, so it is counted as input tokens on every row that carries it, exactly like the rest of your text. A short instruction on a corpus of short rows is a real fraction of the job: 12 tokens of instruction on rows averaging 40 tokens is about a quarter of the bill.
Comparing across formats needs one convention.
Vectors are only comparable when both sides were phrased the same way,
at the same model, at the same dimensions width. Changing
an instruction is a re-embed of the whole corpus, not of the queries
alone, so settle the wording before you run the large job.
Decision jobs
Point a job at laya (see models)
and each row asks typed questions about a text instead of prompting
for a reply. The model reads the row's state - a string,
or a JSON object such as {"body": "..."} that it reads
as text - and answers every question in questions in one
pass, with a probability per option. No text is generated: there is
nothing to parse, there are no output tokens, and the answer to a
question is always one of the options you defined.
JSONL only - url must be
/v1/systemone, and body carries
state and questions, plus an optional
lang (below). One row is one text and the questions
asked of it:
{"custom_id":"ticket-1","method":"POST","url":"/v1/systemone","body":{"state":{"body":"We were billed twice for March. Please refund it today."},"questions":{"team":{"type":"choice","instructions":"Which team should handle this?","criteria":{"billing":"invoices, payments, refunds","technical":"bugs and outages","other":"everything else"}},"urgent":{"type":"noul","instructions":"Does the customer need an answer today?"},"frustration":{"type":"score","instructions":"How frustrated is the customer?","criteria":["calm","annoyed","angry"]}}}}
A file's rows must all target /v1/systemone, and a
body with any field beyond state,
questions, lang and model
fails the job naming the field. A row's own model is
never read: the job's model is what answers.
Question types. Each entry of questions
is named by you (the name comes back on the answer) and is an object
with a type, an instructions string
(the question itself, required), and criteria:
- choice
- One option out of several.
criteriais an object of option key to description. The answer carrieschoice, the key of the most probable option, andprobabilities, one per key. - noul
- Yes or no.
criteriais optional; when set, it is keyed exactlytrueandfalse(either or both), each with a description of what that answer means. The answer carriesnoul, the probability that the answer is yes, from 0 to 1. - score
- A level on an ordered scale.
criteriais a list of level descriptions, lowest first. The answer carriesscore, the expected level as a number from 0 to the index of the last level (1.3 on a three-level scale reads "between the second and third level, nearer the second"), andprobabilities, one per level, keyed by its index"0","1", ...
Every answer also carries its type,
confidence (0 to 1, how concentrated the distribution
is; on a yes-or-no question the larger of the two probabilities) and
answer_confidence (the probability of the reported
answer, comparable across the three types). The probabilities are
the model's own estimates and are not calibrated: they rank
options and flag doubt well, so gate on them with a threshold you
have checked on your own data instead of reading 0.9 as nine in ten.
A score answer adds a legend mapping each level index
back to its description; any further field on an answer is the
model's own and not part of this contract.
Limits, per row: at most 64 questions, a
state of at most 50,000 characters (an object counts as
its JSON rendering), and a budget of 192 tokens that a question's
option texts share: the more options a question has, the fewer
tokens of each the model reads, and a set that cannot fit is
refused. A row over the first two fails the job before approval
with failure_code: invalid_input naming the row; a
question the model itself refuses - options over the budget, a
yes-or-no question whose criteria carry a key other
than true or false, a missing
instructions - fails that row alone, with the model's
message in the row's error and
error_type: invalid_request_error, and every other row
is delivered. The text one question is read over is capped by the
model's context (see models): a longer text is
cut to it.
Languages. The model reads more than 100 languages
and picks one of its two internal checkpoints for each row by the
text's script and language; the answer's routing block
names the one it used. A row may carry lang, a language
code such as "de", to settle the choice for a text whose
script alone does not.
Writing questions that work. The model reads the
option keys and the criteria text, so make both descriptive:
{"billing": "invoices, payments, refunds"} is answered
far better than {"b": ""}. Keep the option set short. A
question with dozens of options loses accuracy and crowds the option
budget; split it into a coarse question and a finer one, or ask a
yes-or-no question per candidate. Say what a yes means: a
noul with criteria for both answers reads
better than one with the bare question. The model's publisher reports
ordered scales as its weakest question type (see the
model page); where a choice between named
levels would do, ask a choice.
Generation-only fields don't apply: max_output_tokens,
reasoning, reasoning_budget,
reasoning_effort, response_format and video_fps are rejected
with 422 on a decision model, and so is the embedding
models' dimensions.
Tokens and billing. A decision job bills input tokens
only, at $0.005 per 1M (0.0005 cents per 1k). The model reads the
text once per question, so a row's billed tokens are, for each
question, the text plus that question's instructions and option
texts, summed over the questions: a 300-token text asked ten short
questions bills a little over 3,000 tokens. The estimate comes from
an exact token count rather than a sample, so a decision job reaches
cost_ready in seconds, and its
est_output_tok is always 0. Each result row
reports the tokens it was billed in usage.input_tokens;
usage.output_tokens is always 0. The
minimum job price is $0.05 (see the
pricing page). See
decision output for the result
file.
POST /inputs - ingest an input from a URL
The API can copy your input file into storage from any HTTPS URL - a presigned GET URL from your own bucket is the simplest source. (The console uploads from the browser instead; this endpoint is API-only.)
curl -X POST "$BASE/inputs" -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{
"source_url": "https://your-bucket.s3.amazonaws.com/path/input.jsonl?X-Amz-...",
"format": "jsonl",
"callback_url": "https://you.example.com/hooks/ingest"
}'
{"id":"<ingest-id>","status":"pending","input_s3_uri":"s3://.../<ingest-id>.jsonl",
"callback_secret":"whsec_..."}
This endpoint requires credits: an organization whose balance is zero
or below is answered 402 here, before an upload target is
minted or a source URL is contacted, so nothing can be placed in
storage before anything is paid for.
format (jsonl) is optional.
callback_url is optional, and
callback_secret comes back only when you supply one - it is
returned on this response only, never again.
data_residency is optional: "world" or
"eu". Omit it and an organization with EU data routing
gets "eu", any other "world"; set
"world" to keep one input in the world region. Asking for
"eu" without EU data routing answers 403
(see turning it on).
"eu" copies the file into storage in the EEA, and only the EU API
(https://api-eu.anex.sh/batch/v1) takes that request: the main
base URL answers it with 421. A job submitted against the
input must have the same data_residency - see
POST /jobs.
Then poll until the copy has finished:
curl "$BASE/inputs/<ingest-id>" -H "Authorization: Bearer $TOKEN"
{"id":"<ingest-id>","status":"ready","input_s3_uri":"s3://...","error":null}
status goes pending →
copying → ready, or failed with
an error. With a callback_url registered you
get an input.ready or input.failed callback
instead of polling.
Keep the returned id (or input_s3_uri): input_id is preferred for job submission.
Both source_url and callback_url must
be https://, and a source that advertises a size over the
input cap is refused with 413 before any bytes are copied. A source that
does not advertise one, or is too slow to answer, is accepted here and
refused during the copy instead: the ingest then settles
as failed with the reason in error.
The copied object is kept for 7 days, the same window an uploaded one
gets, so submit the job well inside it. Once the window lapses the
input reports expired and a submit against it is refused -
create a new input instead. While ready, the response's
expires_at carries that deadline.
POST /inputs - or upload a file directly
If the file lives on your machine rather than somewhere we can fetch it
from, call the same endpoint with no source_url - just a
format:
curl -X POST "$BASE/inputs" -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"format": "jsonl"}'
{"id":"<id>","status":"awaiting_upload",
"upload_url":"https://...s3.amazonaws.com/",
"fields":{"key":"...","policy":"...",...},
"expires_at":"2026-07-26T15:04:05Z",
"input_s3_uri":"s3://.../uploads/.../<id>.jsonl"}
The response is a presigned upload target rather than a copy job. Upload
with a form POST: send every entry of fields as a form
field, then the file itself as the file field, last.
data_residency is optional here too: "world" or
"eu". Omit it and an organization with EU data routing
gets "eu", any other "world"; set
"world" to keep one upload in the world region. Asking for
"eu" without EU data routing answers 403
(see turning it on). With
"eu" the presigned target is storage in the EEA, and the
file goes straight there from your machine. Ask the EU API
(https://api-eu.anex.sh/batch/v1) for that target; the main base
URL answers with 421. A job submitted against the upload
must have the same data_residency - see
POST /jobs.
curl "$UPLOAD_URL" \
-F key=... -F policy=... -F x-amz-algorithm=... \
-F x-amz-credential=... -F x-amz-date=... -F x-amz-signature=... \
-F file=@data.jsonl
The upload must start before expires_at (about an hour
out); a lapsed window can't be refreshed - create a new input instead.
A file over the input cap is rejected by the upload itself, at the S3
level, rather than by this API. GET /inputs/{id} reports
awaiting_upload until the object lands, then
ready, and echoes expires_at while the
window is open - you can submit the job immediately after
uploading, since submission settles readiness itself. An uploaded
object is kept for 7 days, so submit well inside that window; a
ready input's expires_at carries that
deadline, and past it the input reports expired.
callback_url isn't available in this mode (there's no
server-side copy to report on) and is rejected with 400 if supplied.
POST /jobs - submit a job
curl -X POST "$BASE/jobs" -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{
"input_id": "<id>",
"model": "<model-id>"
}'
{"id":"<job-id>","status":"submitted"}
- input_id
- Required unless
input_s3_uriis given. Theidfrom an ingest or upload. Preferred: it lets this call settle an upload for you, so you can submit right after uploading with no separate poll. Submitting one that isn'treadyyet answers 409. - input_s3_uri
- Required unless
input_idis given. Theinput_s3_urifrom an ingest. Give exactly one ofinput_id/input_s3_uri. - model
- Required. A model id from the allowlist; the console's model picker lists the ids on offer. An unknown id is rejected before anything is created.
- max_output_tokens
- Optional, 1–32,768. Omit it and the cap is estimated from a sample
of your rows: the value every row is expected to stay under with
95% confidence, which is also what the job is priced on. Supply it
and your number is the cap and the price basis; on a reasoning
model the sample is still taken to set the thinking budget unless
you set
reasoning_budgetyourself or turn reasoning off. - max_tokens
- The former name of
max_output_tokens. Still accepted (and echoed back) so existing integrations keep working; new code should usemax_output_tokens, which says what it is: an output cap, thinking included, with the input not part of it. - reasoning
- Optional boolean, for the models marked "text (reasoning)" in the
models table above. It is the job-level
default; a JSONL row that sets
chat_template_kwargs.enable_thinkingitself keeps its own setting. Omit the field to leave the model's own default in place.trueon a model that produces no reasoning trace is refused (422 at submit, or a rejected row naming the line), rather than accepted and answered without one;falsestays valid on every generation model, so one body submits across a mixed model set. See thinking. - reasoning_budget
- Optional, 0–32,768, for the models marked "text (reasoning)"
in the models table above. Omit it and it is
estimated from the same sample as the output cap: the thinking
length every row is expected to stay under with 95% confidence, so
rows that think longer are cut off gracefully and still answer.
Set a number to cap thinking yourself (at least 64 below an
explicit
max_output_tokens), or0for no cap. A budget needs thinking on: a value of 1 or higher withreasoning: falseis refused (422), and a JSONL row that setschat_template_kwargs.enable_thinkingto false in a job with your budget fails the job, naming the line. An estimated budget is never applied to a row with thinking off. When the sample shows that even the answers alone would not fit under yourmax_output_tokens, estimation fails and says so; raisemax_output_tokensor setreasoning_budget.0is accepted on any generation model, reasoning or not, so the same submit body works across a mixed set; a value of 1 or higher is only for the models marked "text (reasoning)". See thinking. - reasoning_effort
- Optional, for the models that offer thinking effort levels (see
the note under the models table): how hard
the model thinks before it answers. qwen3.8-27b takes
low,mediumorxhigh; leave the field out and it thinks atxhigh, its default. A lower level usually spends fewer thinking tokens, and thinking tokens are billed as output, so it lowers cost and time per row. A JSONL row can set its own level withchat_template_kwargs.reasoning_effort, which wins for that row. Refused (422) on a model without levels, for a value that is not one of the model's levels (they are case-sensitive), and together withreasoning: false. It works together withreasoning_budget. See thinking. - video_fps
- Optional number above 0, for generation models that take a video
sampling rate: frames sampled per second of clip, applied to every
video row of the job. Omit it and each clip is sampled at the
model's default rate.
qwen3.8-27bdefaults to 1 and takes up to 20;nemotron-3-nano-omni-30bdefaults to 2 and takes up to 20;gemma4-12bdoes not take a rate yet. Each sampled frame is processed at the model's full per-frame resolution budget, and is charged as input tokens at the model's media input rate: a clip costs tokens per frame × (seconds × rate, rounded up, + 1), plus the per-clip and per-minute charges, which do not change with the rate. So a higher rate costs proportionally more per second of clip and shortens the longest clip a row can carry; the longest clip per model and rate is in limits. A rate above the model's ceiling, or any rate on a model that takes none, answers422. Embedding models sample video at a fixed 1 frame per second and do not take this field.
A rate is capped by the clip itself. Sampling cannot produce frames a clip does not contain, so a clip shot at 2 frames per second is sampled at 2 however high you set this, and is charged and length-checked at 2. Ask for more than your clips hold and you pay for what you get, not what you asked for. If you want a higher rate, send clips encoded at or above it. - response_format
- Optional. Applied to every row. See structured output.
- dimensions
- Embedding models only, optional. Truncates the output vector to a
narrower width (32 up to the model's native width). A JSONL row's
own
body.dimensionswins over this for that row.422on a non-embedding model. See embedding jobs. - data_residency
- Optional,
"world"or"eu"- where the job's payload is stored and processed. Omit it and an organization with EU data routing gets"eu", any other"world";"world"keeps one job in the world region."eu"without EU data routing answers403."eu"keeps the job's input, output and processing inside the EEA; submit such a job to the EU API (https://api-eu.anex.sh/batch/v1), since the main base URL answers it with421. The input must be stored in the same region as the job, or the submit is rejected with422- see uploading an input.
Format, size, and the shape of response_format are checked
before the job exists; a failure there returns an error and creates
nothing. Row count and the row-by-row checks (content parts, inline
media size, whether a row fits the model) run afterward: row counts
appear once the input is validated after submission, and invalid input
fails the job asynchronously with the same error reasons as before.
Submission has no idempotency key: every POST /jobs that
succeeds creates a new job. If a submit times out without a response,
check for the job (the console's Jobs page lists every job of the
organization) before retrying, or you may create it twice. Nothing
runs without approval either way, so a duplicate costs nothing until
someone approves it.
A credit balance of zero or below answers 402, before any
of your input is read and before a job is ever priced - a free grant or
a paid top-up, either counts; see
credits and billing. The same check
guards uploads, so an account with no balance
cannot place a file either.
This is a coarser check than approval's: it
only asks whether you have any balance at all, not whether it covers
this particular job.
An embedding model rejects max_output_tokens,
reasoning, and response_format with
422 - see embedding jobs for the
full field list. A decision model rejects the same fields, plus
video_fps and dimensions, and takes JSONL
input only - see decision jobs.
An organization can hold at most 10 jobs awaiting approval. Past that
cap, submit answers 429 with a Retry-After
header in seconds until some are approved, cancelled, or
expired; back off for that long, then resubmit.
The same status and header also meet a caller sending requests faster
than the API accepts. The two are told apart by the body: the cap's
message names the approval queue, the rate limit's
error.code is rate_limit_exceeded.
GET /jobs/{id} - poll a job
curl "$BASE/jobs/<job-id>" -H "Authorization: Bearer $TOKEN"
{"id":"<job-id>","status":"cost_ready",
"est_input_tok":120000,"est_output_tok":40000,"est_cost_cents":72,
"est_processing_seconds":900,"max_output_tokens":600,"reasoning_budget":null,
"success_rows":null,"error_rows":null,"format_error_rows":null,
"truncated_rows":null,"pending_rows":null,
"created_at":"...","approved_at":null,"data_expires_at":null,"partial":null,
"failure_code":null,"failure_reason":null,
"model":"qwen3.8-27b","reasoning":true,"task_type":"chat","format":"jsonl",
"response_format":null,"dimensions":null,"video_fps":null,"data_residency":"world",
"data_deleted_at":null}
status is the job's current state; the values and the order
they come in are described under
job lifecycle.
- est_input_tok, est_output_tok, est_cost_cents
- The estimate, filled in when the job reaches
cost_ready.est_cost_centsis the amount approval holds and the most the job can consume. A job that does any work costs at least 1 credit, so both this estimate and the final charge are floored there; a job that consumed nothing is not charged. Each model also has a minimum job price - the cost of the GPU machine start a small job forces, never added on top of a job whose token price clears it - so a large job pays pure per-token rates; see the pricing page for the current minimums. An embedding job'sest_output_tokis always0, and itsest_input_tokcomes from an exact token count rather than a sample, so it reachescost_readyin seconds - see embedding jobs. A decision job is estimated the same way, with the text counted once per question. - est_processing_seconds
- Expected wall-clock run time from approval to results, in whole seconds, frozen at estimate time: the fixed startup overhead recent jobs of this model have paid (booting a machine, pulling the weights) plus the job's own expected generation at the model's measured throughput. Advisory; most jobs are dominated by the startup overhead.
- max_output_tokens
- The effective per-row output cap: yours, or the estimated one
(
nulluntil estimated). An embedding job reports0: it has no completion budget. - reasoning_budget
- The effective thinking budget: yours, or the estimated one
(
nulluntil estimated). Alwaysnullon a non-reasoning model, which cannot set this field at all. - model
- The model the job ran against.
- reasoning
- The job-level thinking setting.
nullmeans you set no job-level default and the model's own behaviour applied - it does not mean thinking was off. A JSONL row that carried its ownchat_template_kwargs.enable_thinkingoverrides this for that row, which this job-level field cannot report. - reasoning_effort
- The job-level thinking effort level, or
nullwhen you set none and the model's default level applied. A JSONL row that carried its ownchat_template_kwargs.reasoning_effortoverrides this for that row, which this job-level field cannot report. - task_type
chat(a completion per row) orembed(a vector per row).- format
- The input file's format,
jsonl. - response_format
- The job-wide structured-output schema, if you set one. A JSONL row may carry its own, which wins for that row.
- dimensions
- The embedding width on an
embedjob;nullmeans the model's full width. - video_fps
- The video sampling rate you set at submit;
nullmeans the model's default rate applied. A clip whose own frame rate is lower is sampled, charged and length-checked at its own rate. - data_residency
- Where the job's payload is stored and processed:
"world"or"eu". Either API answers this poll; the result of an"eu"job is downloaded from the EU API. - success_rows, error_rows
- Rows that produced an answer, and rows that failed. Every row count
is
nulluntil the job has finished (done, failed or cancelled); a job in progress reports its status alone. - format_error_rows
- Rows whose constrained generation could not satisfy the
response_format. Counted separately, not inerror_rows. - truncated_rows
- Rows whose generation stopped on the output-token cap. Counted
separately, not in
error_rows. On a reasoning model,reasoning_budgetkeeps thinking from consuming the whole cap, so rows still answer instead of truncating. - pending_rows
- Rows not yet accounted for by any of the four counts above.
- partial
- True on a terminal job that still shipped the rows that finished.
- data_expires_at
- When the job's result is deleted: it is kept for 30 days after the
job finishes, then purged. The input file goes sooner, 7 days after
it was uploaded or ingested, so this timestamp tracks the result. Set once the
job reaches a terminal state (
done,failed,cancelled),nullbefore that. Download the result before this time; afterwards the result endpoint answers410 Gone. The job record itself (statuses, counts, cost) stays queryable. - data_deleted_at
- When the job's data was actually deleted, whether by the 30-day
purge or earlier through
DELETE /jobs/{id}/data;nullwhile the input and result still exist. - failure_code, failure_reason
- Set on a failed job.
failure_codeis a stable value (invalid_inputorinternal_error);failure_reasonis a readable message. Four media-specificinvalid_inputreasons: a content part whose type the model has no encoder for (names the model and modality), a clip whose duration exceeds the model's cap (worded differently for a declared vs. an estimated length - seeduration_seconds), an inline media item over its size cap (same wording as the 400 a bad row gets at submit, naming the line), and a row whose input tokens, media included, do not fit the model's context (see models). The last two media reasons read like this, naming the row:Row 'x': audio duration 45s exceeds the 30s maximum for model 'gemma4-12b'. Shorten or split this audio input.andRow 'x' is too long: 37,500 input tokens (including media) exceeds the 32,768-token maximum that model 'nemotron-3-nano-omni-30b' can serve with audio input. Shorten or split this row.A video clip longer than the model takes at the job's sampling rate reads like this:Row 'clip-7': video is 45 s; at the requested 20 fps the longest clip qwen3.8-27b takes is 38 s.Shorten the clip or lowervideo_fps. A job that fails this way is never billed. Jobs that go through estimation (nomax_output_tokens, or a reasoning model with noreasoning_budget) can also fail during estimation withinvalid_inputwhen the output does not fit the model: the rows need more output tokens than the model can generate, the model's reasoning runs too long on them, the typical answer hits the model's maximum output length and would be cut off, or (with an explicitmax_output_tokenson a reasoning model) the answers alone would not fit once the thinking budget and its cut-off notice are reserved. A job on a model we cannot price automatically fails during estimation too, usually withinvalid_inputand the reasonWe could not prepare a price for this job automatically.Each reason states the fix: set or raisemax_output_tokens, setreasoning_budget(0 for no cap), or disable reasoning.
POST /jobs/{id}/approve - approve a job
Approval is what starts paid GPU work. It holds
est_cost_cents against your prepaid balance; see
credits and billing.
curl -X POST "$BASE/jobs/<job-id>/approve" -H "Authorization: Bearer $TOKEN"
{"id":"<job-id>","status":"in_progress"}
A job must be in cost_ready: any other state answers 409.
A balance that does not cover the estimate answers 402 - top up and
approve again. With auto-approve turned on for the organization, this
call is made for you.
An estimate waits at most 24 hours: a job left in
cost_ready past that window is cancelled automatically,
free of charge. An organization can hold at most 10 jobs awaiting
approval; past that cap, submit answers 429
with a Retry-After header until some are approved,
cancelled, or expired.
POST /jobs/{id}/cancel - cancel a job
curl -X POST "$BASE/jobs/<job-id>/cancel" -H "Authorization: Bearer $TOKEN"
{"id":"<job-id>","status":"cancelling"}
A job with no work in flight stops immediately and answers
cancelled. A job that has already started answers
cancelling: no new work is started, rows already running
finish, and the job then settles as cancelled with the
completed rows downloadable as a partial result. Once a job is
finalizing or already terminal it is too late, and the call answers 409.
A job that is estimating answers 409 too: wait for
cost_ready and simply never approve it - estimates are
free, and an unapproved job cancels itself after 24 hours.
The exception is a job that keeps its data in the EEA. It can take up
to about an hour to estimate when no machine is ready to run its
model, so cancelling it while it is estimating stops it
immediately and answers cancelled.
GET /jobs/{id}/result - download the result
curl "$BASE/jobs/<job-id>/result" -H "Authorization: Bearer $TOKEN"
{"id":"<job-id>","status":"done","result_url":"https://...signed...",
"expires_in":3600,"partial":false,"row_count":1000}
result_url is a presigned link you GET to download the file;
expires_in is how many seconds it stays valid. Request the
endpoint again for a fresh link. row_count is the number of
rows actually in the file. What is inside it is described under
output format.
A result exists for a done job, and for a
failed or cancelled job that delivered the rows
it had finished - that one comes back with
"partial": true. A job with nothing to deliver answers 409.
A result is deleted 30 days after the job finishes. After that this endpoint answers 410 and the data is gone - download anything you want to keep well inside that window. You can also delete the data yourself as soon as the job has finished, see below.
DELETE /jobs/{id}/data - delete a job's data
curl -X DELETE "$BASE/jobs/<job-id>/data" -H "Authorization: Bearer $TOKEN"
{"id":"<job-id>","status":"done","data_deleted_at":"2026-07-24T10:30:00+00:00"}
Removes the job's input file and its result file (and every
intermediate object) right away, instead of at the 30-day mark. The
job record - statuses, row counts, cost, dates - stays, and
data_deleted_at on GET /jobs/{id}
records when. Download the result first if you still need it: after
this call the result endpoint answers 410 Gone, and
nothing can bring the data back.
A finished job - done, failed, or
cancelled - is purged at once and the call answers 200.
A job still in flight is first cancelled, exactly as
POST /jobs/{id}/cancel would (rows already
completed are delivered and billed), and its data is deleted
automatically as soon as the job has stopped; the call answers
202 Accepted with data_deleted_at null and a
message saying so. Poll GET /jobs/{id}
until data_deleted_at is set if you need to know when:
{"id":"<job-id>","status":"cancelling","data_deleted_at":null,
"message":"the job is being cancelled; its input and any result will be deleted once it stops"}
The call is idempotent: repeating it on a job whose data is already
gone answers 200 with the original data_deleted_at, and
202 again while the job is still stopping. The console offers the same
action as Delete data on every job.
Webhooks
An organization can register one endpoint (an owner does this on the console's Settings page). It then receives a signed POST for every job event, for every job in the organization - there is no per-event or per-job subscription.
- submitted
- The job was created.
- estimating
- Output length is being measured to price the job.
- cost_ready
- The estimate is ready. Carries
est_input_tok,est_output_tokandest_cost_cents. - in_progress
- The estimate was approved and the credits are held; the job is being run. No further event is sent until it finishes.
- completed
- The job finished. Carries
status,result_urlandresult_expires_in, so the event is directly actionable. - failed
- The job failed. Carries
status,partial, thefailure_code/failure_reasonwhen one was recorded, and the result URL when there is something to download. - cancelled
- The job was cancelled. Same payload as
failed.
A test event is also sent by the Settings page's test
button. Every event has the same envelope; the ones not listed with a
payload above carry only job_id.
{"type":"completed","timestamp":"2026-07-22T10:31:04.512+00:00",
"id":"<event-id>",
"data":{"job_id":"<job-id>","status":"done",
"result_url":"https://...signed...","result_expires_in":3600}}
Deliveries follow the Standard Webhooks convention. Each POST carries
webhook-id, webhook-timestamp (unix seconds)
and webhook-signature headers. To verify, base64-decode the
secret with its whsec_ prefix stripped, HMAC-SHA256 the
bytes {webhook-id}.{webhook-timestamp}. followed by the raw
request body, base64-encode the digest and compare it with the
v1,-prefixed value in webhook-signature. Verify
against the exact bytes received, not a re-serialized copy.
Only a 2xx response counts as delivered. Anything else is retried on a widening backoff for about a day, then given up. Deliveries are HTTPS only and redirects are never followed.
Ingest callbacks are separate: they go to the callback_url
of one ingest, are signed with that ingest's own
callback_secret, carry input.ready /
input.failed, and their envelope
timestamp is unix seconds rather than a date string.
Realtime embeddings
A separate, OpenAI-compatible surface for single-request embeddings - for a query vector at request time rather than a batch of rows. Point the OpenAI SDK at it directly:
from openai import OpenAI
client = OpenAI(base_url="https://api.anex.sh/realtime/v1", api_key="<your-api-key>")
The base URL is https://api.anex.sh/realtime/v1, and it
takes the same sk- API keys as the batch API above - no
separate key needed. Three endpoints: GET /models lists
the models on offer, POST /embeddings returns a vector
for one request, and POST /systemone answers typed
questions about a text (see
realtime decisions below):
curl https://api.anex.sh/realtime/v1/embeddings \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen3-embedding-8b", "input": "a search query", "dimensions": 1024}'
input is a string or a list of strings; model
is one of the realtime-served models below. Text only - realtime
requests carry no image or video content, and one that does is
rejected; embed media on the batch API instead (see
embedding jobs), and the resulting vectors
share the same space as a realtime text query on the same model at
the same dimensions.
A request may carry at most 128 inputs, 262,144 characters in any one
input, and 2,097,152 characters across all of them together; past any
of those the request is rejected with 400. Split a larger
set of texts across requests, or embed them on the batch API.
dimensions is optional and takes the same range as batch:
32 up to 4096 on qwen3-embedding-8b, 64 up to 4096 on
qwen3-vl-embedding-8b and 64 up to 2048 on
qwen3-vl-embedding-2b (text queries only in realtime -
image and video stay batch-only for the two VL models too).
encoding_format only accepts float.
A realtime query is preprocessed exactly like a batch text row: the string reaches the model verbatim, with no instruction and no chat template added, and the vector comes back at unit length. See how embedding inputs are processed for the details, including why a repeated query is not bit-identical.
Pricing is per input token at each model's realtime rate, listed in the realtime column of the embeddings pricing table, billed from the same credit balance as batch jobs.
Errors worth naming beyond the ones below: 402 when the
organization has no credits, 429 on a concurrency or
upstream rate limit (honor the Retry-After header),
501 when the model exists but isn't served in realtime,
and 503 when the realtime API is frozen or billing is
unavailable.
Realtime decisions
The same base URL and the same API keys serve typed decisions one
request at a time: POST /systemone takes a text and a
set of questions and answers them all in one pass, in tens of
milliseconds, with a probability per option and no generated text.
The body is the body of a
decision job row plus the
model, and the question types, limits and advice there
apply here unchanged:
curl https://api.anex.sh/realtime/v1/systemone \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "laya",
"state": {"body": "We were billed twice for March. Please refund it today."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "invoices, payments, refunds", "technical": "bugs and outages", "other": "everything else"}},
"urgent": {"type": "noul", "instructions": "Does the customer need an answer today?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["calm", "annoyed", "angry"]}}}'
The response carries one entry per question under
answers, keyed by the names you gave them (the figures
below only illustrate the shape):
{"model": "laya",
"answers": {
"team": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.91, "technical": 0.03, "other": 0.06},
"confidence": 0.84, "answer_confidence": 0.91},
"urgent": {"type": "noul", "noul": 0.88, "confidence": 0.88, "answer_confidence": 0.88},
"frustration": {"type": "score", "score": 1.3,
"legend": {"0": "calm", "1": "annoyed", "2": "angry"},
"probabilities": {"0": 0.12, "1": 0.46, "2": 0.42},
"confidence": 0.31, "answer_confidence": 0.46}},
"usage": {"input_tokens": 171, "output_tokens": 0},
"routing": {"model": "english", "reason": "..."}}
usage.input_tokens is what the request bills: the text
counted once per question plus each question's own words, the same
rule as a batch row. routing names the internal
checkpoint the request was answered by and why; it is informational.
GET /models lists this model with
"task": "decide", beside the embedding models'
"task": "embed".
A request may carry at most 64 questions and a text of at most
50,000 characters; past either it is rejected with 400.
A question the model refuses - a yes-or-no question whose
criteria carry a key other than true or
false, a missing instructions, an option set
that cannot fit its 192-token budget at all - comes back as
422 with the model's own message naming the question.
Pricing is per input token at the realtime rate in the decisions pricing table, billed from the same credit balance as batch jobs. A request that was refused bills nothing.
Errors worth naming beyond the ones below: 400 when the
model named is an embedding model (the message points at
POST /realtime/v1/embeddings; naming laya
there points back here), 402 when the organization has
no credits, 429 on the concurrency limit (honor the
Retry-After header), 502 when the decision
service failed to answer, and 503 when the realtime API
is frozen or billing is unavailable.
Errors
An error response carries a detail field describing it:
- 400
- The request or the input file failed validation - a non-HTTPS URL, a
file over the size cap, a file that is not JSONL, a malformed
response_format, a missingformaton a direct upload, or acallback_urlsupplied in upload mode.
A jsonl input's own row count and its rows' content are checked only once the job exists, not as part of this 400: those checks run duringestimating(or, for an explicitmax_output_tokens, while the job is chunked), and a violation fails the job withfailure_code: invalid_inputrather than answering the submit call itself. Two of those checks readunsupported content part type '<type>': supported types are ...- a content part whose type we don't recognize at all.an embedded <image|audio file|video file> exceeds the size limit (max N MB)- an inline media item over the cap in limits.
Row 'clip-7': video is 45 s; at the requested 20 fps the longest clip qwen3.8-27b takes is 38 s. - 401
- The API key is invalid or no longer active.
- 402
- The balance does not cover the estimate on approval.
- 403
- The request asks for
"data_residency": "eu", or uses the EU realtime address, for an organization without EU data routing:{"detail": "EU data routing is not enabled for this organization. Contact us to enable it."}. We turn it on per organization on request; see data residency. - 404
- Unknown id, or an id belonging to another organization.
- 409
- Wrong state - approving a job that is not awaiting approval,
cancelling one that is past the point of stopping, submitting an
input_idthat is notreadyyet, or asking for a result that does not exist yet. - 410
- The result data has been deleted - job payloads are kept for 30 days after completion.
- 413
- The ingest source advertises a size over the input cap. A direct upload over the cap is rejected by the upload itself instead.
- 421
- The data this request reads or stores is kept in the other region,
and this API does not serve it: creating or polling an input,
submitting a job, or downloading a result for data kept in the
EEA, sent to the main base URL (or the reverse). The body names
the region, and
apithe origin to send the same request to:{"detail": "this job is stored in the EU region; use https://api-eu.anex.sh", "residency": "eu", "api": "https://api-eu.anex.sh"}. - 422
- The request does not match the endpoint's shape - a missing
Authorizationheader, a missing or wrongly typed body field, or an unknown model id. Avideo_fpsthe model cannot take answers here too:video_fps 25 is above the 20 fps ceiling of qwen3.8-27b, or on a model that takes no ratemodel 'gemma4-12b' does not support video_fps. A decision job submitted with a generation-only field, answers here as well; on the realtime surface, a question the decision model refuses comes back as422with the model's own message. - 429
- Too many jobs awaiting approval in the organization (see
submit), too many concurrent
realtime requests, or requests
arriving faster than the API accepts. Wait the seconds in the
Retry-Afterheader, then retry. - 503
- A check the request depends on could not run - the key could not be validated, or the balance could not be read, so the approval was not made. Also answered while an operator has the API frozen for maintenance (downloading a result keeps working then). Retry.
5xx responses and timeouts are transient: retry with backoff. Approving and cancelling are guarded, so a retried approval never holds credits twice - but submitting is not (see submit a job): confirm a timed-out submit before repeating it.