anex.sh batch API reference

Every endpoint of the batch API, with its request and response shape.

Base URL and authentication

Every endpoint below is relative to the batch API base URL https://api.anex.sh/batch/v1 and needs an API key, created on the console's Keys page:

Authorization: Bearer <your-api-key>

Keys belong to the organization, not to a person: any key of an organization can see and manage that organization's jobs, and each action records which key performed it. An id that belongs to another organization answers 404, the same as an id that does not exist.

An organization with EU data routing, which we turn on per organization on request, keeps its data in the EEA (European Economic Area) and uses a second base URL, https://api-eu.anex.sh/batch/v1, with the same endpoints and the same keys. That API stores and reads the data kept in the EEA; the address above answers a request for such data with 421, naming the API to use. See data residency.

The examples use $BASE for the base URL and $TOKEN for the key.

The API is these eight endpoints, in the order a job meets them:

POST /inputs
Ingest an input file from a URL, or get a presigned target to upload one directly.
GET /inputs/{id}
Check an input's status until it is ready.
POST /jobs
Submit a job: an input plus a model.
GET /jobs/{id}
Poll a job for its status, estimate, and row counts.
POST /jobs/{id}/approve
Approve the estimate - required before any paid work starts.
POST /jobs/{id}/cancel
Cancel a job.
GET /jobs/{id}/result
Download the result file.
DELETE /jobs/{id}/data
Delete a job's data now rather than waiting for it to expire (a job still in flight is cancelled first).

There is no list-jobs and no list-models endpoint: keep the job ids you submit (the console's Jobs page shows every job of the organization), and see models below for the model ids.

The same API is described machine-readably by an OpenAPI 3.1 document at /openapi.json (YAML): every operation has an id, typed parameters and response schemas, so it drops straight into an SDK generator or an agent's tool list. Every page of these docs also answers Accept: text/markdown with a Markdown rendering of itself.

Models

The model column below is the exact string to pass as the model field when submitting a job - copy it verbatim; there are no aliases or version suffixes. The console's model dropdown selects from the same set.

Every model takes text; the input column lists what a model accepts on top of that. A content part your chosen model has no encoder for is not rejected at submit - the job fails while its full input is being checked and priced (the estimating stage, when the job has one), with failure_code: invalid_input - so pick a model that actually covers what your rows carry.

The context column is the budget one row is served in, counted in tokens: at most that many input tokens including everything the media encodes to, at most that many output tokens, and on the cards these models run on today the input and the output together have to fit inside it as well. Media costs a fixed number of those tokens per model - a second of audio, a sampled video frame, an image - and the per-model rates and per-clip ceilings are in limits. A row that does not fit is rejected, and nothing is billed for it.

modelinputoutputcontextquantizationexact weights served
qwen3.8-27btext + image + videotext (reasoning)32,768fp8Qwen/Qwen3.8-27B-FP8
gemma4-31btext + imagetext (reasoning)32,768fp8RedHatAI/gemma-4-31B-it-FP8-dynamic
gemma4-12btext + image + audio + videotext (reasoning)32,768fp8RedHatAI/gemma-4-12B-it-FP8-Dynamic
qwen3-14btexttext (reasoning)32,768fp8Qwen/Qwen3-14B-FP8
nemotron-3-nano-omni-30btext + image + audio + videotext (reasoning)32,768fp8nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8
qwen3-embedding-8btextembedding vector32,768unquantizedQwen/Qwen3-Embedding-8B
qwen3-vl-embedding-8btext + image + videoembedding vector32,768unquantizedQwen/Qwen3-VL-Embedding-8B
qwen3-vl-embedding-2btext + image + videoembedding vector32,768unquantizedQwen/Qwen3-VL-Embedding-2B
layatexttyped decisions8,192unquantizedconvaiinnovations/laya

qwen3.8-27b takes a thinking effort level: low, medium or xhigh (the default). Set it for a job with reasoning_effort on submit, or for one row with chat_template_kwargs.reasoning_effort.

The weights column names the published build each model is served from, so the version and precision are never a guess: an fp8 model runs that fp8 checkpoint, and an unquantized model runs the original published weights. Which models are actually on offer can vary by deployment - treat this table as the shape, not a guaranteed allowlist.

The context column is the general per-row budget. Rows carrying images or video on qwen3.8-27b may take up to 131,072 input tokens, so its long or densely sampled clips fit. Two models take a job-level video sampling rate (video_fps): qwen3.8-27b and nemotron-3-nano-omni-30b, each up to 20 frames per second; gemma4-12b does not take one yet.

The three embedding models return vectors instead of a text completion, and take a different job shape (see embedding jobs below):

qwen3-embedding-8b
4096-dimension output, native; request any narrower width from 32 up to 4096 with dimensions - 32,768-token context
qwen3-vl-embedding-8b
text, image, video - one vector space across all three; 4096-dimension output, native; request any narrower width from 64 up to 4096 with dimensions - 32,768-token context; per 1M input tokens: $0.02 text-only rows, $0.03 rows with images, $0.10 rows with video
qwen3-vl-embedding-2b
budget sibling of the 8B; text, image, video - one vector space across all three; 2048-dimension output, native; request any narrower width from 64 up to 2048 with dimensions - 32,768-token context; per 1M input tokens: $0.005 text-only rows, $0.01 rows with images, $0.05 rows with video

The two qwen3-vl-embedding-* models take image and video content parts in input alongside text, all mapped into the one vector space above - see embedding jobs below for the parts contract. Like qwen3-embedding-8b they bill input tokens only, with no separate per-image or per-video charge - a media part becomes input tokens through the model's own encoder, the same as text - and a row's rate follows what it carries: text-only, images, or video.

laya is a decision model. It answers questions about a text. A row gives it a text and a set of typed questions (a choice between options, a score on a scale, a yes or no) and gets one answer per question with a probability per option, in a single pass and with no generated text. It reads text only, in more than 100 languages, and bills input tokens only. Its context is the longest text one question is read over; the model picks one of two internal checkpoints by the text's script and language, and the one it uses for Latin-script text reads 512 tokens, so a longer text is cut to that. See decision jobs for the row shape and realtime decisions for the single-request surface.

The input file

The input is a .jsonl file: one OpenAI-batch request object per line. Each line must be a JSON object with custom_id, method, url and body, and url must be /v1/chat/completions; a file that targets any other endpoint is rejected.

{"custom_id":"row-1","method":"POST","url":"/v1/chat/completions","body":{"messages":[{"role":"user","content":"Say hello in French."}],"chat_template_kwargs":{"enable_thinking":false}}}

From a line's body, three things are used: messages (multi-turn is preserved), chat_template_kwargs (the per-row thinking setting and thinking effort, chat_template_kwargs.reasoning_effort), and response_format. The model and the output-token cap are always the job-wide values - a model or max_completion_tokens on the line does not override them - and any other body field is not carried through.

A model whose modalities (see Models above) include image, audio, or video accepts those as extra content parts, the same shapes as the OpenAI API:

{"type":"image_url","image_url":{"url":"https://example.com/photo.jpg"}}
{"type":"input_audio","input_audio":{"data":"<base64>","format":"wav"}}
{"type":"audio_url","audio_url":{"url":"https://example.com/clip.mp3","duration_seconds":42}}
{"type":"video_url","video_url":{"url":"https://example.com/clip.mp4","duration_seconds":30}}

image_url, audio_url, and video_url each accept a public URL or a data: URI (use https:// for a remote host - it's what our SSRF protections and pricing probe are built to assume); input_audio carries its bytes as bare base64 (no data: prefix) plus a format. Inline media is size-capped per modality (see limits). A type we don't recognize at all, and a real type your model simply has no encoder for, both pass submit and fail the job once it is checked (see errors).

Supported formats - audio: wav, mp3, m4a, aac, ogg, opus, flac, aiff (wma is not supported). Video: mp4, webm, mov, mkv, avi. Images: png, jpeg, webp, gif, bmp, tiff. These are container formats: inside an allowlisted video container the codec still has to be one the fleet decodes - H.264, H.265, VP8 and VP9 are the safe choices, and an .mp4 carrying AV1 can still fail at serving time. A part naming an unsupported container fails the job while it is still free, with failure_code: invalid_input naming the offending line.

We work out a part's format from, in order: the format field (input_audio only), the media type of a data: URI, the file extension on a URL, and the first bytes of anything sent inline. Inline bytes that contradict the declared format are rejected. For an audio or video URL with no file extension we read the Content-Type from a HEAD request made while the job is being estimated; if your host answers, and that answer is application/octet-stream or names something off the allowlist, the job is rejected. That probe is best effort and covers the first few dozen URLs in a job, so a clip we could not ask about is admitted and can still fail at serving time. Three ways to fix that: put a file extension on the URL, serve the object with its real content type, or send the bytes inline. Images are never probed this way, so an extensionless image URL with no data: media type is passed through to the model as-is.

duration_seconds - audio and video are priced and bounded by duration, which we otherwise have to work out without downloading your media. Add an optional duration_seconds inside the media object (input_audio, audio_url, or video_url) to state the real length yourself.

Undeclared, we work it out in the order below, and how we got there changes what you pay:

  • A clip that states its own length - MP4, MOV, M4A, WebM and MKV all do - is read exactly, whether you send it inline as a data: URI or host it. The clip is then priced and sampled at its real length.
  • Anything else we can size - an AVI, or a clip whose container we could not read - is estimated from its length in bytes at a conservative bitrate for the format. This is only an estimate: a clip encoded well below that bitrate reads short, and one encoded well above it reads long and can be refused for exceeding the model's clip ceiling.
  • A clip we cannot size at all is assumed to run the model's full duration ceiling, the worst case that still fits the row's context: 40 minutes on nemotron-3-nano-omni-30b, 30 s on gemma4-12b. Such a row is admitted, and it is checked and billed at that ceiling, so a 20-second clip sent that way to nemotron-3-nano-omni-30b is priced as 40 minutes of audio.

Declaring the length skips all of that and is the whole fix:

{"type":"audio_url","audio_url":{"url":"https://example.com/clip.mp3","duration_seconds":42}}

Declared, it's validated against the model's duration cap as stated, and the per-minute surcharge (see credits and billing) is billed on exactly that length. Undeclared, an inline clip is estimated from the byte size (a conservative floor, so the estimate only ever runs long) and a hosted clip is taken at the model's ceiling. Admission always resolves a duration offline, without waiting on your host. A follow-up HEAD request only refines the byte-size estimate for pricing on rows admission already accepted; it never affects whether a row is admitted. Either way, declaring the real duration gets you an accurate check and an accurate charge.

Embedding jobs

Point a job at one of the embedding models and the file carries a vector job's input instead of a chat prompt.

JSONL - url must be /v1/embeddings, and body carries input and an optional dimensions. On every embedding model, input may be a single non-empty string - one row is one embedding:

{"custom_id":"row-1","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":"a plain text string"}}

On qwen3-vl-embedding-8b and qwen3-vl-embedding-2b, input can instead be a list of chat-style content parts - text, image_url, and video_url - so a row embeds an image or a video clip, alone or alongside text, into the same vector space as a plain text row:

{"custom_id":"row-2","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"image_url","image_url":{"url":"https://example.com/photo.jpg"}}]}}
{"custom_id":"row-3","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"video_url","video_url":{"url":"https://example.com/clip.mp4"}}],"dimensions":64}}
{"custom_id":"row-4","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"text","text":"caption for the image"},{"type":"image_url","image_url":{"url":"data:image/png;base64,iVBORw0..."}}]}}

image_url and video_url each accept a public URL or a data: URI, the same shapes as chat content parts. A part's modality must be one the chosen model actually supports - a media part sent to qwen3-embedding-8b (text only) is rejected naming the modality, and no embedding model accepts audio. A row's parts must include at least one with media content or non-empty text. A list that isn't shaped as content parts - not every element an object - is rejected with the same message as a bare list of strings (embeddings input must be a single string per row, not a list (submit one row per input)); submit one row per string instead. A messages key is rejected too (an embeddings request carries 'input', not 'messages'), and so is any encoding_format other than "float". A file's rows must all target the same endpoint - a job can't mix /v1/embeddings rows with /v1/chat/completions rows.

Inline image and video parts are size-capped the same as inline chat media, a video clip's duration is capped the same way too, and a clip has to hold at least 2 frames (see limits). Video on the embedding models is always sampled at 1 frame per second; the job-level video_fps is for generation models only. A row can carry at most 4 images and 1 video clip on a model that takes them; a row over that isn't rejected up front - the surplus is refused at serving time, the same as chat media.

dimensions (optional) truncates the output vector to a narrower width, Matryoshka style - an integer from the model's own floor (32 for qwen3-embedding-8b, 64 for qwen3-vl-embedding-8b and qwen3-vl-embedding-2b) up to the model's native width. Set it job-level at submit to apply to every row, or per-row in a JSONL line's body.dimensions; when both are set, the row's own value wins for that row. Leave it out entirely for the model's full native width.

For what reaches the model - whether any instruction or chat template is applied to your text, how the vector is pooled and normalized, and why two embeddings of the same text are not bit-identical - see how embedding inputs are processed.

Generation-only fields don't apply here: max_output_tokens, reasoning, and response_format are all rejected with 422 on an embedding model. The reverse also holds - dimensions is rejected with 422 on a chat-completions model.

Pricing and estimation work differently too: an embedding job bills input tokens only, and its estimate comes from an exact token count rather than a sample - see embedding output and poll a job below. qwen3-embedding-8b bills $0.03 per 1M input tokens (0.003 cents per 1k). The two qwen3-vl-embedding-* models price a row by what it carries: qwen3-vl-embedding-8b bills $0.03 per 1M input tokens on text-only rows and $0.08 on rows carrying any media (images, video, or both); qwen3-vl-embedding-2b bills $0.007 and $0.02 for the same two. An image or video part becomes input tokens through the model's own encoder, the same as text, and each item also adds a small flat charge for the fetch and decode work it forces regardless of how few tokens it becomes: on qwen3-vl-embedding-8b $4 per million images and $60 per million video clips, on qwen3-vl-embedding-2b $2 per million images and $20 per million video clips.

A job's vectors are only comparable to another job's when both used the same model, at the same dimensions width if either job set one. Every embedding model produces its own vector space; qwen3-vl-embedding-8b and qwen3-embedding-8b, the two production-facing families, are not compatible with each other.

How embedding inputs are processed

This section is for readers comparing our vectors against a local run of the same checkpoint, or deciding how to phrase their inputs. It applies to both embedding jobs and realtime embeddings.

We embed exactly the text you send. The string in input reaches the model verbatim. Nothing is prepended, appended, or rewritten: no instruction prefix, no task description, and no system message. A batch row sends {"model": ..., "input": "<your text>"} and a realtime request sends {"model": ..., "input": ["<your text>"]}, and that is the whole transformation.

This is deliberate. Several embedding models are served under one OpenAI-compatible API, and their instruction conventions differ by model and by task, so the convention stays yours to choose. Instructions for embedding inputs below has each model's convention with a worked example of writing it into your own rows.

Text and media do not take the same path. A row whose input is a string is embedded as written. A row whose input is a list of content parts carrying an image or a video has no single string to send, so it goes to the model as a chat message, and a chat message is rendered through the model's own chat template first. On qwen3-vl-embedding-8b and qwen3-vl-embedding-2b a single-image row becomes:

<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|><|im_end|>\n

The system message there is the template's own default, not ours. The prompt ends after the user turn: no assistant header is appended, because the endpoint embeds the conversation rather than continuing it. The full template for every model is published, with its upstream source, on chat templates.

Both kinds of row land in the same vector space and retrieval across them works. But a text row and an image row of the same subject reach the model in different formats, and if you are searching images with text queries that difference is worth closing. See instructions for embedding inputs for the string to send.

Pooling and normalization. A vector is the model's last-token hidden state, L2 normalized to unit length, which is the pooling these checkpoints are built for. A string row has one end-of-text token appended before it is embedded, so that token is the position pooled; a media row's chat rendering does not get one, and pools its final template token instead. This is the one difference between the two paths you cannot close from your side, and it is a single token at the end of an otherwise identical prompt.

What differs from a local run of the checkpoint's example code is the prompt, not the weights: our vectors will not match a reference run element for element unless the reference is fed the same string. That is why phrasing your inputs the same way on both sides of a comparison matters more than matching us exactly.

dimensions truncates before normalizing, Matryoshka style, so a narrowed vector comes back at unit length rather than as a raw slice of the full-width one. Because every returned vector is unit length, cosine similarity and dot product give the same ranking, and no further normalization on your side is needed.

Weights. Every embedding model runs the unmodified upstream checkpoint at its published precision, with no quantization, distillation, or fine tuning of ours. The exact Hugging Face repository for each model is in the models table, and its chat template, byte for byte as that repository publishes it, is on chat templates.

Vectors are not bit-for-bit reproducible. Embedding the same text twice can return numbers that differ in their last digits: floating point results depend on the GPU, the batch the request landed in, and the order of operations, so identical output is never guaranteed across two machines or two runs. Cosine similarity between two such vectors should still come out very close to 1.0. If you compare a repeated embedding of the same text, at the same model and the same dimensions width, and get meaningfully less than that, send us the examples through support and we will look into it.

Instructions for embedding inputs

Embedding models are trained to be steered by a short instruction in front of the text: what the text is for, what you intend to match it against. Qwen reports that dropping the query-side instruction costs roughly 1 to 5 percent on their retrieval benchmarks, so this is worth doing if you are tuning for retrieval quality.

There is no instruction field to set. We add nothing to your text, which means the instruction is simply the first part of the string you send, and you get to choose its wording and its format. The three recipes below are the ones worth copying.

1. A task instruction on qwen3-embedding-8b. This model's own convention is a two-line prefix on the query side, with documents left bare:

{"custom_id":"q1","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-embedding-8b","input":"Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: how do I cancel a running job"}}
{"custom_id":"d1","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-embedding-8b","input":"Cancelling a job stops new chunks from being dispatched. Work already running is drained and billed."}}

Replace the task sentence with your own: it describes the retrieval task, not the query. Keep it identical across every query in a corpus, because a query embedded under one instruction and a query embedded under another are not scored on the same footing.

2. A text query phrased like your image rows. On qwen3-vl-embedding-8b and qwen3-vl-embedding-2b, media rows are rendered through the model's chat template and text rows are not. To put a text query in the same format your images are in, write the template's own output into the string yourself:

<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\nYOUR QUERY<|im_end|>\n

which in a JSONL row, next to the image row it is meant to match, is:

{"custom_id":"img1","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"image_url","image_url":{"url":"https://example.com/bike.jpg"}}]}}
{"custom_id":"q2","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":"<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\na red bicycle leaning on a fence<|im_end|>\n"}}

Those marker strings are the model's own special tokens and are recognized as single tokens, so the query row's prompt matches the image row's exactly, other than the one end-of-text token described under pooling. Send this format for every text row in the corpus or none of them.

3. Your own instruction, on both sides. A media row is sent as a single user message, so its system turn always carries the template's default and there is no way to replace it. Put your instruction in the user turn instead, where a text row and a media row can both carry it: on the media row as a leading text part, on the text row as the same words in the same position.

{"custom_id":"img2","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":[{"type":"text","text":"Describe this image for searching purposes.\n"},{"type":"image_url","image_url":{"url":"https://example.com/bike.jpg"}}]}}
{"custom_id":"q3","method":"POST","url":"/v1/embeddings","body":{"model":"qwen3-vl-embedding-8b","input":"<|im_start|>system\nRepresent the user's input.<|im_end|>\n<|im_start|>user\nDescribe this image for searching purposes.\na red bicycle leaning on a fence<|im_end|>\n"}}

The trailing \n on the instruction is doing real work. Nothing is inserted between your content parts, so an instruction that does not end in a separator runs straight into the image placeholder that follows it.

Instructions that suit this model are short and say what the vector is for: Describe this image for searching purposes., Represent this product photo for retrieval by shoppers., Summarize this frame for matching against text captions. Which one wins is an empirical question on your own data, and the cost of trying three is three embedding runs.

You are billed for the instruction. It is part of the prompt, so it is counted as input tokens on every row that carries it, exactly like the rest of your text. A short instruction on a corpus of short rows is a real fraction of the job: 12 tokens of instruction on rows averaging 40 tokens is about a quarter of the bill.

Comparing across formats needs one convention. Vectors are only comparable when both sides were phrased the same way, at the same model, at the same dimensions width. Changing an instruction is a re-embed of the whole corpus, not of the queries alone, so settle the wording before you run the large job.

Decision jobs

Point a job at laya (see models) and each row asks typed questions about a text instead of prompting for a reply. The model reads the row's state - a string, or a JSON object such as {"body": "..."} that it reads as text - and answers every question in questions in one pass, with a probability per option. No text is generated: there is nothing to parse, there are no output tokens, and the answer to a question is always one of the options you defined.

JSONL only - url must be /v1/systemone, and body carries state and questions, plus an optional lang (below). One row is one text and the questions asked of it:

{"custom_id":"ticket-1","method":"POST","url":"/v1/systemone","body":{"state":{"body":"We were billed twice for March. Please refund it today."},"questions":{"team":{"type":"choice","instructions":"Which team should handle this?","criteria":{"billing":"invoices, payments, refunds","technical":"bugs and outages","other":"everything else"}},"urgent":{"type":"noul","instructions":"Does the customer need an answer today?"},"frustration":{"type":"score","instructions":"How frustrated is the customer?","criteria":["calm","annoyed","angry"]}}}}

A file's rows must all target /v1/systemone, and a body with any field beyond state, questions, lang and model fails the job naming the field. A row's own model is never read: the job's model is what answers.

Question types. Each entry of questions is named by you (the name comes back on the answer) and is an object with a type, an instructions string (the question itself, required), and criteria:

choice
One option out of several. criteria is an object of option key to description. The answer carries choice, the key of the most probable option, and probabilities, one per key.
noul
Yes or no. criteria is optional; when set, it is keyed exactly true and false (either or both), each with a description of what that answer means. The answer carries noul, the probability that the answer is yes, from 0 to 1.
score
A level on an ordered scale. criteria is a list of level descriptions, lowest first. The answer carries score, the expected level as a number from 0 to the index of the last level (1.3 on a three-level scale reads "between the second and third level, nearer the second"), and probabilities, one per level, keyed by its index "0", "1", ...

Every answer also carries its type, confidence (0 to 1, how concentrated the distribution is; on a yes-or-no question the larger of the two probabilities) and answer_confidence (the probability of the reported answer, comparable across the three types). The probabilities are the model's own estimates and are not calibrated: they rank options and flag doubt well, so gate on them with a threshold you have checked on your own data instead of reading 0.9 as nine in ten. A score answer adds a legend mapping each level index back to its description; any further field on an answer is the model's own and not part of this contract.

Limits, per row: at most 64 questions, a state of at most 50,000 characters (an object counts as its JSON rendering), and a budget of 192 tokens that a question's option texts share: the more options a question has, the fewer tokens of each the model reads, and a set that cannot fit is refused. A row over the first two fails the job before approval with failure_code: invalid_input naming the row; a question the model itself refuses - options over the budget, a yes-or-no question whose criteria carry a key other than true or false, a missing instructions - fails that row alone, with the model's message in the row's error and error_type: invalid_request_error, and every other row is delivered. The text one question is read over is capped by the model's context (see models): a longer text is cut to it.

Languages. The model reads more than 100 languages and picks one of its two internal checkpoints for each row by the text's script and language; the answer's routing block names the one it used. A row may carry lang, a language code such as "de", to settle the choice for a text whose script alone does not.

Writing questions that work. The model reads the option keys and the criteria text, so make both descriptive: {"billing": "invoices, payments, refunds"} is answered far better than {"b": ""}. Keep the option set short. A question with dozens of options loses accuracy and crowds the option budget; split it into a coarse question and a finer one, or ask a yes-or-no question per candidate. Say what a yes means: a noul with criteria for both answers reads better than one with the bare question. The model's publisher reports ordered scales as its weakest question type (see the model page); where a choice between named levels would do, ask a choice.

Generation-only fields don't apply: max_output_tokens, reasoning, reasoning_budget, reasoning_effort, response_format and video_fps are rejected with 422 on a decision model, and so is the embedding models' dimensions.

Tokens and billing. A decision job bills input tokens only, at $0.005 per 1M (0.0005 cents per 1k). The model reads the text once per question, so a row's billed tokens are, for each question, the text plus that question's instructions and option texts, summed over the questions: a 300-token text asked ten short questions bills a little over 3,000 tokens. The estimate comes from an exact token count rather than a sample, so a decision job reaches cost_ready in seconds, and its est_output_tok is always 0. Each result row reports the tokens it was billed in usage.input_tokens; usage.output_tokens is always 0. The minimum job price is $0.05 (see the pricing page). See decision output for the result file.

POST /inputs - ingest an input from a URL

The API can copy your input file into storage from any HTTPS URL - a presigned GET URL from your own bucket is the simplest source. (The console uploads from the browser instead; this endpoint is API-only.)

curl -X POST "$BASE/inputs" -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' -d '{
  "source_url": "https://your-bucket.s3.amazonaws.com/path/input.jsonl?X-Amz-...",
  "format": "jsonl",
  "callback_url": "https://you.example.com/hooks/ingest"
}'

{"id":"<ingest-id>","status":"pending","input_s3_uri":"s3://.../<ingest-id>.jsonl",
 "callback_secret":"whsec_..."}

This endpoint requires credits: an organization whose balance is zero or below is answered 402 here, before an upload target is minted or a source URL is contacted, so nothing can be placed in storage before anything is paid for.

format (jsonl) is optional. callback_url is optional, and callback_secret comes back only when you supply one - it is returned on this response only, never again.

data_residency is optional: "world" or "eu". Omit it and an organization with EU data routing gets "eu", any other "world"; set "world" to keep one input in the world region. Asking for "eu" without EU data routing answers 403 (see turning it on). "eu" copies the file into storage in the EEA, and only the EU API (https://api-eu.anex.sh/batch/v1) takes that request: the main base URL answers it with 421. A job submitted against the input must have the same data_residency - see POST /jobs.

Then poll until the copy has finished:

curl "$BASE/inputs/<ingest-id>" -H "Authorization: Bearer $TOKEN"

{"id":"<ingest-id>","status":"ready","input_s3_uri":"s3://...","error":null}

status goes pending → copying → ready, or failed with an error. With a callback_url registered you get an input.ready or input.failed callback instead of polling.

Keep the returned id (or input_s3_uri): input_id is preferred for job submission. Both source_url and callback_url must be https://, and a source that advertises a size over the input cap is refused with 413 before any bytes are copied. A source that does not advertise one, or is too slow to answer, is accepted here and refused during the copy instead: the ingest then settles as failed with the reason in error.

The copied object is kept for 7 days, the same window an uploaded one gets, so submit the job well inside it. Once the window lapses the input reports expired and a submit against it is refused - create a new input instead. While ready, the response's expires_at carries that deadline.

POST /inputs - or upload a file directly

If the file lives on your machine rather than somewhere we can fetch it from, call the same endpoint with no source_url - just a format:

curl -X POST "$BASE/inputs" -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' -d '{"format": "jsonl"}'

{"id":"<id>","status":"awaiting_upload",
 "upload_url":"https://...s3.amazonaws.com/",
 "fields":{"key":"...","policy":"...",...},
 "expires_at":"2026-07-26T15:04:05Z",
 "input_s3_uri":"s3://.../uploads/.../<id>.jsonl"}

The response is a presigned upload target rather than a copy job. Upload with a form POST: send every entry of fields as a form field, then the file itself as the file field, last.

data_residency is optional here too: "world" or "eu". Omit it and an organization with EU data routing gets "eu", any other "world"; set "world" to keep one upload in the world region. Asking for "eu" without EU data routing answers 403 (see turning it on). With "eu" the presigned target is storage in the EEA, and the file goes straight there from your machine. Ask the EU API (https://api-eu.anex.sh/batch/v1) for that target; the main base URL answers with 421. A job submitted against the upload must have the same data_residency - see POST /jobs.

curl "$UPLOAD_URL" \
  -F key=... -F policy=... -F x-amz-algorithm=... \
  -F x-amz-credential=... -F x-amz-date=... -F x-amz-signature=... \
  -F file=@data.jsonl

The upload must start before expires_at (about an hour out); a lapsed window can't be refreshed - create a new input instead. A file over the input cap is rejected by the upload itself, at the S3 level, rather than by this API. GET /inputs/{id} reports awaiting_upload until the object lands, then ready, and echoes expires_at while the window is open - you can submit the job immediately after uploading, since submission settles readiness itself. An uploaded object is kept for 7 days, so submit well inside that window; a ready input's expires_at carries that deadline, and past it the input reports expired. callback_url isn't available in this mode (there's no server-side copy to report on) and is rejected with 400 if supplied.

POST /jobs - submit a job

curl -X POST "$BASE/jobs" -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' -d '{
  "input_id": "<id>",
  "model": "<model-id>"
}'

{"id":"<job-id>","status":"submitted"}
input_id
Required unless input_s3_uri is given. The id from an ingest or upload. Preferred: it lets this call settle an upload for you, so you can submit right after uploading with no separate poll. Submitting one that isn't ready yet answers 409.
input_s3_uri
Required unless input_id is given. The input_s3_uri from an ingest. Give exactly one of input_id / input_s3_uri.
model
Required. A model id from the allowlist; the console's model picker lists the ids on offer. An unknown id is rejected before anything is created.
max_output_tokens
Optional, 1–32,768. Omit it and the cap is estimated from a sample of your rows: the value every row is expected to stay under with 95% confidence, which is also what the job is priced on. Supply it and your number is the cap and the price basis; on a reasoning model the sample is still taken to set the thinking budget unless you set reasoning_budget yourself or turn reasoning off.
max_tokens
The former name of max_output_tokens. Still accepted (and echoed back) so existing integrations keep working; new code should use max_output_tokens, which says what it is: an output cap, thinking included, with the input not part of it.
reasoning
Optional boolean, for the models marked "text (reasoning)" in the models table above. It is the job-level default; a JSONL row that sets chat_template_kwargs.enable_thinking itself keeps its own setting. Omit the field to leave the model's own default in place. true on a model that produces no reasoning trace is refused (422 at submit, or a rejected row naming the line), rather than accepted and answered without one; false stays valid on every generation model, so one body submits across a mixed model set. See thinking.
reasoning_budget
Optional, 0–32,768, for the models marked "text (reasoning)" in the models table above. Omit it and it is estimated from the same sample as the output cap: the thinking length every row is expected to stay under with 95% confidence, so rows that think longer are cut off gracefully and still answer. Set a number to cap thinking yourself (at least 64 below an explicit max_output_tokens), or 0 for no cap. A budget needs thinking on: a value of 1 or higher with reasoning: false is refused (422), and a JSONL row that sets chat_template_kwargs.enable_thinking to false in a job with your budget fails the job, naming the line. An estimated budget is never applied to a row with thinking off. When the sample shows that even the answers alone would not fit under your max_output_tokens, estimation fails and says so; raise max_output_tokens or set reasoning_budget. 0 is accepted on any generation model, reasoning or not, so the same submit body works across a mixed set; a value of 1 or higher is only for the models marked "text (reasoning)". See thinking.
reasoning_effort
Optional, for the models that offer thinking effort levels (see the note under the models table): how hard the model thinks before it answers. qwen3.8-27b takes low, medium or xhigh; leave the field out and it thinks at xhigh, its default. A lower level usually spends fewer thinking tokens, and thinking tokens are billed as output, so it lowers cost and time per row. A JSONL row can set its own level with chat_template_kwargs.reasoning_effort, which wins for that row. Refused (422) on a model without levels, for a value that is not one of the model's levels (they are case-sensitive), and together with reasoning: false. It works together with reasoning_budget. See thinking.
video_fps
Optional number above 0, for generation models that take a video sampling rate: frames sampled per second of clip, applied to every video row of the job. Omit it and each clip is sampled at the model's default rate. qwen3.8-27b defaults to 1 and takes up to 20; nemotron-3-nano-omni-30b defaults to 2 and takes up to 20; gemma4-12b does not take a rate yet. Each sampled frame is processed at the model's full per-frame resolution budget, and is charged as input tokens at the model's media input rate: a clip costs tokens per frame × (seconds × rate, rounded up, + 1), plus the per-clip and per-minute charges, which do not change with the rate. So a higher rate costs proportionally more per second of clip and shortens the longest clip a row can carry; the longest clip per model and rate is in limits. A rate above the model's ceiling, or any rate on a model that takes none, answers 422. Embedding models sample video at a fixed 1 frame per second and do not take this field.

A rate is capped by the clip itself. Sampling cannot produce frames a clip does not contain, so a clip shot at 2 frames per second is sampled at 2 however high you set this, and is charged and length-checked at 2. Ask for more than your clips hold and you pay for what you get, not what you asked for. If you want a higher rate, send clips encoded at or above it.
response_format
Optional. Applied to every row. See structured output.
dimensions
Embedding models only, optional. Truncates the output vector to a narrower width (32 up to the model's native width). A JSONL row's own body.dimensions wins over this for that row. 422 on a non-embedding model. See embedding jobs.
data_residency
Optional, "world" or "eu" - where the job's payload is stored and processed. Omit it and an organization with EU data routing gets "eu", any other "world"; "world" keeps one job in the world region. "eu" without EU data routing answers 403. "eu" keeps the job's input, output and processing inside the EEA; submit such a job to the EU API (https://api-eu.anex.sh/batch/v1), since the main base URL answers it with 421. The input must be stored in the same region as the job, or the submit is rejected with 422 - see uploading an input.

Format, size, and the shape of response_format are checked before the job exists; a failure there returns an error and creates nothing. Row count and the row-by-row checks (content parts, inline media size, whether a row fits the model) run afterward: row counts appear once the input is validated after submission, and invalid input fails the job asynchronously with the same error reasons as before.

Submission has no idempotency key: every POST /jobs that succeeds creates a new job. If a submit times out without a response, check for the job (the console's Jobs page lists every job of the organization) before retrying, or you may create it twice. Nothing runs without approval either way, so a duplicate costs nothing until someone approves it.

A credit balance of zero or below answers 402, before any of your input is read and before a job is ever priced - a free grant or a paid top-up, either counts; see credits and billing. The same check guards uploads, so an account with no balance cannot place a file either. This is a coarser check than approval's: it only asks whether you have any balance at all, not whether it covers this particular job.

An embedding model rejects max_output_tokens, reasoning, and response_format with 422 - see embedding jobs for the full field list. A decision model rejects the same fields, plus video_fps and dimensions, and takes JSONL input only - see decision jobs.

An organization can hold at most 10 jobs awaiting approval. Past that cap, submit answers 429 with a Retry-After header in seconds until some are approved, cancelled, or expired; back off for that long, then resubmit. The same status and header also meet a caller sending requests faster than the API accepts. The two are told apart by the body: the cap's message names the approval queue, the rate limit's error.code is rate_limit_exceeded.

GET /jobs/{id} - poll a job

curl "$BASE/jobs/<job-id>" -H "Authorization: Bearer $TOKEN"

{"id":"<job-id>","status":"cost_ready",
 "est_input_tok":120000,"est_output_tok":40000,"est_cost_cents":72,
 "est_processing_seconds":900,"max_output_tokens":600,"reasoning_budget":null,
 "success_rows":null,"error_rows":null,"format_error_rows":null,
 "truncated_rows":null,"pending_rows":null,
 "created_at":"...","approved_at":null,"data_expires_at":null,"partial":null,
 "failure_code":null,"failure_reason":null,
 "model":"qwen3.8-27b","reasoning":true,"task_type":"chat","format":"jsonl",
 "response_format":null,"dimensions":null,"video_fps":null,"data_residency":"world",
 "data_deleted_at":null}

status is the job's current state; the values and the order they come in are described under job lifecycle.

est_input_tok, est_output_tok, est_cost_cents
The estimate, filled in when the job reaches cost_ready. est_cost_cents is the amount approval holds and the most the job can consume. A job that does any work costs at least 1 credit, so both this estimate and the final charge are floored there; a job that consumed nothing is not charged. Each model also has a minimum job price - the cost of the GPU machine start a small job forces, never added on top of a job whose token price clears it - so a large job pays pure per-token rates; see the pricing page for the current minimums. An embedding job's est_output_tok is always 0, and its est_input_tok comes from an exact token count rather than a sample, so it reaches cost_ready in seconds - see embedding jobs. A decision job is estimated the same way, with the text counted once per question.
est_processing_seconds
Expected wall-clock run time from approval to results, in whole seconds, frozen at estimate time: the fixed startup overhead recent jobs of this model have paid (booting a machine, pulling the weights) plus the job's own expected generation at the model's measured throughput. Advisory; most jobs are dominated by the startup overhead.
max_output_tokens
The effective per-row output cap: yours, or the estimated one (null until estimated). An embedding job reports 0: it has no completion budget.
reasoning_budget
The effective thinking budget: yours, or the estimated one (null until estimated). Always null on a non-reasoning model, which cannot set this field at all.
model
The model the job ran against.
reasoning
The job-level thinking setting. null means you set no job-level default and the model's own behaviour applied - it does not mean thinking was off. A JSONL row that carried its own chat_template_kwargs.enable_thinking overrides this for that row, which this job-level field cannot report.
reasoning_effort
The job-level thinking effort level, or null when you set none and the model's default level applied. A JSONL row that carried its own chat_template_kwargs.reasoning_effort overrides this for that row, which this job-level field cannot report.
task_type
chat (a completion per row) or embed (a vector per row).
format
The input file's format, jsonl.
response_format
The job-wide structured-output schema, if you set one. A JSONL row may carry its own, which wins for that row.
dimensions
The embedding width on an embed job; null means the model's full width.
video_fps
The video sampling rate you set at submit; null means the model's default rate applied. A clip whose own frame rate is lower is sampled, charged and length-checked at its own rate.
data_residency
Where the job's payload is stored and processed: "world" or "eu". Either API answers this poll; the result of an "eu" job is downloaded from the EU API.
success_rows, error_rows
Rows that produced an answer, and rows that failed. Every row count is null until the job has finished (done, failed or cancelled); a job in progress reports its status alone.
format_error_rows
Rows whose constrained generation could not satisfy the response_format. Counted separately, not in error_rows.
truncated_rows
Rows whose generation stopped on the output-token cap. Counted separately, not in error_rows. On a reasoning model, reasoning_budget keeps thinking from consuming the whole cap, so rows still answer instead of truncating.
pending_rows
Rows not yet accounted for by any of the four counts above.
partial
True on a terminal job that still shipped the rows that finished.
data_expires_at
When the job's result is deleted: it is kept for 30 days after the job finishes, then purged. The input file goes sooner, 7 days after it was uploaded or ingested, so this timestamp tracks the result. Set once the job reaches a terminal state (done, failed, cancelled), null before that. Download the result before this time; afterwards the result endpoint answers 410 Gone. The job record itself (statuses, counts, cost) stays queryable.
data_deleted_at
When the job's data was actually deleted, whether by the 30-day purge or earlier through DELETE /jobs/{id}/data; null while the input and result still exist.
failure_code, failure_reason
Set on a failed job. failure_code is a stable value (invalid_input or internal_error); failure_reason is a readable message. Four media-specific invalid_input reasons: a content part whose type the model has no encoder for (names the model and modality), a clip whose duration exceeds the model's cap (worded differently for a declared vs. an estimated length - see duration_seconds), an inline media item over its size cap (same wording as the 400 a bad row gets at submit, naming the line), and a row whose input tokens, media included, do not fit the model's context (see models). The last two media reasons read like this, naming the row: Row 'x': audio duration 45s exceeds the 30s maximum for model 'gemma4-12b'. Shorten or split this audio input. and Row 'x' is too long: 37,500 input tokens (including media) exceeds the 32,768-token maximum that model 'nemotron-3-nano-omni-30b' can serve with audio input. Shorten or split this row. A video clip longer than the model takes at the job's sampling rate reads like this: Row 'clip-7': video is 45 s; at the requested 20 fps the longest clip qwen3.8-27b takes is 38 s. Shorten the clip or lower video_fps. A job that fails this way is never billed. Jobs that go through estimation (no max_output_tokens, or a reasoning model with no reasoning_budget) can also fail during estimation with invalid_input when the output does not fit the model: the rows need more output tokens than the model can generate, the model's reasoning runs too long on them, the typical answer hits the model's maximum output length and would be cut off, or (with an explicit max_output_tokens on a reasoning model) the answers alone would not fit once the thinking budget and its cut-off notice are reserved. A job on a model we cannot price automatically fails during estimation too, usually with invalid_input and the reason We could not prepare a price for this job automatically. Each reason states the fix: set or raise max_output_tokens, set reasoning_budget (0 for no cap), or disable reasoning.

POST /jobs/{id}/approve - approve a job

Approval is what starts paid GPU work. It holds est_cost_cents against your prepaid balance; see credits and billing.

curl -X POST "$BASE/jobs/<job-id>/approve" -H "Authorization: Bearer $TOKEN"

{"id":"<job-id>","status":"in_progress"}

A job must be in cost_ready: any other state answers 409. A balance that does not cover the estimate answers 402 - top up and approve again. With auto-approve turned on for the organization, this call is made for you.

An estimate waits at most 24 hours: a job left in cost_ready past that window is cancelled automatically, free of charge. An organization can hold at most 10 jobs awaiting approval; past that cap, submit answers 429 with a Retry-After header until some are approved, cancelled, or expired.

POST /jobs/{id}/cancel - cancel a job

curl -X POST "$BASE/jobs/<job-id>/cancel" -H "Authorization: Bearer $TOKEN"

{"id":"<job-id>","status":"cancelling"}

A job with no work in flight stops immediately and answers cancelled. A job that has already started answers cancelling: no new work is started, rows already running finish, and the job then settles as cancelled with the completed rows downloadable as a partial result. Once a job is finalizing or already terminal it is too late, and the call answers 409.

A job that is estimating answers 409 too: wait for cost_ready and simply never approve it - estimates are free, and an unapproved job cancels itself after 24 hours.

The exception is a job that keeps its data in the EEA. It can take up to about an hour to estimate when no machine is ready to run its model, so cancelling it while it is estimating stops it immediately and answers cancelled.

GET /jobs/{id}/result - download the result

curl "$BASE/jobs/<job-id>/result" -H "Authorization: Bearer $TOKEN"

{"id":"<job-id>","status":"done","result_url":"https://...signed...",
 "expires_in":3600,"partial":false,"row_count":1000}

result_url is a presigned link you GET to download the file; expires_in is how many seconds it stays valid. Request the endpoint again for a fresh link. row_count is the number of rows actually in the file. What is inside it is described under output format.

A result exists for a done job, and for a failed or cancelled job that delivered the rows it had finished - that one comes back with "partial": true. A job with nothing to deliver answers 409.

A result is deleted 30 days after the job finishes. After that this endpoint answers 410 and the data is gone - download anything you want to keep well inside that window. You can also delete the data yourself as soon as the job has finished, see below.

DELETE /jobs/{id}/data - delete a job's data

curl -X DELETE "$BASE/jobs/<job-id>/data" -H "Authorization: Bearer $TOKEN"

{"id":"<job-id>","status":"done","data_deleted_at":"2026-07-24T10:30:00+00:00"}

Removes the job's input file and its result file (and every intermediate object) right away, instead of at the 30-day mark. The job record - statuses, row counts, cost, dates - stays, and data_deleted_at on GET /jobs/{id} records when. Download the result first if you still need it: after this call the result endpoint answers 410 Gone, and nothing can bring the data back.

A finished job - done, failed, or cancelled - is purged at once and the call answers 200. A job still in flight is first cancelled, exactly as POST /jobs/{id}/cancel would (rows already completed are delivered and billed), and its data is deleted automatically as soon as the job has stopped; the call answers 202 Accepted with data_deleted_at null and a message saying so. Poll GET /jobs/{id} until data_deleted_at is set if you need to know when:

{"id":"<job-id>","status":"cancelling","data_deleted_at":null,
 "message":"the job is being cancelled; its input and any result will be deleted once it stops"}

The call is idempotent: repeating it on a job whose data is already gone answers 200 with the original data_deleted_at, and 202 again while the job is still stopping. The console offers the same action as Delete data on every job.

Webhooks

An organization can register one endpoint (an owner does this on the console's Settings page). It then receives a signed POST for every job event, for every job in the organization - there is no per-event or per-job subscription.

submitted
The job was created.
estimating
Output length is being measured to price the job.
cost_ready
The estimate is ready. Carries est_input_tok, est_output_tok and est_cost_cents.
in_progress
The estimate was approved and the credits are held; the job is being run. No further event is sent until it finishes.
completed
The job finished. Carries status, result_url and result_expires_in, so the event is directly actionable.
failed
The job failed. Carries status, partial, the failure_code/failure_reason when one was recorded, and the result URL when there is something to download.
cancelled
The job was cancelled. Same payload as failed.

A test event is also sent by the Settings page's test button. Every event has the same envelope; the ones not listed with a payload above carry only job_id.

{"type":"completed","timestamp":"2026-07-22T10:31:04.512+00:00",
 "id":"<event-id>",
 "data":{"job_id":"<job-id>","status":"done",
         "result_url":"https://...signed...","result_expires_in":3600}}

Deliveries follow the Standard Webhooks convention. Each POST carries webhook-id, webhook-timestamp (unix seconds) and webhook-signature headers. To verify, base64-decode the secret with its whsec_ prefix stripped, HMAC-SHA256 the bytes {webhook-id}.{webhook-timestamp}. followed by the raw request body, base64-encode the digest and compare it with the v1,-prefixed value in webhook-signature. Verify against the exact bytes received, not a re-serialized copy.

Only a 2xx response counts as delivered. Anything else is retried on a widening backoff for about a day, then given up. Deliveries are HTTPS only and redirects are never followed.

Ingest callbacks are separate: they go to the callback_url of one ingest, are signed with that ingest's own callback_secret, carry input.ready / input.failed, and their envelope timestamp is unix seconds rather than a date string.

Realtime embeddings

A separate, OpenAI-compatible surface for single-request embeddings - for a query vector at request time rather than a batch of rows. Point the OpenAI SDK at it directly:

from openai import OpenAI
client = OpenAI(base_url="https://api.anex.sh/realtime/v1", api_key="<your-api-key>")

The base URL is https://api.anex.sh/realtime/v1, and it takes the same sk- API keys as the batch API above - no separate key needed. Three endpoints: GET /models lists the models on offer, POST /embeddings returns a vector for one request, and POST /systemone answers typed questions about a text (see realtime decisions below):

curl https://api.anex.sh/realtime/v1/embeddings \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3-embedding-8b", "input": "a search query", "dimensions": 1024}'

input is a string or a list of strings; model is one of the realtime-served models below. Text only - realtime requests carry no image or video content, and one that does is rejected; embed media on the batch API instead (see embedding jobs), and the resulting vectors share the same space as a realtime text query on the same model at the same dimensions.

A request may carry at most 128 inputs, 262,144 characters in any one input, and 2,097,152 characters across all of them together; past any of those the request is rejected with 400. Split a larger set of texts across requests, or embed them on the batch API.

dimensions is optional and takes the same range as batch: 32 up to 4096 on qwen3-embedding-8b, 64 up to 4096 on qwen3-vl-embedding-8b and 64 up to 2048 on qwen3-vl-embedding-2b (text queries only in realtime - image and video stay batch-only for the two VL models too). encoding_format only accepts float.

A realtime query is preprocessed exactly like a batch text row: the string reaches the model verbatim, with no instruction and no chat template added, and the vector comes back at unit length. See how embedding inputs are processed for the details, including why a repeated query is not bit-identical.

Pricing is per input token at each model's realtime rate, listed in the realtime column of the embeddings pricing table, billed from the same credit balance as batch jobs.

Errors worth naming beyond the ones below: 402 when the organization has no credits, 429 on a concurrency or upstream rate limit (honor the Retry-After header), 501 when the model exists but isn't served in realtime, and 503 when the realtime API is frozen or billing is unavailable.

Realtime decisions

The same base URL and the same API keys serve typed decisions one request at a time: POST /systemone takes a text and a set of questions and answers them all in one pass, in tens of milliseconds, with a probability per option and no generated text. The body is the body of a decision job row plus the model, and the question types, limits and advice there apply here unchanged:

curl https://api.anex.sh/realtime/v1/systemone \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "laya",
       "state": {"body": "We were billed twice for March. Please refund it today."},
       "questions": {
         "team": {"type": "choice", "instructions": "Which team should handle this?",
                  "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs and outages", "other": "everything else"}},
         "urgent": {"type": "noul", "instructions": "Does the customer need an answer today?"},
         "frustration": {"type": "score", "instructions": "How frustrated is the customer?",
                         "criteria": ["calm", "annoyed", "angry"]}}}'

The response carries one entry per question under answers, keyed by the names you gave them (the figures below only illustrate the shape):

{"model": "laya",
 "answers": {
   "team": {"type": "choice", "choice": "billing",
            "probabilities": {"billing": 0.91, "technical": 0.03, "other": 0.06},
            "confidence": 0.84, "answer_confidence": 0.91},
   "urgent": {"type": "noul", "noul": 0.88, "confidence": 0.88, "answer_confidence": 0.88},
   "frustration": {"type": "score", "score": 1.3,
                   "legend": {"0": "calm", "1": "annoyed", "2": "angry"},
                   "probabilities": {"0": 0.12, "1": 0.46, "2": 0.42},
                   "confidence": 0.31, "answer_confidence": 0.46}},
 "usage": {"input_tokens": 171, "output_tokens": 0},
 "routing": {"model": "english", "reason": "..."}}

usage.input_tokens is what the request bills: the text counted once per question plus each question's own words, the same rule as a batch row. routing names the internal checkpoint the request was answered by and why; it is informational. GET /models lists this model with "task": "decide", beside the embedding models' "task": "embed".

A request may carry at most 64 questions and a text of at most 50,000 characters; past either it is rejected with 400. A question the model refuses - a yes-or-no question whose criteria carry a key other than true or false, a missing instructions, an option set that cannot fit its 192-token budget at all - comes back as 422 with the model's own message naming the question.

Pricing is per input token at the realtime rate in the decisions pricing table, billed from the same credit balance as batch jobs. A request that was refused bills nothing.

Errors worth naming beyond the ones below: 400 when the model named is an embedding model (the message points at POST /realtime/v1/embeddings; naming laya there points back here), 402 when the organization has no credits, 429 on the concurrency limit (honor the Retry-After header), 502 when the decision service failed to answer, and 503 when the realtime API is frozen or billing is unavailable.

Errors

An error response carries a detail field describing it:

400
The request or the input file failed validation - a non-HTTPS URL, a file over the size cap, a file that is not JSONL, a malformed response_format, a missing format on a direct upload, or a callback_url supplied in upload mode.

A jsonl input's own row count and its rows' content are checked only once the job exists, not as part of this 400: those checks run during estimating (or, for an explicit max_output_tokens, while the job is chunked), and a violation fails the job with failure_code: invalid_input rather than answering the submit call itself. Two of those checks read
  • unsupported content part type '<type>': supported types are ... - a content part whose type we don't recognize at all.
  • an embedded <image|audio file|video file> exceeds the size limit (max N MB) - an inline media item over the cap in limits.
naming the offending line. A few more surface the same way: a content part whose type your model has no encoder for, a clip whose duration (declared or estimated) exceeds the model's cap, and a row that exceeds the model's context, media included - see the input file above. Nothing is billed either way. A video clip longer than the model takes at the job's sampling rate is refused with a message naming the clip length, the rate and the longest clip the model takes at it: Row 'clip-7': video is 45 s; at the requested 20 fps the longest clip qwen3.8-27b takes is 38 s.
401
The API key is invalid or no longer active.
402
The balance does not cover the estimate on approval.
403
The request asks for "data_residency": "eu", or uses the EU realtime address, for an organization without EU data routing: {"detail": "EU data routing is not enabled for this organization. Contact us to enable it."}. We turn it on per organization on request; see data residency.
404
Unknown id, or an id belonging to another organization.
409
Wrong state - approving a job that is not awaiting approval, cancelling one that is past the point of stopping, submitting an input_id that is not ready yet, or asking for a result that does not exist yet.
410
The result data has been deleted - job payloads are kept for 30 days after completion.
413
The ingest source advertises a size over the input cap. A direct upload over the cap is rejected by the upload itself instead.
421
The data this request reads or stores is kept in the other region, and this API does not serve it: creating or polling an input, submitting a job, or downloading a result for data kept in the EEA, sent to the main base URL (or the reverse). The body names the region, and api the origin to send the same request to: {"detail": "this job is stored in the EU region; use https://api-eu.anex.sh", "residency": "eu", "api": "https://api-eu.anex.sh"}.
422
The request does not match the endpoint's shape - a missing Authorization header, a missing or wrongly typed body field, or an unknown model id. A video_fps the model cannot take answers here too: video_fps 25 is above the 20 fps ceiling of qwen3.8-27b, or on a model that takes no rate model 'gemma4-12b' does not support video_fps. A decision job submitted with a generation-only field, answers here as well; on the realtime surface, a question the decision model refuses comes back as 422 with the model's own message.
429
Too many jobs awaiting approval in the organization (see submit), too many concurrent realtime requests, or requests arriving faster than the API accepts. Wait the seconds in the Retry-After header, then retry.
503
A check the request depends on could not run - the key could not be validated, or the balance could not be read, so the approval was not made. Also answered while an operator has the API frozen for maintenance (downloading a result keeps working then). Retry.

5xx responses and timeouts are transient: retry with backoff. Approving and cancelling are guarded, so a retried approval never holds credits twice - but submitting is not (see submit a job): confirm a timed-out submit before repeating it.