# anex.sh > LLM batch inference for open models. Submit a file with one prompt per row, get a cost > estimate, approve it, and download one result file with a response per row. Payment > is prepaid credits; an approved job never consumes more than the estimate you > approved. Media-capable models also take images, audio, and video in the rows, > and a decision model answers typed questions about each row instead of writing. Last updated: 2026-09-26. Three ways in, on the same organization and the same API keys: - Browser console: sign in, submit, approve, and download from the browser. - Batch HTTP API: base URL https://api.anex.sh/batch/v1, with an API key from the console's Keys page sent as `Authorization: Bearer `. - Realtime API: base URL https://api.anex.sh/realtime/v1. OpenAI-compatible embeddings at POST /embeddings, and typed decisions at POST /systemone. The only surface that answers immediately; everything else is batch. ## When to use anex.sh Use it when you have many prompts to run through an open-weights model and can wait for the whole set rather than needing each answer at once. Typical fits: - Classifying, tagging, extracting fields from, or summarizing thousands to millions of documents, tickets, product listings, transcripts, or web pages. - Generating structured JSON per row (response_format with a JSON schema) for a data pipeline, at a known ceiling price. - Describing, captioning, or transcribing large sets of images, audio, or video with a media-capable model. - Computing embeddings for a whole corpus (batch), with the same model then serving query-time vectors on the realtime endpoint so both live in one vector space. - Routing, triaging or tagging thousands to millions of texts with the decision model (laya): one answer per typed question (a choice, a score, a yes or no) with a probability per option, no text to parse, at the lowest rate here; the same model then answers live traffic on the realtime endpoint. - Running an evaluation or a synthetic-data job over a fixed dataset with a reasoning model, with the thinking budget capped per row. Do not use it for a single interactive chat turn or for a request that must answer in seconds: the batch API estimates, waits for approval, rents a GPU, and then runs the file (minutes to hours). Embeddings and decisions are the two exceptions, on the realtime endpoint above. How an agent should call it: create an API key on the console's Keys page (a human does this once), then POST /inputs (upload the file to the presigned target if no source_url), POST /jobs with {"input_id": ..., "model": ...}, poll GET /jobs/{id} until status is cost_ready, read est_cost_cents, POST /jobs/{id}/approve, poll until status is done, GET /jobs/{id}/result and download result_url. Call DELETE /jobs/{id}/data afterwards if the data must not outlive the job. There is no list-jobs and no batch list-models endpoint; the full operation list with schemas is the OpenAPI document below. ## Developer resources - OpenAPI 3.1 document: https://anex.sh/openapi.json (also /openapi.yaml). Every operation has an operationId (createInput, getInput, createJob, getJob, approveJob, cancelJob, getJobResult, deleteJobData, listRealtimeModels, createEmbedding, createDecision), typed parameters, and response schemas, ready for function calling or an SDK generator. - API reference (prose): https://anex.sh/docs/api. Authentication: https://anex.sh/docs/api#basics. Errors: https://anex.sh/docs/api#errors. - Every documentation page answers `Accept: text/markdown` with a Markdown rendering of itself, at the same URL. - Documentation search index: https://anex.sh/docs/search.json. One JSON entry per section of every docs and model page - `p` the page path, `a` its anchor, `h` the heading, `t` the section's plain text - so a whole-corpus lookup costs one fetch and a result addresses as `

#`. It is what the docs search box reads. - Console (sign in, keys, credits): https://console.anex.sh/dashboard - Site map: https://anex.sh/sitemap.xml. Support: support@anex.sh. ## Docs - [Overview](/docs): what the service is, the job lifecycle (submitted, estimating, cost_ready, in_progress, finalizing, done), and the input limits. A job is in_progress from the moment it is approved. Batches run at high throughput but take tens of minutes to come back: a machine is rented and the model loaded before any row is processed, so a job that sits in in_progress for a while is normal and is not a stalled job - [Console guide](/docs/console): organization roles, API keys, credits and billing, submitting, approving and cancelling, auto-approve, webhooks, results - [API reference](/docs/api): every endpoint with request and response shapes, the model table (ids, exact weights builds, quantization, accepted inputs), how embedding inputs are processed (no instruction or chat template is added to your text, vectors are last-token pooled and L2 normalized to unit length; media rows are rendered through the model's chat template and text rows are not), worked examples of writing your own instruction into a row and of phrasing a text query the way an image row is phrased, and decision jobs (the /v1/systemone row shape, the three question types, their limits and how their tokens are billed) with the realtime decision endpoint beside them - [Output and formats](/docs/output): the result file for JSONL and parquet jobs, embedding output, decision output, structured output (response_format), thinking/reasoning - [Chat templates](/docs/chat-templates): the Jinja chat template each model renders rows with, downloadable per model at /static/chat-templates/.jinja, each linked to the upstream file it was copied from at a pinned commit (also in /static/chat-templates/index.json); and the placeholder string an image, audio clip or video part becomes. The templates are served exactly as the checkpoint publishes them, never patched or overridden. One template reorders your content parts: nemotron-3-nano-omni-30b hoists every media placeholder to the front of the message, images then video then audio, with all of your text after them - [Pricing details](/docs/pricing): per-model input, cached-input, output, per-image and per-clip rates, embedding rates, the per-model minimum job price, and example job costs. Most media-capable models carry two rate rows, one for text-only rows and a higher one for rows carrying any image, audio or video, because media rows cost more GPU time per token; qwen3.6-27b is the exception and prices both the same. Every rate and every flat per-item charge is set on the middle card of the range a model's rows actually run on, at that card's current market price. Per 1M tokens, gemma4-12b is $0.12 input / $0.03 cached / $0.60 output on text rows and $0.30 / $0.075 / $1.50 on media rows ($0.0002 per image, $0.0002 per clip, $0.05 minimum job price); nemotron-3-nano-omni-30b is $0.12 / $0.03 / $0.60 on text and $0.30 / $0.075 / $1.50 on media ($0.0011 per image, $0.0002 per clip, $0.06 minimum job price). Neither carries a per-minute media surcharge. qwen3.8-27b is $0.18 / $0.045 / $0.90 on text rows and $0.30 / $0.075 / $1.20 on image and video rows ($0.0001 per image, $0.0004 per clip, $0.0003 per video minute, $0.12 minimum job price). A video row is charged per sampled frame, so a higher video_fps costs proportionally more per second of clip. The decision model laya bills input tokens only: $0.005 per 1M in batch, $0.033 per 1M on the realtime endpoint, $0.05 minimum job price, with the text counted once per question - [Price calculator](/price-calculator): what a job of a given shape costs. Enter a typical input and output, the number of rows, and any images, audio or video per row, and it counts the tokens on the chosen model's own tokenizer and prices them on the same rate cards that bill a real job. It reports both the settled bill and the credit hold placed at approval ## Models The eleven ids to pass as `model`, and what each takes and returns: - glm-5.2: text in, reasoning text out. The largest model we serve - qwen3.8-27b: text, image, video in; reasoning text out. Video sampled at 1 frame/s by default, up to 20 with video_fps - qwen3.6-27b: text, image in; reasoning text out - gemma4-31b: text, image in; reasoning text out - gemma4-12b: text, image, audio, video in; reasoning text out. Short clips only, audio 30 s and video 60 s. Does not take video_fps yet - qwen3-14b: text in, reasoning text out - nemotron-3-nano-omni-30b: text, image, audio, video in; reasoning text out. The longest audio here, up to 40 min per clip. Video sampled at 2 frames/s by default, up to 20 with video_fps. A short prompt it answers without thinking still lands in content (the output column), never in the trace - qwen3-embedding-8b: text in, embedding vector out - qwen3-vl-embedding-8b: text, image, video in; embedding vector out - qwen3-vl-embedding-2b: text, image, video in; embedding vector out - laya: text in; typed decisions out. Each row is a state (a string or a JSON object) plus named questions of type choice (criteria: option key to description), noul (yes or no; criteria optional, keyed true/false) or score (criteria: a list of level descriptions, lowest first), and the answer to each carries its probabilities and a confidence. No text is generated and no output tokens are billed. More than 100 languages. Batch rows are JSONL only, url /v1/systemone; the same body answers at POST /realtime/v1/systemone Each model's exact weights build and quantization is on [/docs/api#models](/docs/api#models), and the rates are on [/docs/pricing](/docs/pricing). Custom models are available on request: an open-weight model not listed here, or a customer's own fine-tune, served to that organization only through the same API. Ask at support@anex.sh. ## API endpoints All relative to https://api.anex.sh/batch/v1 and documented on [/docs/api](/docs/api): - POST /inputs: ingest an input file from an HTTPS source_url, or with no source_url get a presigned target to upload the file directly - GET /inputs/{id}: the input's status; submit once it is ready - POST /jobs: submit a job, body {"input_id": "...", "model": "..."} plus optional max_output_tokens, reasoning, reasoning_budget (omit = estimated, 0 = no cap; 1+ needs thinking on: 422 with reasoning false, and a JSONL row with enable_thinking false fails the job), response_format, video_fps (frames sampled per second of every video clip; omit = the model's default rate; qwen3.8-27b and nemotron-3-nano-omni-30b up to 20, gemma4-12b not yet; 422 above the ceiling) (embedding models: dimensions, video fixed at 1 frame/s). A parquet input needs two fixed columns, custom_id (unique per row) and input (a string column); any other columns are echoed into the result file - GET /jobs/{id}: status, the estimate (est_cost_cents is the approval hold and the cost ceiling), row counts (null until the job has finished), and the job's video_fps (null = the model's default rate) - POST /jobs/{id}/approve: approve the estimate. Required before any paid work starts: a job waits at cost_ready until this call (or the organization's auto-approve) happens. 402 if the balance does not cover the estimate. - POST /jobs/{id}/cancel: cancel; rows already completed are delivered and billed - GET /jobs/{id}/result: a presigned download link for the result file - DELETE /jobs/{id}/data: delete a job's input and result files now (the job record stays); otherwise the input is purged 7 days after it arrived and the result 30 days after the job finishes. A job still in flight is cancelled first and its data deleted once it stops (202) Realtime, relative to https://api.anex.sh/realtime/v1, same API keys: - GET /models: the realtime models; each entry's task is embed or decide - POST /embeddings: OpenAI-compatible, text only, {"model": "...", "input": "..." or ["..."], "dimensions": optional}; at most 128 inputs per request - POST /systemone: typed decisions, {"model": "laya", "state": , "questions": {"": {"type": "choice" | "noul" | "score", "instructions": "", "criteria": }}} -> {"model": "laya", "answers": {"": {"type": ..., "choice" | "noul" | "score": ..., "probabilities": {...}, "confidence": ..., "answer_confidence": ...}}, "usage": {"input_tokens": N, "output_tokens": 0}, "routing": {...}}. A batch decision job's JSONL row is {"custom_id": "...", "method": "POST", "url": "/v1/systemone", "body": {"state": ..., "questions": {...}}} and its result row's response.body is the same answer object; the result is always JSONL, since a decision job takes JSONL input only ## Limits - Input file: 1 GiB, .jsonl or .parquet, up to 10,000,000 rows per job - 10 jobs awaiting approval per organization; approval window 24 hours, then the job is cancelled free of charge - Inline media per item: image 20 MB, audio 90 MB on nemotron-3-nano-omni-30b and 20 MB on gemma4-12b, video 150 MB - Row context: 32,768 tokens per row on every generative model, and on the embedding models. Input including media at most 32,768, output at most 32,768, and input plus output within 32,768 on the cards in service today. The exception is a row carrying images or video on qwen3.8-27b: input including media up to 131,072 tokens, output at most 32,768, and both within 131,072 - Media duration, per model: audio 2,400 s (40 min) on nemotron-3-nano-omni-30b at 12.5 tokens/s (30,000 tokens, leaving room for a prompt and a short answer) and 30 s on gemma4-12b at 25 tokens/s; video 60 s on both (nemotron 2 frames/s at 256 tokens a frame, about 31,000 tokens; gemma4-12b 1 frame/s at 74 tokens a frame), 120 s on qwen3.8-27b (1 frame/s at 105 tokens a frame, at most 768 frames per clip). Images: 280 tokens on gemma4-12b and gemma4-31b, up to 3,328 on nemotron-3-nano-omni-30b (256 per 512 px tile, 13 tiles), up to 1,280 on qwen3.6-27b, up to 200 on qwen3.8-27b - Shortest video clip: a clip has to hold at least 2 frames. A one-frame video (a sub-second scene re-encoded at 1 frame per second) is refused before any paid work starts, with the row named; send that frame as an image part instead - Video sampling rate (video_fps): every sampled frame is processed at the model's full per-frame resolution budget and charged tokens per frame x (ceil(seconds x fps) + 1), so a higher rate fits a shorter clip. Longest clip at the default rate / 5 / 10 / 20 fps: qwen3.8-27b 120 s (1 fps) / 120 s / 76 s / 38 s; nemotron-3-nano-omni-30b 60 s (2 fps) / 25 s / 12 s / 6 s; gemma4-12b 60 s at 1 fps, no other rate yet - video_fps is capped by the clip's own frame rate: sampling cannot produce frames a clip does not contain, so a 2 fps clip is sampled, CHARGED and length-checked at 2 fps however high video_fps is set. Encode clips at or above the rate you intend to ask for. - Clip length, when duration_seconds is not declared: mp4/mov/m4a and webm/mkv state their own duration and we read it exactly, inline or hosted; an avi, or anything we cannot read, is estimated from its byte size at a per-format bitrate, which reads short for a clip encoded below it and long for one above it; a clip we cannot size at all is charged and gated at the model's full duration ceiling. Declaring duration_seconds skips all of it. - A row over a duration cap or over its row context is rejected and never billed: 400 at submit when it is one of the rows sampled there (a clip we cannot size is assumed worst case), otherwise the job fails before approval with failure_code invalid_input naming the row. A clip too long at the job's rate reads "Row 'clip-7': video is 45 s; at the requested 20 fps the longest clip qwen3.8-27b takes is 38 s." - Media formats: audio wav/mp3/m4a/aac/ogg/opus/flac/aiff; video mp4/webm/mov/mkv/avi; image png/jpeg/webp/gif/bmp/tiff. Anything else is rejected before the job is priced. - Realtime: at most 8 requests in flight per organization across both realtime endpoints; a ninth answers 429 with Retry-After: 1. Embedding requests: 128 inputs, 262,144 characters per input, 2,097,152 across the request - Decision requests (laya), a batch row and a realtime request alike: at most 64 questions, a state of at most 50,000 characters, and a 192-token budget shared by one question's option texts (fewer options read better). Every question needs instructions; a noul's criteria may only be keyed true/false. A row over the first two caps fails the job with failure_code invalid_input (400 in realtime); a question the model refuses fails that row alone as a status-400 invalid_request_error (422 in realtime). Billed input tokens are the state once per question plus each question's words; output tokens are always 0. Batch input is JSONL only (parquet 422), and max_output_tokens, reasoning, reasoning_budget, response_format, video_fps and dimensions are refused (422) ## Errors Every error the API answers with carries a `detail` field; two bodies come from outside its routes and carry an `error` object instead: the maintenance-freeze 503 ({"error": "frozen", "message": ...}) and the edge rate-limit 429 ({"error": {"code": "rate_limit_exceeded", ...}}). Three call for an action rather than a report: 402, the balance does not cover the estimate - top up and approve again; 409, the job is not in the state the call needs; 429, wait the seconds in the Retry-After header and retry. Any 5xx is transient: retry with backoff, and note that approve and cancel are guarded, so a retry never double-charges. A job that fails after it was accepted carries a `failure_code` instead; `invalid_input` names the offending line. Every status is listed under Optional below, and in full on [/docs/api#errors](/docs/api#errors). Worked designs and cost arithmetic for pipelines built on this API are on the blog: index at [/blog](/blog), Atom feed at [/blog/feed.xml](/blog/feed.xml), posts listed under Optional below. Support: [/support](/support). Terms, including refunds, cancellation and the delivery period: [/terms](/terms). Uploaded datasets and job results are never used to train models; what is stored and for how long is on [/privacy](/privacy). ## Optional Background. Nothing above depends on it, so skip this section on a tight context budget. - [Qwen](/models/qwen), [Gemma](/models/gemma), [GLM](/models/glm), [Nemotron](/models/nemotron), [Laya](/models/laya): one page per weights family, with the benchmark scores the publisher reports - [LLM terms explained](/blog/llm-terms-explained): tokens, context, embeddings and quantization in plain language, for a reader with no technical background - [Self-hosting LLMs: When does it make sense?](/blog/self-hosting-llms): when running your own GPU beats a public API, with a worked cost example, the token arithmetic of coding agents, and a GSO benchmark run on a self-hosted 27B model - [How the Coyote vs Acme audience chose hope over burnout, TikTok marketing analysis with anex.sh](/blog/coyote-vs-acme-on-tiktok): what a batch of open models found in 2,264 TikTok posts about one film, with video read at 1 fps and again at 5 fps, and what the whole study cost - [How to spend less on tokens](/blog/spend-less-on-tokens): eight ways to cut an LLM bill, in the order to try them, with what a million images, 100,000 videos and 10M text tokens cost on Sonnet 5 against qwen3.8-27b here Every status the API answers with: - 400: request or input file failed validation (bad URL, over a cap, a parquet input missing custom_id or input, a malformed response_format, a rejected media part) - 401: the API key is invalid or no longer active - 402: the balance does not cover the estimate - 404: unknown id, or one belonging to another organization - 409: wrong state - approving a job not awaiting approval, cancelling one past stopping, submitting an input that is not ready yet - 410: the result data has been deleted (results are kept 30 days) - 413: the ingest source advertises a size over the input cap - 422: the request does not match the endpoint's shape, including an unknown model id, a video_fps above the model's ceiling ("video_fps 25 is above the 20 fps ceiling of qwen3.8-27b"), a video_fps on a model that takes none ("model 'gemma4-12b' does not support video_fps"), a parquet input or a generation field on a decision model, or, on POST /realtime/v1/systemone, a question the decision model refuses (the message names the question) - 429: too many jobs awaiting approval, too many concurrent realtime requests, or requests arriving faster than the API accepts (body error.code rate_limit_exceeded) - 502: the realtime decision service failed to answer; retry - 503: a check the request depends on could not run, or the API is frozen for maintenance (downloading a result keeps working)