anex.sh documentation

What the service does, how a job flows, and where to go next.

What anex.sh is

anex.sh runs an open language model over a whole dataset as one job, instead of one request at a time. You give it a file with one row per prompt, choose a model, and get back one result file with a response per row. Rows are text, and on a media-capable model they can also carry images, audio clips, and video. Two kinds of model return something other than text: the embedding models return a vector per row, and the decision model answers typed questions about each row (a choice, a score, a yes or no) with a probability per option.

Before anything runs, the job is estimated and held at an approval gate. Estimates are free, and an approved job never consumes more credits than the estimate you approved. Payment is prepaid credits, not an invoice after the fact.

It fits work that arrives as a file of rows rather than as live chat traffic - anything you would otherwise run as a long loop of per-row API calls.

Two ways to use it

Browser console
Sign in and submit, approve, track, and download jobs from the browser. Nothing to install. See the console guide.
HTTP API
Call the batch API with an API key from a script or from CI. It can also ingest an input file from a URL, which the console cannot. See the API reference.

Both work on the same jobs and the same organization: a job submitted from a script shows up in the console, and a job submitted in the console can be polled with an API key.

Where to start

  1. Create an account. The first person to sign in for a new organization becomes its owner. The owner can then invite teammates into the same organization and promote them - up to owner - rather than each person creating an organization of their own; roles and invitations are covered in the console guide.
  2. Buy credits. The service is prepaid, and an owner buys them from the Settings page; what a job costs is on the pricing page.
  3. Submit a job from the Jobs page: an input file and a model.
  4. Approve the estimate when the job reaches cost_ready, or turn on auto-approve so that happens for you.
  5. Preview and download the result once the job is done.

From a script, the same path is five calls: POST /inputs to upload or ingest the file, POST /jobs to submit, GET /jobs/{id} until the job is cost_ready, POST /jobs/{id}/approve to start the paid work, and GET /jobs/{id}/result once it is done. The API reference covers each one.

Job lifecycle

A job reports one status at a time, the same value in the console and in the API. The normal path runs top to bottom:

submitted
Just created. The input has passed its format, size, and row-count checks and is queued for costing.
estimating
The output length is being measured from a sample of your rows, so the job can be priced. A job that was submitted with an explicit max tokens per row is priced from that number instead and can go straight to cost_ready.
cost_ready
Estimated and waiting for approval - yours, or an auto-approve rule. No credits are held yet.
in_progress
Approved and running. The estimated cost is held against your balance and the job is run as a batch. Batches run at high throughput but take tens of minutes to come back: a machine is rented and the model loaded onto it before any row is processed, and on a short job that setup is most of the elapsed time. A job that sits here for a while is still on track, so give it time before you cancel and resubmit. Row counts are reported once the job has finished, not while it runs.
finalizing
GPU work is done; the per-row outputs are being combined into one result file. Too late to cancel from here on.
done
Finished. The result can be previewed and downloaded.

Off that main line, a job can also end up:

cancelling
A cancel was requested after work had started; rows already in flight are finishing, then the job stops.
cancelled
Stopped by a cancel. If nothing had started, no charge. If rows had already run, those are delivered and billed and the rest is dropped.
failed
Something went wrong; the reason is reported with the status. If some rows had already completed, they are still available as a partial result.

Limits

Input file size
1 GiB per input.
Rows per job
10,000,000.
Jobs awaiting approval
10 per organization across submitted and cost_ready; submit answers 429 past the cap.
Approval window
24 hours in cost_ready, then the job is cancelled automatically, free of charge.
Inline images
20 MB per image, for JSONL rows that embed an image directly in the request.
Inline audio
Per model, for an input_audio part or a data: URI in audio_url, in line with the clip lengths each one takes: nemotron-3-nano-omni-30b allows 90 MB, gemma4-12b 20 MB.
Inline video
150 MB per clip, for a data: URI in video_url.
Audio formats
wav, mp3, m4a, aac, ogg, opus, flac, aiff. Anything else is rejected at submit.
Video formats
mp4, webm, mov, mkv, avi.
Image formats
png, jpeg, webp, gif, bmp, tiff.
Media duration and row context
Per model, in the tables below. Every row is served in a 32,768-token context, except a row carrying images or video on qwen3.8-27b, which may take up to 131,072 input tokens. Each second of audio, each sampled video frame and each image costs a fixed number of those tokens.
Media items per row
Up to 2 images; 1 audio clip and 1 video clip, on the chat-completions models that take them. An embedding model that takes media (qwen3-vl-embedding-8b, qwen3-vl-embedding-2b - see embedding jobs) allows up to 4 images and 1 video clip instead, and never audio. A row carrying more than its model allows is not rejected up front - the surplus is refused at serving time and can fail that row.
Shortest video clip
A clip has to hold at least 2 frames. A one-frame video (a sub-second scene re-encoded at 1 frame per second, for example) is refused before any paid work starts, with the row named; send that frame as an image part instead.
Max tokens per row
32,768 - the highest per-row output cap you can request when you set max tokens yourself instead of letting the platform estimate it. A submit with a larger value is rejected.
modelaudio per clipvideo per clipimagerow context
gemma4-12b30 s, at 25 tokens per second60 s, 1 frame per second at 74 tokens a frame280 tokens32,768
nemotron-3-nano-omni-30b2,400 s (40 min), at 12.5 tokens per second60 s, 2 frames per second at 256 tokens a frame256 tokens per 512 px tile, up to 13 tiles (3,328)32,768
qwen3.8-27bnot accepted120 s, 1 frame per second at 105 tokens a frame, at most 768 framesup to 200 tokens32,768; 131,072 input with images or video
gemma4-31bnot acceptednot accepted280 tokens32,768

Video is sampled at the model's default rate unless the job sets video_fps. Every sampled frame is processed at the model's full per-frame resolution budget and costs the same tokens at any rate, so a higher rate costs more per second of clip and fits a shorter clip. The longest clip each model takes:

modeldefault rateat the default5 fps10 fps20 fps
qwen3.8-27b1 fps120 s120 s76 s38 s
nemotron-3-nano-omni-30b2 fps60 s25 s12 s6 s
gemma4-12b1 fps60 snot yet selectable

20 fps is the highest rate either model takes; a higher rate, or a rate on a model that takes none, is refused at submit. The embedding models always sample video at 1 frame per second.

The row context is the whole budget for a row: at most 32,768 input tokens including everything the media encodes to, at most 32,768 output tokens, and on the cards these models run on today the input and the output together have to fit inside 32,768 as well. That is what narrows the long clips: 40 minutes of audio on nemotron-3-nano-omni-30b is 30,000 tokens, leaving room for a prompt and a short answer, and 60 s of video on it is about 31,000. A row carrying images or video on qwen3.8-27b is the exception: up to 131,072 input tokens, at most 32,768 output tokens, and both together within 131,072. qwen3-14b takes text only and is served in the same 32,768 tokens; the embedding models carry their own context, listed on the API reference.

A row over any of these limits is rejected and nothing is billed for it. If it is one of the rows sampled at submit - a declared duration, an inline clip, or a hosted URL with no duration_seconds, which is assumed to be worst case - submit answers 400 with the reason. Otherwise the job fails before it can be approved, with failure_code: invalid_input and a failure_reason naming the row.

These are enforced on the input itself, whichever way the job was submitted. A file that clears the browser's own check can still be rejected once the server reads it. The image/audio/video sizes cap the bytes accepted inline per media item; a remote URL is never size-checked - its fetch is bounded by a timeout at serving time instead. The inline caps are checked on a sample of the file's first rows at submit, then on every row while the job is estimating: a violation past the sampled rows fails the job with failure_code: invalid_input and the same message the submit-time rejection uses. A model with no encoder for a modality (see the API reference) isn't caught by the size check - it fails the job later instead. An embedding model's image and video parts are checked the same way; an embedding model never accepts audio, so an audio part on one is rejected as an unsupported modality rather than a size violation. A media part's format is read from what the row already carries, in order: the format field on input_audio, a data: URI's media type, or the file extension on a URL. We also check the first bytes of anything sent inline, so a clip whose contents contradict what it claims is rejected rather than failing on a GPU later. An audio or video URL with no file extension is resolved from the Content-Type its host returns on a HEAD request made while the job is being estimated; when a host answers with application/octet-stream, or names something off the list, that leaves us nothing to go on and the job is rejected with failure_code: invalid_input. The probe covers the first few dozen URLs in a job, so a clip we could not ask about is admitted and can still fail at serving time. Give the URL a file extension, serve it with a real content type, or send the bytes inline. Images are never probed this way, so an extensionless image URL with no data: media type passes through to the model as-is.