Batch inference · open models · estimates up front
AI Batch API: Run massive jobs faster
LLM batch inference for open models like Qwen, Gemma and Nemotron. All modalities including audio, video, embeddings. Your job gets its own GPUs, so it isn't throttled or queued behind others, thousands of tokens per second, parallelized across the fleet. Anex prices the job before it runs, and an approved job never costs more than its estimate. 10–20× below Google and Anthropic.
Batch jobs save money and time
Processing large amounts of data with AI models is best done in batches. It is cheaper and faster than standard realtime APIs thanks to better HW utilization and our dedicated optimizations.
Thousands of tokens per second
Every job runs on GPUs dedicated to it, so it moves at the speed of the hardware instead of a shared queue. No rate limits to spread a large dataset over days.
Capped estimate with no dynamic pricing
We tokenize your dataset and estimate input, output, images, and audio/video minutes in credits before you commit. Per token prices are fixed, no dynamic pricing causing surprising bills.
No hidden quantization
We never serve a quantized model without clearly specifying it.
Pricing
You buy prepaid credits (1 credit = 1 US cent, USD) and jobs consume them. Estimates are free, and the credit amount shown at approval is a hard cap - an approved job never consumes more, and unused credits stay on your balance.
What will your job cost?
Enter a row that looks like your data and the number of rows you have. Tokens are counted with the model's own tokenizer and priced on the rate cards below.
Text output models
| model | input modality | input / 1M tokens | cached input / 1M | output / 1M tokens | price per image | price per clip |
|---|---|---|---|---|---|---|
| qwen3.8-27b | text | $0.18 | $0.045 | $0.90 | - | - |
| image + video | $0.30 | $0.075 | $1.20 | $0.0001 | $0.0004 | |
| gemma4-31b | text | $0.20 | $0.05 | $1.00 | - | - |
| image | $0.24 | $0.06 | $1.20 | $0.0002 | - | |
| gemma4-12b | text | $0.12 | $0.03 | $0.60 | - | - |
| image + audio + video | $0.30 | $0.075 | $1.50 | $0.0002 | $0.0002 | |
| qwen3-14b | text | $0.025 | $0.006 | $0.125 | - | - |
| nemotron-3-nano-omni-30b | text | $0.12 | $0.03 | $0.60 | - | - |
| image + audio + video | $0.30 | $0.075 | $1.50 | $0.0011 | $0.0002 |
Embedding output models
| model | embeds | text rows / 1M input tokens | media rows / 1M | per 1M images | per 1M video clips | realtime queries / 1M |
|---|---|---|---|---|---|---|
| qwen3-vl-embedding-8b | text + image + video | $0.03 | $0.08 | $4 | $60 | $0.121 |
| qwen3-vl-embedding-2b | text + image + video | $0.007 | $0.02 | $2 | $20 | $0.0077 |
| qwen3-embedding-8b | text | $0.03 | - | - | - | $0.033 |
Realtime embedding requests (see realtime embeddings) bill per input token at the realtime rate in the last column, from the same credit balance as batch jobs. Text queries only: image and video embedding stays on the batch API.
Decision output models
| model | answers | batch jobs / 1M input tokens | realtime queries / 1M | minimum job price |
|---|---|---|---|---|
| laya | choice, score, yes/no | $0.005 | $0.033 | $0.05 |
A decision model answers typed questions about a text (a choice, a score, a yes or no) with a probability per option instead of writing a reply, and bills input tokens only: the text is read once per question. See decision jobs and realtime decisions.
What jobs actually cost
| example job | model | assumed per row | estimate |
|---|---|---|---|
| Classify 1M support tickets | qwen3-14b | ≈300 tokens in, 10 out | ≈ $9 |
| Caption 100,000 photos | gemma4-31b | 1 image, ≈200 tokens in, 60 out | ≈ $32 |
| Transcribe 1,000 hours of audio | nemotron-3-nano-omni-30b | audio at 12.5 tokens/s, ≈12k transcript tokens out per hour | ≈ $32 |
| Describe 10,000 one-minute videos | nemotron-3-nano-omni-30b | video at 512 tokens/s, ≈500 tokens out | ≈ $101 |
Your job gets its own estimate based on current prices; these are representative examples. What the columns mean and what else shapes a bill is on the pricing page, and the price calculator prices a job with media, caching and the approval hold. The exact version (weights build) and quantization behind every model name here are on the model list.
The batch lifecycle
You can use our OpenAI compatible API or submit the jobs through your user dashboard.
Submit a dataset
Point the API at a JSONL file - upload it to us, or have us ingest it from any HTTPS URL.
Get the estimate
We tokenize your rows and measure real output length, then hold the job here with a priced estimate: input, output, and media credits. Nothing has run yet.
Approve it
Nothing runs until you (or your auto-approve rule) accept the estimate. This is the spend gate; the moment you approve, the job is in progress.
Collect results
The job runs on the lowest-cost hardware that fits it, and the merged result file, with per-row usage, is ready to download. A signed webhook tells you it's ready.
You can see all job statuses in the job lifecycle reference.
Submit, review the estimate, approve
Mint an API key in the dashboard. Estimates are free; approval is what starts paid compute.
# submit a batch job against qwen3.8-27b
curl -X POST https://api.anex.sh/batch/v1/jobs \
-H "Authorization: Bearer $ANEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input_s3_uri": "s3://.../in/ingest/.../input.jsonl",
"model": "qwen3.8-27b",
"max_output_tokens": 512,
"response_format": {"type": "json_object"}
}'
# poll until "status": "cost_ready", review est_cost_cents, then approve
curl -X POST https://api.anex.sh/batch/v1/jobs/<job-id>/approve \
-H "Authorization: Bearer $ANEX_API_KEY"About us
We are a small team of AI and infrastructure engineers (Tonda, Radek, Jan) based in Prague, with 30+ years of combined experience building production ML systems. We built anex.sh because we kept paying API prices for dataset work, and knew from running our own GPUs exactly how much cheaper the same tokens could be.
The GPU fleet itself runs on ANEX, our open-source engine for integrating GPU containers into Kubernetes clusters - the project this domain started as, and still free and open on GitHub.
When you write to support@anex.sh, you talk to the people who wrote the code.
FAQ
How are you so cheap?
We automatically analyze each job and then select the best way to run it. We optimize model serving extensively to achieve high GPU utilization, which makes our low prices possible.
Are the models modified in any way?
No. We run the publicly available weights with clearly specified quantization. So a model served at fp8 (8-bit floating point) runs that published fp8 checkpoint and an unquantized model runs the original published weights. Nothing is fine tuned, distilled, or quantized by us.
The chosen weights behind every model name are in the model list.
How quickly do you provide the results?
We offer a standard 24h SLA, but we are able to deliver the results much faster, even in less than 1h.
If you have specific needs for delivery speed guarantees per job, we are happy to discuss it - just message us at support@anex.sh.
How does the batch API differ from a standard real-time API?
Real-time APIs are built for use cases where you need the response immediately because the user is waiting for it: chatbots, copilots, interactive AI applications.
At the same time, they are expensive and have strict rate limits that slow down the processing of large amounts of data. This is exactly where a batch API becomes useful: it is both cheaper and faster at scale.
What are the use cases for the batch API?
The batch API is optimized to process large amounts of data efficiently. If you have any case where you would be repeatedly calling a standard real-time API for a list of inputs (text, audio, video), then the batch API is the perfect fit.
Some examples:
- Categorizing and tagging millions of products.
- Generating product descriptions.
- Summarizing news articles, videos, or audio.
- Call transcription or categorization for sales and customer support teams.
- Vectorizing product data to implement AI search in ecommerce.
- Generating video descriptions or vectorizing video segments for an AI video search.
- Evaluating your product on a large dataset or benchmark.
Can you serve a custom model?
Yes. Alongside the catalogue we can serve a model you bring: an open-weight model we do not list yet, or your own fine-tune of one. It runs for your organization only, through the same API and console as everything else.
Write to support@anex.sh with the model and roughly how much data you expect to run through it, and we will come back with a price and a date it can be serving.
Do credits expire, and how is tax handled?
Credits do not expire. Unused ones stay on your balance for the next job. They are prepaid and non-refundable, except where our Terms of Service say otherwise.
Credit prices exclude tax. Tax is worked out from the billing address you give at checkout and added to the payment, so the credits you receive always match the amount you entered. Where a reverse charge applies, a business that enters a valid tax ID is billed with no tax added and the invoice records that. Every paid top-up is invoiced, and the console lists your purchases with a PDF invoice for each one.
Can you keep our data inside the EEA?
Yes, on request. We turn EU data routing on for your organization, and every new upload and job after that keeps its content inside the EEA (European Economic Area): the files you upload, your prompts and the results are stored there, and the servers and GPU machines that process them are there too. Write to support@anex.sh to have it turned on.
Job details such as ids, status and row counts, your usage, billing and account stay in the US. Prices are the same. Data residency lists exactly what is covered and how to use it.
Can we delete our data right after a job finishes?
Yes. Once a job has finished, delete its input and output with one click in the console or one API call, so that we store none of your data longer than is necessary.
By default, inputs are deleted 7 days after upload and outputs 30 days after the job finishes to prevent accidentally losing your data and results. Only the job records (status, counts, cost) stay.
Do you train on our data?
No. We never use your datasets or your results to train models, and we do not sell any customer data. Your job data is scoped to your organization and is processed to run the job, bill it, and support you.
What we store, who processes it, and for how long is in our privacy policy.
Know the price. Then run it.
Create an organization, invite your team, and submit your first batch job. Estimates are free.
Create account