Batch inference · open models · estimates up front

AI Batch API: Run massive jobs faster

LLM batch inference for open models like Qwen, Gemma and Nemotron. All modalities including audio, video, embeddings. Your job gets its own GPUs, so it isn't throttled or queued behind others, thousands of tokens per second, parallelized across the fleet. Anex prices the job before it runs, and an approved job never costs more than its estimate. 10–20× below Google and Anthropic.

Batch job estimate job 8f21c6…9e4a
modelqwen3.8-27b
input48,210 rows · jsonl
est. input tokens19,441,208
est. output tokens6,102,400
estimated credits899 ($8.99)
submittedestimatingcost_readyin_progressfinalizingdone
The job holds at cost_ready until you approve. Nothing has been spent yet.

Batch jobs save money and time

Processing large amounts of data with AI models is best done in batches. It is cheaper and faster than standard realtime APIs thanks to better HW utilization and our dedicated optimizations.

speed

Thousands of tokens per second

Every job runs on GPUs dedicated to it, so it moves at the speed of the hardware instead of a shared queue. No rate limits to spread a large dataset over days.

certainty

Capped estimate with no dynamic pricing

We tokenize your dataset and estimate input, output, images, and audio/video minutes in credits before you commit. Per token prices are fixed, no dynamic pricing causing surprising bills.

transparency

No hidden quantization

We never serve a quantized model without clearly specifying it.

Pricing

You buy prepaid credits (1 credit = 1 US cent, USD) and jobs consume them. Estimates are free, and the credit amount shown at approval is a hard cap - an approved job never consumes more, and unused credits stay on your balance.

What will your job cost?

Enter a row that looks like your data and the number of rows you have. Tokens are counted with the model's own tokenizer and priced on the rate cards below.

Text output models

modelinput modalityinput
/ 1M tokens
cached input
/ 1M
output
/ 1M tokens
price
per image
price
per clip
qwen3.8-27btext$0.18$0.045$0.90--
image + video$0.30$0.075$1.20$0.0001$0.0004
gemma4-31btext$0.20$0.05$1.00--
image$0.24$0.06$1.20$0.0002-
gemma4-12btext$0.12$0.03$0.60--
image + audio + video$0.30$0.075$1.50$0.0002$0.0002
qwen3-14btext$0.025$0.006$0.125--
nemotron-3-nano-omni-30btext$0.12$0.03$0.60--
image + audio + video$0.30$0.075$1.50$0.0011$0.0002

Embedding output models

modelembedstext rows
/ 1M input tokens
media rows
/ 1M
per 1M
images
per 1M
video clips
realtime
queries / 1M
qwen3-vl-embedding-8btext + image + video$0.03$0.08$4$60$0.121
qwen3-vl-embedding-2btext + image + video$0.007$0.02$2$20$0.0077
qwen3-embedding-8btext$0.03---$0.033

Realtime embedding requests (see realtime embeddings) bill per input token at the realtime rate in the last column, from the same credit balance as batch jobs. Text queries only: image and video embedding stays on the batch API.

Decision output models

modelanswersbatch jobs
/ 1M input tokens
realtime
queries / 1M
minimum
job price
layachoice, score, yes/no$0.005$0.033$0.05

A decision model answers typed questions about a text (a choice, a score, a yes or no) with a probability per option instead of writing a reply, and bills input tokens only: the text is read once per question. See decision jobs and realtime decisions.

What jobs actually cost

example jobmodelassumed per rowestimate
Classify 1M support ticketsqwen3-14b≈300 tokens in, 10 out≈ $9
Caption 100,000 photosgemma4-31b1 image, ≈200 tokens in, 60 out≈ $32
Transcribe 1,000 hours of audionemotron-3-nano-omni-30baudio at 12.5 tokens/s, ≈12k transcript tokens out per hour≈ $32
Describe 10,000 one-minute videosnemotron-3-nano-omni-30bvideo at 512 tokens/s, ≈500 tokens out≈ $101

Your job gets its own estimate based on current prices; these are representative examples. What the columns mean and what else shapes a bill is on the pricing page, and the price calculator prices a job with media, caching and the approval hold. The exact version (weights build) and quantization behind every model name here are on the model list.

The batch lifecycle

You can use our OpenAI compatible API or submit the jobs through your user dashboard.

01 · submitted

Submit a dataset

Point the API at a JSONL file - upload it to us, or have us ingest it from any HTTPS URL.

02 · cost_ready

Get the estimate

We tokenize your rows and measure real output length, then hold the job here with a priced estimate: input, output, and media credits. Nothing has run yet.

03 · in_progress

Approve it

Nothing runs until you (or your auto-approve rule) accept the estimate. This is the spend gate; the moment you approve, the job is in progress.

04 · done

Collect results

The job runs on the lowest-cost hardware that fits it, and the merged result file, with per-row usage, is ready to download. A signed webhook tells you it's ready.

You can see all job statuses in the job lifecycle reference.

Submit, review the estimate, approve

Mint an API key in the dashboard. Estimates are free; approval is what starts paid compute.

# submit a batch job against qwen3.8-27b curl -X POST https://api.anex.sh/batch/v1/jobs \ -H "Authorization: Bearer $ANEX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "input_s3_uri": "s3://.../in/ingest/.../input.jsonl", "model": "qwen3.8-27b", "max_output_tokens": 512, "response_format": {"type": "json_object"} }' # poll until "status": "cost_ready", review est_cost_cents, then approve curl -X POST https://api.anex.sh/batch/v1/jobs/<job-id>/approve \ -H "Authorization: Bearer $ANEX_API_KEY"

About us

We are a small team of AI and infrastructure engineers (Tonda, Radek, Jan) based in Prague, with 30+ years of combined experience building production ML systems. We built anex.sh because we kept paying API prices for dataset work, and knew from running our own GPUs exactly how much cheaper the same tokens could be.

The GPU fleet itself runs on ANEX, our open-source engine for integrating GPU containers into Kubernetes clusters - the project this domain started as, and still free and open on GitHub.

When you write to support@anex.sh, you talk to the people who wrote the code.

The anex.sh team

FAQ

How are you so cheap?

We automatically analyze each job and then select the best way to run it. We optimize model serving extensively to achieve high GPU utilization, which makes our low prices possible.

Are the models modified in any way?

No. We run the publicly available weights with clearly specified quantization. So a model served at fp8 (8-bit floating point) runs that published fp8 checkpoint and an unquantized model runs the original published weights. Nothing is fine tuned, distilled, or quantized by us.

The chosen weights behind every model name are in the model list.

How quickly do you provide the results?

We offer a standard 24h SLA, but we are able to deliver the results much faster, even in less than 1h.

If you have specific needs for delivery speed guarantees per job, we are happy to discuss it - just message us at support@anex.sh.

How does the batch API differ from a standard real-time API?

Real-time APIs are built for use cases where you need the response immediately because the user is waiting for it: chatbots, copilots, interactive AI applications.

At the same time, they are expensive and have strict rate limits that slow down the processing of large amounts of data. This is exactly where a batch API becomes useful: it is both cheaper and faster at scale.

What are the use cases for the batch API?

The batch API is optimized to process large amounts of data efficiently. If you have any case where you would be repeatedly calling a standard real-time API for a list of inputs (text, audio, video), then the batch API is the perfect fit.

Some examples:

  • Categorizing and tagging millions of products.
  • Generating product descriptions.
  • Summarizing news articles, videos, or audio.
  • Call transcription or categorization for sales and customer support teams.
  • Vectorizing product data to implement AI search in ecommerce.
  • Generating video descriptions or vectorizing video segments for an AI video search.
  • Evaluating your product on a large dataset or benchmark.
Can you serve a custom model?

Yes. Alongside the catalogue we can serve a model you bring: an open-weight model we do not list yet, or your own fine-tune of one. It runs for your organization only, through the same API and console as everything else.

Write to support@anex.sh with the model and roughly how much data you expect to run through it, and we will come back with a price and a date it can be serving.

Do credits expire, and how is tax handled?

Credits do not expire. Unused ones stay on your balance for the next job. They are prepaid and non-refundable, except where our Terms of Service say otherwise.

Credit prices exclude tax. Tax is worked out from the billing address you give at checkout and added to the payment, so the credits you receive always match the amount you entered. Where a reverse charge applies, a business that enters a valid tax ID is billed with no tax added and the invoice records that. Every paid top-up is invoiced, and the console lists your purchases with a PDF invoice for each one.

Can you keep our data inside the EEA?

Yes, on request. We turn EU data routing on for your organization, and every new upload and job after that keeps its content inside the EEA (European Economic Area): the files you upload, your prompts and the results are stored there, and the servers and GPU machines that process them are there too. Write to support@anex.sh to have it turned on.

Job details such as ids, status and row counts, your usage, billing and account stay in the US. Prices are the same. Data residency lists exactly what is covered and how to use it.

Can we delete our data right after a job finishes?

Yes. Once a job has finished, delete its input and output with one click in the console or one API call, so that we store none of your data longer than is necessary.

By default, inputs are deleted 7 days after upload and outputs 30 days after the job finishes to prevent accidentally losing your data and results. Only the job records (status, counts, cost) stay.

Do you train on our data?

No. We never use your datasets or your results to train models, and we do not sell any customer data. Your job data is scoped to your organization and is processed to run the job, bill it, and support you.

What we store, who processes it, and for how long is in our privacy policy.

Know the price. Then run it.

Create an organization, invite your team, and submit your first batch job. Estimates are free.

Create account