How to spend less on tokens

Save money by following best practices battle tested in production.

The short version

We have been using LLMs for years and built several production systems on top of them. Here are the ways that stayed relevant. The best savings methods are switching to open models, using batch APIs and sending relevant data. In total we list eight ways to reduce your AI token costs, ranked by how we would approach them. The first few methods take very little time and preserve your output quality, while the later ones require more complex setups.

  1. Before you start optimizing
  2. Easy savings with immediate gain
  3. More complex saving options
  4. What to watch out for when using aggregators

Most bloated bills stem from ignoring the first two groups.

Before you start optimizing

Tracking detailed token usage and avoiding unnecessary AI calls is good practice at any point during development.

Track token usage per task

Looking at a massive monthly bill will not help you fix it; you need to track the exact cost of completing one specific task, like classifying a single support ticket. Knowing the price of a single action tells you exactly what variables drive your costs up or down.

Split your bill into inputs and outputs. Output tokens are more expensive than input tokens. E.g. on anex.sh output tokens cost 5x as much as an input token, and 20x a cached one. If your bill just says "tokens," you are likely paying almost entirely for generated text. You can check the token counts and price estimate in our price calculator.

Do not use AI for everything

The absolute best way to save money is to stop using AI for things that normal code can do for free. Any time you are developing a system that uses AI, you should try to call LLMs as little as possible. Classic software engineering brings many advantages like determinism, explainability and no API costs. A couple of examples where calling AI is not necessary:

Routing data
If your database already knows the user's country, use an if statement to route the data. Do not ask an AI to read the profile and guess.
Validating formats
Checking if a string is a valid date or a proper email address is a job for standard code. Using a model to double check another model's work simply doubles your bill.
Handling the obvious cases
Sift your data for easy wins first. Use simple lookups or regular expressions for exact matches, then only send the confusing leftovers to the AI.
Sorting and filtering
If you must use AI for e.g. recommendations or AI search, use embedding models, which cost a tiny fraction of generative models. And even then it is best practice to use LLMs to enrich your data with filters that you can use in a standard database and use LLMs only for the final re-ranking of top candidates.

Easy savings with immediate gain

Start with these simple but impactful changes that don't require using new models or changing your workflow.

Use batch processing

Run non-urgent tasks through a background batch system instead of demanding instant answers to save a massive amount of money. Realtime systems charge a premium to keep servers warm and ready, while batch jobs utilize the servers to 100% reducing the price.

Batch processing is perfect for costly data labeling or running massive evaluations. When picking the Batch API provider make sure to check both their prices and delivery speed. Some providers take days to return the result which is not ideal, others have dynamic pricing that charges you more than you would expect. Another common issue is rate limits that differ from provider to provider.

Here is what three typical jobs cost on Claude Sonnet 5 at Anthropic's list prices and on similarly intelligent Qwen 3.8 27B at anex.sh, with reasoning off (priced 18 September 2026):

jobSonnet 5 realtimeSonnet 5 batchQwen 3.8 27B on anex.shSonnet vs anex.sh (realtime / batch)
1M images$2,150$1,075$3526.1x / 3.1x
100k 60-second videos$3,070$1,535$23113.3x / 6.7x
10M text tokens$30$15$2.7011.1x / 5.6x

Each image is resized to about 400 px. Both including a 100-token prompt, and 150 output tokens. Sonnet 5 does not take video, so each video goes in as 60 frames at one per second. The text job is 10,000 rows of 1,000 input and 100 output tokens.

You also get to maximize your prompt caching discounts since all the requests are bundled together. The only tradeoff is time. You will wait anywhere from a few minutes to a few hours for the results, but the savings are well worth it for offline tasks. At anex.sh we typically deliver results in less than 1h and have upfront pricing without dynamic surcharges. We also run the jobs on dedicated machines which ensures that there is no data leakage to other customers.

Cache your instructions

Put your repeating instructions at the very beginning of your prompt to get massive discounts on your input costs. Systems remember text they have just seen and charge you significantly less to process it the second time around. On anex.sh a prefix already seen in the same job bills at a quarter of the input rate, automatically.

This discount relies on three strict rules. First, the text must be perfectly identical. Adding a unique timestamp or an extra space will break the match and cost you full price. Second, the shared instructions must be at the very top of the prompt before any unique data appears. Third, you must use the exact same model on the exact same server.

In our worked example of classifying a 100,000-product catalog, a 1,000-token instruction prefix is 100M of the job's 140M input tokens: $4.50 cached, against $18.00 at the normal input rate.

If you use an aggregator service, you might hit a snag. Switching models or landing on a different provider means the server has no memory of your prompt, forcing a cold start at full price. OpenRouter documents sticky routing that holds a conversation on one provider for ten minutes of activity, and cache discounts that differ per provider, from 0.1x to 0.5x of the input rate. Keep this in mind when comparing models, since the initial test runs will look more expensive than the cached production runs.

Trim the payload

Clean up your data and remove all the useless text before you send it to the AI. Every word you send costs money, and stripping out boilerplate formatting is the easiest way to lower your input costs while actually improving the model's focus.

Remove the junk
Strip out website menus, legal disclaimers, and hidden code blobs. Just send the raw text.
Send only what matters
If the model only needs to see the changes in a document, just send the edits instead of the entire file.
Summarize chat histories
In long conversations, you pay to re-read the entire history on every single reply. Generate a short summary of older messages and drop it below your cached instructions to save space.
Remove duplicates
Hash your inputs and run identical requests only once.
Resize your images
Big images consume far more tokens. Shrink them to the minimum size the AI requires before uploading them. Across the models we serve, one image ranges from a fixed 280 tokens up to 3,328.

Cap the output

Generating text is the most expensive part of any AI request, so you must strictly limit how much the model is allowed to talk. Cutting down on unnecessary thinking steps and forcing short, structured answers will drastically shrink your bill.

  1. Disable reasoning if you ignore it. Tasks like basic tagging or routing do not need a lengthy explanation. Turn the reasoning feature off completely.
  2. Set a thinking budget. If you do need reasoning, cap it. Setting a limit ensures the model stops rambling and provides an answer before it burns through your funds.
  3. Limit the final answer. Always set a hard ceiling on the total response length. Without this, a model might get stuck in a loop and drain your account.

On anex.sh those three controls are "reasoning": false, reasoning_budget and max_output_tokens. Classifying a million support tickets on qwen3-14b at 300 input and 10 output tokens per row costs about $9; leaving reasoning on for 500 tokens a row takes the same job, with the same answers, to about $71.

Finally, stop asking for conversational prose. If a computer script is going to read the AI's output, force the model to reply in strict JSON format.

More complex saving options

Switching to a different model can bring an order of magnitude cost reduction but requires evaluating and choosing the right model for your use case.

Switch to open models

The flagship models from OpenAI, Anthropic and Google cost significantly more than the open models. The cost of open models is lower because it's based on the hardware needed to run them. For example, Qwen 3.8 27B that costs 0.9 USD per 1M output tokens has the same intelligence score as Sonnet 5 (xhigh) from Anthropic that costs 5 USD when using Batch API pricing.

Most production work is classification, extraction, tagging, summarization and similar. Open models handle these tasks without any issues in dozens of languages. Unless you are running very complex analyses that require state of the art intelligence, you will find the new open models like Qwen 3.8 27B perfectly capable.

Use smaller models for simpler tasks

Switching to a smaller model can slash your costs to a fraction of your current bill, provided you test it thoroughly to ensure the quality remains acceptable. Smaller models are incredibly cheap, but you must verify they can actually handle your specific workflow.

The price difference can be orders of magnitude. Moving from a massive reasoning model to a compact 14 billion parameter model can drop your input costs by a factor of 28.

Our list prices (updated 18 September 2026):

modelinput / 1Moutput / 1Mbest for
qwen3-14b$0.025$0.125simple classification
gemma4-12b$0.12$0.60general and audio understanding
qwen3.8-27b$0.18$0.90Sonnet intelligence, video understanding
glm-5.2$0.70$2.20tasks requiring higher capability

To do this safely, you need an evaluation set. Pick a few hundred real examples from your own data and grade the correct answers manually. Run several cheaper models against this test set and pick the most affordable one that passes your quality bar.

If the cheap model only struggles with a few tricky cases, you can set up a cascading system. Let the cheap model handle the bulk of the easy work, and only escalate the failures to the expensive flagship model. Just make sure you use a reliable trigger to decide what gets escalated, like a missing required field.

Host the open models yourself

Renting your own AI servers is only a good idea if you have a massive, never ending stream of data that keeps the machines running constantly. It offers the biggest potential savings, but it is also the easiest way to waste money because you pay for the hardware every second it sits idle.

If your developers are writing code, testing logic, or taking a lunch break, that rented server is burning cash while doing absolutely nothing. Interactive human workflows rarely justify dedicated hardware. We wrote up the numbers, with a worked example at $3.50 an hour and a benchmark run on a self-hosted 27B model, in Self-hosting LLMs: When does it make sense?.

If you want the benefits of a rented server without the idle costs, go back to batch processing. A batch API effectively rents a machine for the exact duration of your job and immediately returns it when the work is done, letting someone else foot the bill for the downtime.

What to watch out for when using aggregators

Aggregators like OpenRouter are great for testing out different models but introduce some unique traps in production.

Hand-pick trusted providers

If you use a third party aggregator to access AI models, force them to use specific providers to avoid hidden quality drops or dangerous security risks. Letting the router automatically pick the cheapest server can result in terrible performance or expose your private data to malicious operators.

Some providers run heavily compressed (quantized) versions of open source models to save on computing power. These compressed versions perform noticeably worse, even though they share the exact same model name. OpenRouter exposes a quantizations filter in provider routing for exactly this reason, and some providers report their quantization as unknown. The Silent Hyperparameter (arXiv 2605.19537, May 2026) measured that the choice of inference backend alone shifts benchmark scores by up to 16.6 percentage points on the same weights. By locking in a specific provider, you ensure you get the exact quality level you tested. Every model we serve names its exact published checkpoint and precision in the models table, for the same reason.

You also need to watch your settings. Aggregators often ignore your strict cost controls if the chosen provider does not support them. On OpenRouter, allow_fallbacks defaults to true and require_parameters defaults to false, so a provider that drops your reasoning cap is still eligible to answer. More alarmingly, some shady providers actively harvest the plaintext data passing through their servers. Your Agent Is Mine (arXiv 2604.08407, April 2026) bought 28 paid routers and collected 400 free ones: 9 injected code into the responses they relayed, 17 reached for AWS credentials planted as bait, and one drained a wallet whose private key had passed through it. One of that paper's authors has since said he bought 6TB of invocation logs from an unnamed Chinese router, an unverified claim that no operator has answered. Treat any unverified route as a potential leak, and keep sensitive credentials on platforms where you have a legally binding contract.

Verify the price first

Never trust a catalog price blindly; always run a tiny test query and check the actual billed rate before processing massive datasets. Third party aggregators can have outdated pricing pages or apply hidden multipliers that will absolutely destroy your budget if you aren't paying attention.

We learned this the hard way when a model advertised at less than a dollar ended up costing three times as much in reality. On 1 August 2026, OpenRouter's /api/v1/models endpoint advertised that model at about $0.87 per million completion tokens, and the usage.cost on real calls worked out to about $3.00, roughly 3.4 times the listed price. Because our script generated massive amounts of text, that hidden multiplier spiked our projected costs well past our safety caps.

To protect yourself, implement three safety checks. First, run a micro test call and read the exact cost data it returns. Second, project that real cost across your entire dataset and refuse to run the job if it breaks your budget. Finally, build a live kill switch that checks your actual spending every few seconds and halts the script the moment you hit your spending limit.

This is why we never use dynamic pricing. Our jobs are also priced before they run and capped at the estimate you approve. You can check our price calculator to estimate the real cost for free.