Qwen models on anex.sh

Alibaba's open-weight family, and the largest part of our catalogue.

The family

Qwen is Alibaba's open-weight line, published under Apache 2.0 and released as whole checkpoints rather than API access. It is the widest family in our catalogue: dense text and vision models in the size range batch work actually pays for, and a set of embedding towers that put several modalities into one vector space.

Models that answer in text

modelsizeinputquantizationweights served
qwen3.8-27b27B densetext + image + videofp8Qwen/Qwen3.8-27B-FP8
qwen3-14b14B densetextfp8Qwen/Qwen3-14B-FP8
qwen3.8-27b
A 27B dense checkpoint that reads text, images and video, the only chat model in this family that takes media. Its vision band caps an image at 200 tokens, which puts its published per-image ceiling at the bottom of the price table, and video is charged per sampled frame at one frame per second by default, or up to 20 when a job asks for it, on a clip of up to two minutes. It emits a reasoning trace, which a job can cap with a reasoning budget.
qwen3-14b
Text only, and the cheapest chat model we serve. A 14B dense reasoning model is enough for classification, extraction, routing and rewriting at volume, which is most of what a batch queue actually carries.

Models that answer with a vector

modelsizeinputnative dimensionscontext
qwen3-embedding-8b8Btext4096 (32-4096)32,768
qwen3-vl-embedding-8b8Btext + image + video4096 (64-4096)32,768
qwen3-vl-embedding-2b2Btext + image + video2048 (64-2048)32,768

The bracketed range is what you may ask for with the dimensions field; the first number is what the model returns if you ask for nothing. The two qwen3-vl-embedding-* models put text, images and video into one vector space, so a text query can retrieve a page image or a clip without a second index. See embedding jobs for the request shape and pricing for the rates.

Benchmark scores

Published figures, each from the source named beside it. They are not comparable across rows: different benchmarks, different harnesses, and in most cases the model's own publisher doing the measuring. Use them to tell these models apart from each other, not to rank them against something scored elsewhere.

modelbenchmarkscoresource
qwen3.8-27bSWE-bench Pro61.7model card, via Yotta Labs
qwen3.8-27bOSWorld-Verified84.3model card, via Yotta Labs
qwen3-14bGPQA Diamond (thinking)64.0Qwen3 report
qwen3-14bMMLU-Redux (thinking)88.6Qwen3 report
qwen3-14bLiveCodeBench v5 (thinking)63.5Qwen3 report
qwen3-14bAIME 2025 (thinking)70.4Qwen3 report
qwen3-embedding-8bMTEB multilingual70.58Qwen3 Embedding
qwen3-vl-embedding-8bMMEB-V2 overall77.8Qwen3-VL-Embedding report
qwen3-vl-embedding-2bMMEB-V2 overall73.2Qwen3-VL-Embedding report

Before you read too much into the table: the qwen3-14b figures are its thinking-mode scores. With thinking off the same model reads 54.8 on GPQA Diamond and 29.0 on LiveCodeBench, so which mode you run is a bigger lever than which of these models you pick.

Which one to pick

Simple classification, extraction, routing, rewriting
qwen3-14b. The cheapest per token and the cheapest minimum job price. If the task is more nuanced, or the rows carry images, go for qwen3.8-27b instead.
Documents, screenshots, product photos
qwen3.8-27b: it reads the page and its image encoding is the cheapest in the price table.
Video
qwen3.8-27b if the clip is silent. If the soundtrack carries meaning, two models outside this family read audio and video together: nemotron-3-nano-omni-30b and gemma4-12b.
Audio, transcription, call analytics
Nothing in this family hears. Two models outside it do: nemotron-3-nano-omni-30b for recordings up to 40 minutes, and gemma4-12b for clips under 30 seconds that need a stronger answer than a transcript.
Search and retrieval indexes
qwen3-embedding-8b for text alone; qwen3-vl-embedding-8b when pages, images or clips have to sit in the same index as the text; qwen3-vl-embedding-2b when the corpus is large enough that a quarter of the price matters more than the last few points of recall.

Pass the model id exactly as it is written above when you submit a job. Rates for every model are on the pricing page, and every job is estimated and priced before it runs, so trying a second model costs you an estimate rather than a bill.