The family
Qwen is Alibaba's open-weight line, published under Apache 2.0 and released as whole checkpoints rather than API access. It is the widest family in our catalogue: dense text and vision models in the size range batch work actually pays for, and a set of embedding towers that put several modalities into one vector space.
Models that answer in text
| model | size | input | quantization | weights served |
|---|---|---|---|---|
| qwen3.8-27b | 27B dense | text + image + video | fp8 | Qwen/Qwen3.8-27B-FP8 |
| qwen3-14b | 14B dense | text | fp8 | Qwen/Qwen3-14B-FP8 |
- qwen3.8-27b
- A 27B dense checkpoint that reads text, images and video, the only chat model in this family that takes media. Its vision band caps an image at 200 tokens, which puts its published per-image ceiling at the bottom of the price table, and video is charged per sampled frame at one frame per second by default, or up to 20 when a job asks for it, on a clip of up to two minutes. It emits a reasoning trace, which a job can cap with a reasoning budget.
- qwen3-14b
- Text only, and the cheapest chat model we serve. A 14B dense reasoning model is enough for classification, extraction, routing and rewriting at volume, which is most of what a batch queue actually carries.
Models that answer with a vector
| model | size | input | native dimensions | context |
|---|---|---|---|---|
| qwen3-embedding-8b | 8B | text | 4096 (32-4096) | 32,768 |
| qwen3-vl-embedding-8b | 8B | text + image + video | 4096 (64-4096) | 32,768 |
| qwen3-vl-embedding-2b | 2B | text + image + video | 2048 (64-2048) | 32,768 |
The bracketed range is what you may ask for with the
dimensions field; the first number is what the model
returns if you ask for nothing. The two
qwen3-vl-embedding-* models put text, images and video
into one vector space, so a text query can retrieve a page image or a
clip without a second index. See
embedding jobs for the request shape
and pricing for the rates.
Benchmark scores
Published figures, each from the source named beside it. They are not comparable across rows: different benchmarks, different harnesses, and in most cases the model's own publisher doing the measuring. Use them to tell these models apart from each other, not to rank them against something scored elsewhere.
| model | benchmark | score | source |
|---|---|---|---|
| qwen3.8-27b | SWE-bench Pro | 61.7 | model card, via Yotta Labs |
| qwen3.8-27b | OSWorld-Verified | 84.3 | model card, via Yotta Labs |
| qwen3-14b | GPQA Diamond (thinking) | 64.0 | Qwen3 report |
| qwen3-14b | MMLU-Redux (thinking) | 88.6 | Qwen3 report |
| qwen3-14b | LiveCodeBench v5 (thinking) | 63.5 | Qwen3 report |
| qwen3-14b | AIME 2025 (thinking) | 70.4 | Qwen3 report |
| qwen3-embedding-8b | MTEB multilingual | 70.58 | Qwen3 Embedding |
| qwen3-vl-embedding-8b | MMEB-V2 overall | 77.8 | Qwen3-VL-Embedding report |
| qwen3-vl-embedding-2b | MMEB-V2 overall | 73.2 | Qwen3-VL-Embedding report |
Before you read too much into the table: the qwen3-14b figures are its thinking-mode scores. With thinking off the same model reads 54.8 on GPQA Diamond and 29.0 on LiveCodeBench, so which mode you run is a bigger lever than which of these models you pick.
Which one to pick
- Simple classification, extraction, routing, rewriting
- qwen3-14b. The cheapest per token and the cheapest minimum job price. If the task is more nuanced, or the rows carry images, go for qwen3.8-27b instead.
- Documents, screenshots, product photos
- qwen3.8-27b: it reads the page and its image encoding is the cheapest in the price table.
- Video
- qwen3.8-27b if the clip is silent. If the soundtrack carries meaning, two models outside this family read audio and video together: nemotron-3-nano-omni-30b and gemma4-12b.
- Audio, transcription, call analytics
- Nothing in this family hears. Two models outside it do: nemotron-3-nano-omni-30b for recordings up to 40 minutes, and gemma4-12b for clips under 30 seconds that need a stronger answer than a transcript.
- Search and retrieval indexes
- qwen3-embedding-8b for text alone; qwen3-vl-embedding-8b when pages, images or clips have to sit in the same index as the text; qwen3-vl-embedding-2b when the corpus is large enough that a quarter of the price matters more than the last few points of recall.
Pass the model id exactly as it is written above when you submit a job. Rates for every model are on the pricing page, and every job is estimated and priced before it runs, so trying a second model costs you an estimate rather than a bill.