Gemma models on anex.sh

Google's open-weight family, one of the few open western models.

The family

Gemma is Google's open-weight line, built from the research behind Gemini and published as downloadable checkpoints rather than API access. Where Qwen ships many checkpoints, Gemma ships fewer and larger dense models, and puts the effort into what a single checkpoint does well.

Two Gemma models are on offer here: the multimodal 31B instruct checkpoint and the 12B, a third of the size, which reads audio and video on top of images and is published under Apache 2.0. Both are served from a published FP8 quantization of the original weights rather than one we made, by the same publisher and with the same recipe.

What we serve

modelsizeinputquantizationweights served
gemma4-31b30.7B densetext + imagefp8RedHatAI/gemma-4-31B-it-FP8-dynamic
gemma4-12b12B densetext + image + audio + videofp8RedHatAI/gemma-4-12B-it-FP8-Dynamic
gemma4-31b
30.7B dense parameters - every one of them active on every token, unlike a mixture-of-experts model of nominally similar size. It is the largest dense checkpoint we serve, and the FP8 build fits a 48 GB-class card. Text and images only: a row carrying audio or video goes to gemma4-12b below, or to Qwen. Output is text, with a reasoning trace a job can cap with a reasoning budget.
gemma4-12b
The same family at a third of the weights, with the audio and video encoders the 31B does not have. Served from the same publisher's FP8 build as the 31B: the linear layers carry FP8 weights with dynamic FP8 activations, while the vision and audio embedders and the output layer keep the original BF16 (bfloat16) weights, so the encoders that read a picture or a clip and the layer that scores each next token are the base model's own. One 15 GB checkpoint serves every card we run it on, with no quantization step at load. It reads speech and short clips directly: there is no separate encoder tower in front of the model, raw audio and raw frames are projected into the decoder. Also a reasoning model, with the same reasoning budget. Its weights are Apache 2.0, which is the most permissive licence in our catalogue.
Images cost a fixed amount on both
This is the practical reason to choose either. Gemma encodes every image to exactly 280 tokens, whatever its size, instead of tiling a large picture into many more. A page scan and a thumbnail bill the same, which makes an image-heavy job's estimate almost entirely a function of how many images it carries rather than what they happen to be.
What gemma4-12b takes as media
Audio clips up to 30 seconds, at 25 tokens per second of clip. Video up to 60 seconds, sampled at one frame per second and each frame charged as an image. The audio ceiling is far narrower than the omni checkpoint on Nemotron, which takes 40 minutes in one row, so a long recording has to be cut into 30-second pieces here or sent there instead. The numbers are under limits below.

Limits

Both models serve a row in a 32,768-token context: at most 32,768 input tokens with the media counted in, at most 32,768 output tokens, and on the cards they run on today the input and the output together have to fit inside 32,768 as well. Media spends that same budget at a fixed rate, so the arithmetic is worth doing before a job is submitted.

Audio, gemma4-12b
30 seconds per clip, at 25 tokens per second. A full clip is 750 tokens, a small corner of the row, so the context is never what stops you here - the 30-second ceiling is. Longer recordings go to Nemotron or get cut into pieces.
Video, gemma4-12b
60 seconds per clip, sampled at one frame per second at 280 tokens a frame. A full minute is about 17,000 tokens, half the row, so leave room for the prompt and the answer.
Images, both models
280 tokens each, whatever the picture's size, and up to 2 images in a row. gemma4-31b reads text and images only, in the same 32,768 tokens.
A row that does not fit
It is rejected, and nothing is billed for it. If it is one of the rows sampled when the job is submitted, submit answers 400 with the reason; otherwise the job fails before it can be approved, with failure_code: invalid_input and a reason naming the row and the number it went over. The per-model table is under limits in the docs.

Benchmark scores

Published figures, each from the source named beside it. Different benchmarks and different harnesses, so read down the column rather than across it.

modelbenchmarkscoresource
gemma4-31bAA Intelligence Index39Artificial Analysis
gemma4-31bGPQA Diamond84.3%BenchLM
gemma4-31bMMLU-Pro85.2%BenchLM
gemma4-31bSWE-bench Verified~75%BenchLM
gemma4-31bAIME 202689.2%BenchLM
gemma4-31bagentic coding41.6BenchLM
gemma4-12bAA Intelligence Index (reasoning)22Artificial Analysis
gemma4-12bAA Intelligence Index (non-reasoning)20Artificial Analysis

The shape of the 31B's table is the point: strong on maths, knowledge and assistant-style work, weaker on agentic coding. Artificial Analysis also measured it as unusually terse for its score - it reached index 39 on 39M output tokens where a higher-scoring sub-32B peer needed 98M. On a batch bill, where output tokens are the expensive half, a model that says less for the same answer is worth more than its index suggests.

Two models here read audio, and the 12B's 22 is the higher index of the two: nemotron-3-nano-omni-30b scores 15 on the same index. So a job that has to both hear a clip and reason about what it heard is on stronger weights here than on the omni checkpoint, as long as 30 seconds of audio at a time is enough. Reasoning is what buys the two points between 22 and 20, and it is billed as output tokens, so the cheaper number is available to a job that caps its reasoning budget at zero.

Which one to pick

Image-heavy jobs with a budget to hold
Either one. The fixed 280-token image cost makes the estimate predictable in a way tiling encoders are not, especially when the images vary wildly in size. gemma4-12b is the cheaper of the two per token and has the lower minimum job price, so it is the default until an answer needs more of the 31B.
Maths, science and assistant-style answers
gemma4-31b. This is where it leads the models we serve at its size.
Short clips, and reasoning about what is in them
gemma4-12b. Speech up to 30 seconds and video up to a minute, on the stronger general-purpose weights of the two models here that read audio.
Long, repetitive runs where verbosity is the bill
gemma4-31b. Terse output at a competitive score is exactly the trade a batch queue wants.
Not for
Recordings longer than half a minute, video past a minute, or code and tool-use heavy work. Send long audio to Nemotron, which takes 40 minutes in one row, and a longer silent clip to qwen3.8-27b.

Pass gemma4-31b or gemma4-12b exactly as written when you submit a job. Their rates and their minimum job prices are on the pricing page.