The family
Gemma is Google's open-weight line, built from the research behind Gemini and published as downloadable checkpoints rather than API access. Where Qwen ships many checkpoints, Gemma ships fewer and larger dense models, and puts the effort into what a single checkpoint does well.
Two Gemma models are on offer here: the multimodal 31B instruct checkpoint and the 12B, a third of the size, which reads audio and video on top of images and is published under Apache 2.0. Both are served from a published FP8 quantization of the original weights rather than one we made, by the same publisher and with the same recipe.
What we serve
| model | size | input | quantization | weights served |
|---|---|---|---|---|
| gemma4-31b | 30.7B dense | text + image | fp8 | RedHatAI/gemma-4-31B-it-FP8-dynamic |
| gemma4-12b | 12B dense | text + image + audio + video | fp8 | RedHatAI/gemma-4-12B-it-FP8-Dynamic |
- gemma4-31b
-
30.7B dense parameters - every one of them active on every token,
unlike a mixture-of-experts model of nominally similar size. It is
the largest dense checkpoint we serve, and the FP8 build fits a
48 GB-class card. Text and images only: a row carrying audio or
video goes to
gemma4-12bbelow, or to Qwen. Output is text, with a reasoning trace a job can cap with a reasoning budget. - gemma4-12b
- The same family at a third of the weights, with the audio and video encoders the 31B does not have. Served from the same publisher's FP8 build as the 31B: the linear layers carry FP8 weights with dynamic FP8 activations, while the vision and audio embedders and the output layer keep the original BF16 (bfloat16) weights, so the encoders that read a picture or a clip and the layer that scores each next token are the base model's own. One 15 GB checkpoint serves every card we run it on, with no quantization step at load. It reads speech and short clips directly: there is no separate encoder tower in front of the model, raw audio and raw frames are projected into the decoder. Also a reasoning model, with the same reasoning budget. Its weights are Apache 2.0, which is the most permissive licence in our catalogue.
- Images cost a fixed amount on both
- This is the practical reason to choose either. Gemma encodes every image to exactly 280 tokens, whatever its size, instead of tiling a large picture into many more. A page scan and a thumbnail bill the same, which makes an image-heavy job's estimate almost entirely a function of how many images it carries rather than what they happen to be.
- What gemma4-12b takes as media
- Audio clips up to 30 seconds, at 25 tokens per second of clip. Video up to 60 seconds, sampled at one frame per second and each frame charged as an image. The audio ceiling is far narrower than the omni checkpoint on Nemotron, which takes 40 minutes in one row, so a long recording has to be cut into 30-second pieces here or sent there instead. The numbers are under limits below.
Limits
Both models serve a row in a 32,768-token context: at most 32,768 input tokens with the media counted in, at most 32,768 output tokens, and on the cards they run on today the input and the output together have to fit inside 32,768 as well. Media spends that same budget at a fixed rate, so the arithmetic is worth doing before a job is submitted.
- Audio, gemma4-12b
- 30 seconds per clip, at 25 tokens per second. A full clip is 750 tokens, a small corner of the row, so the context is never what stops you here - the 30-second ceiling is. Longer recordings go to Nemotron or get cut into pieces.
- Video, gemma4-12b
- 60 seconds per clip, sampled at one frame per second at 280 tokens a frame. A full minute is about 17,000 tokens, half the row, so leave room for the prompt and the answer.
- Images, both models
-
280 tokens each, whatever the picture's size, and up to 2 images
in a row.
gemma4-31breads text and images only, in the same 32,768 tokens. - A row that does not fit
-
It is rejected, and nothing is billed for it. If it is one of the
rows sampled when the job is submitted, submit answers
400with the reason; otherwise the job fails before it can be approved, withfailure_code: invalid_inputand a reason naming the row and the number it went over. The per-model table is under limits in the docs.
Benchmark scores
Published figures, each from the source named beside it. Different benchmarks and different harnesses, so read down the column rather than across it.
| model | benchmark | score | source |
|---|---|---|---|
| gemma4-31b | AA Intelligence Index | 39 | Artificial Analysis |
| gemma4-31b | GPQA Diamond | 84.3% | BenchLM |
| gemma4-31b | MMLU-Pro | 85.2% | BenchLM |
| gemma4-31b | SWE-bench Verified | ~75% | BenchLM |
| gemma4-31b | AIME 2026 | 89.2% | BenchLM |
| gemma4-31b | agentic coding | 41.6 | BenchLM |
| gemma4-12b | AA Intelligence Index (reasoning) | 22 | Artificial Analysis |
| gemma4-12b | AA Intelligence Index (non-reasoning) | 20 | Artificial Analysis |
The shape of the 31B's table is the point: strong on maths, knowledge and assistant-style work, weaker on agentic coding. Artificial Analysis also measured it as unusually terse for its score - it reached index 39 on 39M output tokens where a higher-scoring sub-32B peer needed 98M. On a batch bill, where output tokens are the expensive half, a model that says less for the same answer is worth more than its index suggests.
Two models here read audio, and the 12B's 22 is the higher index of
the two: nemotron-3-nano-omni-30b scores 15 on the same
index. So a job that has to both hear a clip and reason about what it
heard is on stronger weights here than on the omni checkpoint, as long
as 30 seconds of audio at a time is enough. Reasoning is what buys the
two points between 22 and 20, and it is billed as output tokens, so
the cheaper number is available to a job that caps its reasoning
budget at zero.
Which one to pick
- Image-heavy jobs with a budget to hold
- Either one. The fixed 280-token image cost makes the estimate predictable in a way tiling encoders are not, especially when the images vary wildly in size. gemma4-12b is the cheaper of the two per token and has the lower minimum job price, so it is the default until an answer needs more of the 31B.
- Maths, science and assistant-style answers
- gemma4-31b. This is where it leads the models we serve at its size.
- Short clips, and reasoning about what is in them
- gemma4-12b. Speech up to 30 seconds and video up to a minute, on the stronger general-purpose weights of the two models here that read audio.
- Long, repetitive runs where verbosity is the bill
- gemma4-31b. Terse output at a competitive score is exactly the trade a batch queue wants.
- Not for
- Recordings longer than half a minute, video past a minute, or code and tool-use heavy work. Send long audio to Nemotron, which takes 40 minutes in one row, and a longer silent clip to qwen3.8-27b.
Pass gemma4-31b or gemma4-12b exactly as
written when you submit a job. Their
rates and their minimum job prices are on the
pricing page.