SAVRN Model Hub · Comparisons
gemma-4-31B-it vs Qwen3.8-27B
Gemma-4-31B-it has 31.3B parameters and Qwen3.8-27B has 27.8B parameters; both are released under Apache License 2.0; at 16-bit, gemma-4-31B-it needs about 75.1 GB (1x MI300X from $1.85 an hour) and Qwen3.8-27B about 66.7 GB (1x MI300X from $1.85 an hour).
| Field | gemma-4-31B-it google/gemma-4-31B-it | Qwen3.8-27B Qwen/Qwen3.8-27B |
|---|---|---|
| Publisher | Qwen | |
| Task | Image and text to text | Image and text to text |
| Modality | Image and text | Image and text |
| Parameters, as reported | 31.3B parameters | 27.8B parameters |
| Architecture | Gemma4ForConditionalGeneration | Qwen3_5ForConditionalGeneration |
| Library | transformers | transformers |
| Context length | 262,144 tokens | 262,144 tokens |
| Repository size | 62.6 GB | 55.6 GB |
| Artifact formats | safetensors | safetensors |
| License | apache-2.0 | apache-2.0 |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 75.1 GB | 66.7 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Memory at 4-bit (weights and margin) | 18.8 GB | 16.7 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | 842da3794eaa | 1d4bf0f2ff60 |
| Downloads reported by the hub | 9M | 7.4M |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
Other Reported Results
These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.
gemma-4-31B-it
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Idavidrein/gpqa | Task diamondMetric diamondComparison conditions not established | 84.3 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-04-02 |
| LiquidAI/ifstruct-v1.0 | Task ifstruct_v1Metric ifstruct_v1Comparison conditions not established | 95.9 | Liquid AI — IFStruct v1.0 blog (gemma-4-31B-it) Reported by a third party |
Evaluated revision not stated | 2026-06-30 |
| MMMU/MMMU_Pro | Task mmmu_pro_visionMetric mmmu_pro_visionComparison conditions not established | 76.9 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-05-12 |
| TIGER-Lab/MMLU-Pro | Task mmlu_proMetric mmlu_proComparison conditions not established | 85.2 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-04-02 |
| cais/hle | Task hleMetric hleSetup With searchComparison conditions not established | 26.5 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-04-02 |
| joelniklaus/LEXam-hard | Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established | 37.41 | SwissLegalEvals per-sample details (lighteval) Reported by a third party |
Evaluated revision not stated | 2026-06-11 |
| llamaindex/ParseBench | Task chartMetric chartSetup Pipeline name: gemma4_31b_vllm_with_layoutComparison conditions not established | 15 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-17 |
| llamaindex/ParseBench | Task layoutMetric layoutSetup Pipeline name: gemma4_31b_vllm_with_layoutComparison conditions not established | 57.4 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-17 |
| llamaindex/ParseBench | Task meanMetric meanSetup Pipeline name: gemma4_31b_vllm_with_layoutComparison conditions not established | 62.4 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-17 |
| llamaindex/ParseBench | Task tableMetric tableSetup Pipeline name: gemma4_31b_vllm_with_layoutComparison conditions not established | 80.6 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-17 |
| llamaindex/ParseBench | Task text_contentMetric text_contentSetup Pipeline name: gemma4_31b_vllm_with_layoutComparison conditions not established | 89.9 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-17 |
| llamaindex/ParseBench | Task text_formattingMetric text_formattingSetup Pipeline name: gemma4_31b_vllm_with_layoutComparison conditions not established | 69.3 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-17 |
| sapbot/ask-my-agent-bench-2 | Task fallbackMetric fallbackComparison conditions not established | 72.1 | Not named Reported by a third party |
Evaluated revision not stated | 2026-06-23 |
Qwen3.8-27B
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Idavidrein/gpqa | Task diamondMetric diamondComparison conditions not established | 89.2 | Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated | 2026-08-14 |
| ScaleAI/SWE-bench_Pro | Task SWE_Bench_ProMetric SWE_Bench_ProSetup Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, 256K context; baseline models re-evaluated on the same refined task set.Comparison conditions not established | 61.7 | Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated | 2026-08-14 |
| cais/hle | Task hleMetric hleSetup Judged by GPT-4o.Comparison conditions not established | 30.8 | Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated | 2026-08-14 |
| claw-eval/Claw-Eval | Task multimodalMetric multimodalSetup Reported as ClawEval-MM. Pass@3, the benchmark's own Pass³ methodology; card also reports a secondary 56.9 'Average' metric, not included here.Comparison conditions not established | 57.4 | Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated | 2026-08-14 |
| datacurve/deep-swe | Task deep_sweMetric deep_sweComparison conditions not established | 42.2 | Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated | 2026-08-14 |
| harborframework/terminal-bench-2.1 | Task terminalbench_2_1Metric terminalbench_2_1Setup Row labeled "(Terminus)" as the harness; no further hyperparameter footnote given for this row.Comparison conditions not established | 73 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-08-14 |
| internlm/WildClawBench | Task avg_timeMetric avg_timeComparison conditions not established | 516 | WildClawBench Reported by a third party |
Evaluated revision not stated | 2026-08-16 |
| internlm/WildClawBench | Task overallMetric overallComparison conditions not established | 48.0152 | WildClawBench Reported by a third party |
Evaluated revision not stated | 2026-08-16 |
| llamaindex/ExtractBench | Task longMetric longSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 38.45 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-24 |
| llamaindex/ExtractBench | Task meanMetric meanSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 89.75 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-24 |
| llamaindex/ExtractBench | Task mediumMetric mediumSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 87.54 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-24 |
| llamaindex/ExtractBench | Task shortMetric shortSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 94.68 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-24 |
| llamaindex/ParseBench | Task chartMetric chartSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 69.17 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-28 |
| llamaindex/ParseBench | Task layoutMetric layoutSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 69.9 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-28 |
| llamaindex/ParseBench | Task meanMetric meanSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 70.79 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-28 |
| llamaindex/ParseBench | Task tableMetric tableSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 66.82 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-28 |
| llamaindex/ParseBench | Task text_contentMetric text_contentSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 88.28 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-28 |
| llamaindex/ParseBench | Task text_formattingMetric text_formattingSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established | 59.77 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-28 |
SAVRN's Notes on gemma-4-31B-it
One MI300X at $1.85 per hour carries this 31.3 billion parameter Gemma at 16-bit: 62.5 GB of weights and 75.1 GB needed against the card's 192 GB. Quantizing brings that to 37.5 GB at 8-bit and 18.8 GB at 4-bit, but here it buys concurrent context, not fit. With 262,144 tokens of context and image input, we would stay at 16-bit and give the spare memory to long documents.
Apache 2.0 permits commercial use, modification and redistribution, provided the license and NOTICE file stay attached and significant changes are stated. It is the instruction-tuned build of google/gemma-4-31B, described at arXiv:2607.02770; audio input belongs to the E2B, E4B and 12B sizes, not this one. DeepInfra hosts it at $0.13 in and $0.38 out per million tokens, Novita at $0.14 and $0.40; the output rate is where your own facility makes its case.
SAVRN's Notes on Qwen3.8-27B
Plan around 66.7 GB for Qwen3.8-27B at 16-bit. That fits one MI300X with 192 GB, which the SAVRN Index prices at $1.85 an hour on demand, and 8-bit brings it to 33.3 GB, 4-bit to 16.7 GB, all on that one card. It reads images and text and writes text, with a 262,144-token context, so it belongs where one accelerator per instance is the budget and the inputs include pictures or long documents.
Apache 2.0 allows commercial use, modification and redistribution, provided the license and notice files stay with the weights and significant changes are stated. Weigh the card against renting tokens on the Index: Cerebras at $0.99 in and $1.49 out per million, DeepInfra at $0.20 and $2.50, Novita at $0.42 and $3.00, OVHcloud at $0.47 and $3.19. The benchmark figures on the page are reported by the model card and outside evaluators, not measured by SAVRN.
Questions
Which is larger, gemma-4-31B-it or Qwen3.8-27B?
gemma-4-31B-it (31.3B parameters) is larger than Qwen3.8-27B (27.8B parameters), by the parameter counts their publishers report.
Which is cheaper to run, gemma-4-31B-it or Qwen3.8-27B?
At 4-bit, gemma-4-31B-it fits on 1x MI300X from $1.85 an hour and Qwen3.8-27B on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use gemma-4-31B-it commercially?
Yes. gemma-4-31B-it is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Can I use Qwen3.8-27B commercially?
Yes. Qwen3.8-27B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.