SAVRN Model Hub · Comparisons
Gemma-4-31B-IT-NVFP4 vs gpt-oss-20b
Gemma-4-31B-IT-NVFP4 has 20.9B parameters and gpt-oss-20b has 20.9B parameters; Gemma-4-31B-IT-NVFP4 is released under other and gpt-oss-20b under Apache License 2.0; at 16-bit, Gemma-4-31B-IT-NVFP4 needs about 50.1 GB (1x MI300X from $1.85 an hour) and gpt-oss-20b about 50.2 GB (1x MI300X from $1.85 an hour).
| Field | Gemma-4-31B-IT-NVFP4 nvidia/Gemma-4-31B-IT-NVFP4 | gpt-oss-20b openai/gpt-oss-20b |
|---|---|---|
| Publisher | NVIDIA | OpenAI |
| Task | Text generation | Text generation |
| Modality | Text | Text |
| Parameters, as reported | 20.9B parameters | 20.9B parameters |
| Architecture | Gemma4ForConditionalGeneration | GptOssForCausalLM |
| Library | Model Optimizer | transformers |
| Context length | 262,144 tokens | 131,072 tokens |
| Repository size | 32.7 GB | 41.3 GB |
| Artifact formats | safetensors | safetensors |
| License | other | apache-2.0 |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 50.1 GB | 50.2 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Memory at 4-bit (weights and margin) | 12.5 GB | 12.5 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | 4135a98a9b72 | 6cee5e81ee83 |
| Downloads reported by the hub | 1.6M | 6.7M |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
Other Reported Results
These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.
gpt-oss-20b
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Idavidrein/gpqa | Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established | 58.5859 | EvalEval Reported by a third party |
Evaluated revision not stated | 2026-04-19 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: mediumComparison conditions not established | 66 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: high, With toolsComparison conditions not established | 74.2 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: highComparison conditions not established | 71.5 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: low, With toolsComparison conditions not established | 58 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: medium, With toolsComparison conditions not established | 67.1 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: lowComparison conditions not established | 56.8 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| LEXam-Benchmark/LEXam | Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established | 40.78 | LEXam Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-06-02 |
| LEXam-Benchmark/LEXam | Task open_questionMetric open_questionComparison conditions not established | 32.12 | LEXam Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-06-02 |
| LiquidAI/ifstruct-v1.0 | Task ifstruct_v1Metric ifstruct_v1Comparison conditions not established | 91.95 | Liquid AI — IFStruct v1.0 blog (gpt-oss-20b) Reported by a third party |
Evaluated revision not stated | 2026-06-30 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: lowComparison conditions not established | 37.4 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: mediumComparison conditions not established | 53.2 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: highComparison conditions not established | 60.7 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| TIGER-Lab/MMLU-Pro | Task mmlu_proMetric mmlu_proComparison conditions not established | 73.6 | EvalEval Reported by a third party |
Evaluated revision not stated | 2026-06-30 |
| cais/hle | Task hleMetric hleSetup Reasoning: lowComparison conditions not established | 4.2 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: medium, With toolsComparison conditions not established | 8.8 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: mediumComparison conditions not established | 7 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: highComparison conditions not established | 10.9 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: high, With toolsComparison conditions not established | 17.3 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: low, With toolsComparison conditions not established | 6.3 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
SAVRN's Notes on Gemma-4-31B-IT-NVFP4
Two numbers on this page pull against each other. The context window is 262,144 tokens, while the 4-bit row, the one that fits an NVFP4 quantization of google/gemma-4-31B-it, needs 12.5 GB of memory for 10.4 GB of weights. On the 192 GB MI300X our Index lists as the cheapest host at $1.85 per hour, the weights are a small tenant; the key-value cache for long inputs is what fills the card, so we size this one by cache, not by weights. The publisher lists text and image input and over 140 languages.
The license is recorded as other, with no commercial-use line on file, so legal reads the terms attached to the gemma-4-31B-it lineage before anything ships. No reported evaluations and no Index host prices exist for this quantization, so a buyer runs their own tests and weighs the 16-bit row's 50.1 GB against the 12.5 GB saved.
SAVRN's Notes on gpt-oss-20b
Five hosts on the SAVRN Index serve gpt-oss-20b by the token, from $0.03 in and $0.14 out per million at DeepInfra to $0.10 and $0.50 at Groq. Owned, this text generator fits on one card: 12.5 GB at 4-bit, 25.1 GB at 8-bit, 50.2 GB at 16-bit, and the least expensive card listed, a 192 GB MI300X, rents for $1.85 an hour, so even full precision leaves room for its 131,072-token context.
You can modify and resell this one; Apache 2.0 allows commercial use, modification and redistribution as long as the notices stay intact. The publisher says it was trained on its harmony response format and should only be used with it, so your serving stack has to speak harmony. And settle rent-or-own yourself: whether $0.14 per million output tokens beats $1.85 an hour depends on how many tokens a day you push through.
Questions
Which is larger, Gemma-4-31B-IT-NVFP4 or gpt-oss-20b?
gpt-oss-20b (20.9B parameters) is larger than Gemma-4-31B-IT-NVFP4 (20.9B parameters), by the parameter counts their publishers report.
Which is cheaper to run, Gemma-4-31B-IT-NVFP4 or gpt-oss-20b?
At 4-bit, Gemma-4-31B-IT-NVFP4 fits on 1x MI300X from $1.85 an hour and gpt-oss-20b on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use gpt-oss-20b commercially?
Yes. gpt-oss-20b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.