SAVRN Model Hub · Comparisons
gpt-oss-120b vs NVIDIA-Nemotron-3-Super-120B-A12B-BF16
Gpt-oss-120b has 116.8B parameters and NVIDIA-Nemotron-3-Super-120B-A12B-BF16 has 123.6B parameters; gpt-oss-120b is released under Apache License 2.0 and NVIDIA-Nemotron-3-Super-120B-A12B-BF16 under other; at 16-bit, gpt-oss-120b needs about 280.4 GB (1x MI355X from $2.59 an hour) and NVIDIA-Nemotron-3-Super-120B-A12B-BF16 about 296.7 GB (2x MI300X from $3.70 an hour).
| Field | gpt-oss-120b openai/gpt-oss-120b | NVIDIA-Nemotron-3-Super-120B-A12B-BF16 nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 |
|---|---|---|
| Publisher | OpenAI | NVIDIA |
| Task | Text generation | Text generation |
| Modality | Text | Text |
| Parameters, as reported | 116.8B parameters | 123.6B parameters |
| Architecture | GptOssForCausalLM | NemotronHForCausalLM |
| Library | transformers | transformers |
| Context length | 131,072 tokens | 262,144 tokens |
| Repository size | 195.8 GB | 247.2 GB |
| Artifact formats | safetensors | safetensors, pytorch |
| License | apache-2.0 | other |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 280.4 GB | 296.7 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI355X, $2.59 | 2x MI300X, $3.70 |
| Memory at 4-bit (weights and margin) | 70.1 GB | 74.2 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | b5c939de8f75 | 2dc98e2afe4f |
| Downloads reported by the hub | 5.2M | 1.3M |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
Other Reported Results
These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.
gpt-oss-120b
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Idavidrein/gpqa | Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established | 80.8081 | EvalEval Reported by a third party |
Evaluated revision not stated | 2026-04-16 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: mediumComparison conditions not established | 73.1 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: high, With toolsComparison conditions not established | 80.9 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: highComparison conditions not established | 80.1 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: low, With toolsComparison conditions not established | 68.1 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: medium, With toolsComparison conditions not established | 73.5 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: lowComparison conditions not established | 67.1 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| LEXam-Benchmark/LEXam | Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established | 47.71 | LEXam Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-06-02 |
| LEXam-Benchmark/LEXam | Task open_questionMetric open_questionComparison conditions not established | 51.74 | LEXam Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-06-02 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: lowComparison conditions not established | 47.9 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: mediumComparison conditions not established | 52.6 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: highComparison conditions not established | 62.4 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| ScaleAI/SWE-bench_Pro | Task SWE_Bench_ProMetric SWE_Bench_ProComparison conditions not established | 16.2 | SWE-Bench Pro official evaluation results Reported by a third party |
Evaluated revision not stated | 2026-02-28 |
| TIGER-Lab/MMLU-Pro | Task mmlu_proMetric mmlu_proComparison conditions not established | 80.8 | EvalEval Reported by a third party |
Evaluated revision not stated | 2026-06-30 |
| cais/hle | Task hleMetric hleSetup Reasoning: lowComparison conditions not established | 5.2 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: medium, With toolsComparison conditions not established | 11.3 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: mediumComparison conditions not established | 8.6 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: highComparison conditions not established | 14.9 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: high, With toolsComparison conditions not established | 19 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: low, With toolsComparison conditions not established | 9.1 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| joelniklaus/LEXam-hard | Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established | 37.55 | SwissLegalEvals per-sample details (lighteval) Reported by a third party |
Evaluated revision not stated | 2026-06-11 |
NVIDIA-Nemotron-3-Super-120B-A12B-BF16
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Idavidrein/gpqa | Task diamondMetric diamondComparison conditions not established | 79.23 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-12 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup With toolsComparison conditions not established | 82.7 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-12 |
| MathArena/aime_2026 | Task MathArena/aime_2026Metric MathArena/aime_2026Comparison conditions not established | 90 | Official MathArena Evaluation Reported by a third party |
Evaluated revision not stated | 2026-03-17 |
| MathArena/hmmt_feb_2026 | Task MathArena/hmmt_feb_2026Metric MathArena/hmmt_feb_2026Comparison conditions not established | 84.85 | Official MathArena Evaluation Reported by a third party |
Evaluated revision not stated | 2026-03-17 |
| SWE-bench/SWE-bench_Multilingual | Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedComparison conditions not established | 45.8 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-08-10 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup OpenCode harnessComparison conditions not established | 59.2 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-23 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Codex harnessComparison conditions not established | 53.73 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-23 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup OpenHands harnessComparison conditions not established | 60.47 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-23 |
| TIGER-Lab/MMLU-Pro | Task mmlu_proMetric mmlu_proComparison conditions not established | 83.73 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-12 |
| cais/hle | Task hleMetric hleComparison conditions not established | 18.26 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-12 |
| cais/hle | Task hleMetric hleSetup With toolsComparison conditions not established | 22.82 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-12 |
| claw-eval/Claw-Eval | Task generalMetric generalSetup Pass³% | N=3 | 161 tasksComparison conditions not established | 6.8 | Claw-Eval Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-23 |
| claw-eval/Claw-Eval | Task multi_turnMetric multi_turnSetup Pass³% | N=3 | 38 tasksComparison conditions not established | 0 | Claw-Eval Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-23 |
| harborframework/terminal-bench-2.0 | Task terminalbench_2Metric terminalbench_2Comparison conditions not established | 31 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-03-12 |
SAVRN's Notes on gpt-oss-120b
Memory decides the hardware. gpt-oss-120b carries 116.8 billion parameters; at 16-bit the working footprint is 280.4 GB, meaning one MI355X with 288 GB at $2.59 an hour. At 8-bit it drops to 140.2 GB and one MI300X with 192 GB at $1.85 an hour covers it, and 4-bit needs 70.1 GB on the same card. For general text generation with a 131,072-token window, we would start on the 8-bit single card.
Apache 2.0 permits commercial use, modification and redistribution, provided the license and copyright notices stay attached and significant changes are stated, plus an express patent grant. Before committing, read the model card at arXiv:2508.10925 and price the hosted route: on the Index, DeepInfra lists $0.037 in and $0.17 out per million tokens, Cerebras $0.35 and $0.75. Once a month of tokens at those rates costs more than an MI300X at $1.85 an hour, run it yourself.
Questions
Which is larger, gpt-oss-120b or NVIDIA-Nemotron-3-Super-120B-A12B-BF16?
NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (123.6B parameters) is larger than gpt-oss-120b (116.8B parameters), by the parameter counts their publishers report.
Which is cheaper to run, gpt-oss-120b or NVIDIA-Nemotron-3-Super-120B-A12B-BF16?
At 4-bit, gpt-oss-120b fits on 1x MI300X from $1.85 an hour and NVIDIA-Nemotron-3-Super-120B-A12B-BF16 on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use gpt-oss-120b commercially?
Yes. gpt-oss-120b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.