SAVRN Model Hub · Comparisons
gpt-neox-20b vs gpt-oss-20b
Gpt-neox-20b has 20.7B parameters and gpt-oss-20b has 20.9B parameters; both are released under Apache License 2.0; at 16-bit, gpt-neox-20b needs about 49.8 GB (1x MI300X from $1.85 an hour) and gpt-oss-20b about 50.2 GB (1x MI300X from $1.85 an hour).
| Field | gpt-neox-20b EleutherAI/gpt-neox-20b | gpt-oss-20b openai/gpt-oss-20b |
|---|---|---|
| Publisher | EleutherAI | OpenAI |
| Task | Text generation | Text generation |
| Modality | Text | Text |
| Parameters, as reported | 20.7B parameters | 20.9B parameters |
| Architecture | GPTNeoXForCausalLM | GptOssForCausalLM |
| Library | transformers | transformers |
| Context length | 2,048 tokens | 131,072 tokens |
| Repository size | 82.6 GB | 41.3 GB |
| Artifact formats | safetensors, pytorch | safetensors |
| License | apache-2.0 | apache-2.0 |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 49.8 GB | 50.2 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Memory at 4-bit (weights and margin) | 12.4 GB | 12.5 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | c292233c833e | 6cee5e81ee83 |
| Downloads reported by the hub | 707.1k | 6.7M |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
Other Reported Results
These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.
gpt-oss-20b
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Idavidrein/gpqa | Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established | 58.5859 | EvalEval Reported by a third party |
Evaluated revision not stated | 2026-04-19 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: mediumComparison conditions not established | 66 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: high, With toolsComparison conditions not established | 74.2 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: highComparison conditions not established | 71.5 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: low, With toolsComparison conditions not established | 58 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: medium, With toolsComparison conditions not established | 67.1 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| Idavidrein/gpqa | Task diamondMetric diamondSetup Reasoning: lowComparison conditions not established | 56.8 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| LEXam-Benchmark/LEXam | Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established | 40.78 | LEXam Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-06-02 |
| LEXam-Benchmark/LEXam | Task open_questionMetric open_questionComparison conditions not established | 32.12 | LEXam Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-06-02 |
| LiquidAI/ifstruct-v1.0 | Task ifstruct_v1Metric ifstruct_v1Comparison conditions not established | 91.95 | Liquid AI — IFStruct v1.0 blog (gpt-oss-20b) Reported by a third party |
Evaluated revision not stated | 2026-06-30 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: lowComparison conditions not established | 37.4 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: mediumComparison conditions not established | 53.2 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| SWE-bench/SWE-bench_Verified | Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: highComparison conditions not established | 60.7 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| TIGER-Lab/MMLU-Pro | Task mmlu_proMetric mmlu_proComparison conditions not established | 73.6 | EvalEval Reported by a third party |
Evaluated revision not stated | 2026-06-30 |
| cais/hle | Task hleMetric hleSetup Reasoning: lowComparison conditions not established | 4.2 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: medium, With toolsComparison conditions not established | 8.8 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: mediumComparison conditions not established | 7 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: highComparison conditions not established | 10.9 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: high, With toolsComparison conditions not established | 17.3 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
| cais/hle | Task hleMetric hleSetup Reasoning: low, With toolsComparison conditions not established | 6.3 | GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated | 2025-08-05 |
SAVRN's Notes on gpt-neox-20b
20.7 billion parameters, a 2,048-token window, and a training set with its own datasheet. EleutherAI released it on April 7, 2022 as a general-purpose English text generator whose design tracks GPT-3 and sits close to GPT-J-6B. Memory is comfortable: 49.8 GB at 16-bit, 24.9 GB at 8-bit, 12.4 GB at 4-bit, all inside the single MI300X at $1.85 per hour in the cheapest listed setup. The 2,048-token context sets the job: short prompts, completion and fine-tuning experiments rather than long documents.
Apache 2.0 gives you commercial use, modification and redistribution with notice obligations and a patent grant, so a fine-tuned derivative can be redistributed if the notices travel with it and significant changes are stated. Read what you inherit: the Pile is described in arXiv:2101.00027 and its datasheet in arXiv:2201.07311, and the model paper is arXiv:2204.06745, which answer the provenance questions a review board will raise.
SAVRN's Notes on gpt-oss-20b
Five hosts on the SAVRN Index serve gpt-oss-20b by the token, from $0.03 in and $0.14 out per million at DeepInfra to $0.10 and $0.50 at Groq. Owned, this text generator fits on one card: 12.5 GB at 4-bit, 25.1 GB at 8-bit, 50.2 GB at 16-bit, and the least expensive card listed, a 192 GB MI300X, rents for $1.85 an hour, so even full precision leaves room for its 131,072-token context.
You can modify and resell this one; Apache 2.0 allows commercial use, modification and redistribution as long as the notices stay intact. The publisher says it was trained on its harmony response format and should only be used with it, so your serving stack has to speak harmony. And settle rent-or-own yourself: whether $0.14 per million output tokens beats $1.85 an hour depends on how many tokens a day you push through.
Questions
Which is larger, gpt-neox-20b or gpt-oss-20b?
gpt-oss-20b (20.9B parameters) is larger than gpt-neox-20b (20.7B parameters), by the parameter counts their publishers report.
Which is cheaper to run, gpt-neox-20b or gpt-oss-20b?
At 4-bit, gpt-neox-20b fits on 1x MI300X from $1.85 an hour and gpt-oss-20b on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use gpt-neox-20b commercially?
Yes. gpt-neox-20b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Can I use gpt-oss-20b commercially?
Yes. gpt-oss-20b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.