SAVRN
Search Contact SAVRN

SAVRN Model Hub · Comparisons

DeepSeek-V4-Flash-DSpark vs gpt-oss-120b

DeepSeek-V4-Flash-DSpark has 165.3B parameters and gpt-oss-120b has 116.8B parameters; DeepSeek-V4-Flash-DSpark is released under MIT License and gpt-oss-120b under Apache License 2.0; at 16-bit, DeepSeek-V4-Flash-DSpark needs about 396.6 GB (2x MI325X from $4.00 an hour) and gpt-oss-120b about 280.4 GB (1x MI355X from $2.59 an hour).

Published metadata for 2 models, each read from its own repository.
Field DeepSeek-V4-Flash-DSpark
deepseek-ai/DeepSeek-V4-Flash-DSpark
gpt-oss-120b
openai/gpt-oss-120b
Publisher DeepSeek OpenAI
Task Text generation Text generation
Modality Text Text
Parameters, as reported 165.3B parameters 116.8B parameters
Architecture DeepseekV4ForCausalLM GptOssForCausalLM
Library transformers transformers
Context length 1,048,576 tokens 131,072 tokens
Repository size 166.9 GB 195.8 GB
Artifact formats safetensors safetensors
License mit apache-2.0
Access Open weights, no gate Open weights, no gate
Memory at 16-bit (weights and margin) 396.6 GB 280.4 GB
Cheapest GPUs at 16-bit, per hour 2x MI325X, $4.00 1x MI355X, $2.59
Memory at 4-bit (weights and margin) 99.2 GB 70.1 GB
Cheapest GPUs at 4-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Revision viewed 62af8fffb2f7 b5c939de8f75
Downloads reported by the hub 1M 5.2M
Last observed 2026-09-18 2026-09-18

An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.

Other Reported Results

These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.

gpt-oss-120b

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established 80.8081 EvalEval
Reported by a third party
Evaluated revision not stated 2026-04-16
Idavidrein/gpqa Task diamondMetric diamondSetup Reasoning: mediumComparison conditions not established 73.1 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
Idavidrein/gpqa Task diamondMetric diamondSetup Reasoning: high, With toolsComparison conditions not established 80.9 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
Idavidrein/gpqa Task diamondMetric diamondSetup Reasoning: highComparison conditions not established 80.1 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
Idavidrein/gpqa Task diamondMetric diamondSetup Reasoning: low, With toolsComparison conditions not established 68.1 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
Idavidrein/gpqa Task diamondMetric diamondSetup Reasoning: medium, With toolsComparison conditions not established 73.5 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
Idavidrein/gpqa Task diamondMetric diamondSetup Reasoning: lowComparison conditions not established 67.1 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
LEXam-Benchmark/LEXam Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established 47.71 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
LEXam-Benchmark/LEXam Task open_questionMetric open_questionComparison conditions not established 51.74 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: lowComparison conditions not established 47.9 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: mediumComparison conditions not established 52.6 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: highComparison conditions not established 62.4 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
ScaleAI/SWE-bench_Pro Task SWE_Bench_ProMetric SWE_Bench_ProComparison conditions not established 16.2 SWE-Bench Pro official evaluation results
Reported by a third party
Evaluated revision not stated 2026-02-28
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 80.8 EvalEval
Reported by a third party
Evaluated revision not stated 2026-06-30
cais/hle Task hleMetric hleSetup Reasoning: lowComparison conditions not established 5.2 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
cais/hle Task hleMetric hleSetup Reasoning: medium, With toolsComparison conditions not established 11.3 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
cais/hle Task hleMetric hleSetup Reasoning: mediumComparison conditions not established 8.6 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
cais/hle Task hleMetric hleSetup Reasoning: highComparison conditions not established 14.9 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
cais/hle Task hleMetric hleSetup Reasoning: high, With toolsComparison conditions not established 19 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
cais/hle Task hleMetric hleSetup Reasoning: low, With toolsComparison conditions not established 9.1 GPT-OSS Model Card
Reported by a third party
Evaluated revision not stated 2025-08-05
joelniklaus/LEXam-hard Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established 37.55 SwissLegalEvals per-sample details (lighteval)
Reported by a third party
Evaluated revision not stated 2026-06-11

SAVRN's Notes on DeepSeek-V4-Flash-DSpark

The publisher says it outright: DSpark is not a new model. It is the DeepSeek-V4-Flash checkpoint with a speculative decoding module attached, so you are evaluating a serving path, not fresh weights. Those weights run 165.3 billion parameters across 256 routed experts with 6 active per token. Precision picks the hardware: 16-bit needs 396.6 GB and two MI325X cards at $4.00 per hour, 8-bit needs 198.3 GB on one MI325X at $2.00, and 4-bit needs 99.2 GB on one MI300X at $1.85.

MIT covers it: commercial use, modification and redistribution with the notices kept. Two checks before buying cards. The memory figures are the floor; a request that uses the 1,048,576-token context adds cache on top, sized from the single key/value head at 512 dimensions. And the speculative decoding path is the point of this release, so confirm your serving stack runs the publisher's inference example.

SAVRN's Notes on gpt-oss-120b

Memory decides the hardware. gpt-oss-120b carries 116.8 billion parameters; at 16-bit the working footprint is 280.4 GB, meaning one MI355X with 288 GB at $2.59 an hour. At 8-bit it drops to 140.2 GB and one MI300X with 192 GB at $1.85 an hour covers it, and 4-bit needs 70.1 GB on the same card. For general text generation with a 131,072-token window, we would start on the 8-bit single card.

Apache 2.0 permits commercial use, modification and redistribution, provided the license and copyright notices stay attached and significant changes are stated, plus an express patent grant. Before committing, read the model card at arXiv:2508.10925 and price the hosted route: on the Index, DeepInfra lists $0.037 in and $0.17 out per million tokens, Cerebras $0.35 and $0.75. Once a month of tokens at those rates costs more than an MI300X at $1.85 an hour, run it yourself.

Questions

Which is larger, DeepSeek-V4-Flash-DSpark or gpt-oss-120b?

DeepSeek-V4-Flash-DSpark (165.3B parameters) is larger than gpt-oss-120b (116.8B parameters), by the parameter counts their publishers report.

Which is cheaper to run, DeepSeek-V4-Flash-DSpark or gpt-oss-120b?

At 4-bit, DeepSeek-V4-Flash-DSpark fits on 1x MI300X from $1.85 an hour and gpt-oss-120b on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-V4-Flash-DSpark commercially?

Yes. DeepSeek-V4-Flash-DSpark is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Can I use gpt-oss-120b commercially?

Yes. gpt-oss-120b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Related Comparisons