SAVRN
Search Contact SAVRN

SAVRN Model Hub · Comparisons

DeepSeek-R1-0528-Qwen3-8B vs Qwen3-8B

DeepSeek-R1-0528-Qwen3-8B has 8.2B parameters and Qwen3-8B has 8.2B parameters; DeepSeek-R1-0528-Qwen3-8B is released under MIT License and Qwen3-8B under Apache License 2.0; at 16-bit, DeepSeek-R1-0528-Qwen3-8B needs about 19.7 GB (1x MI300X from $1.85 an hour) and Qwen3-8B about 19.7 GB (1x MI300X from $1.85 an hour).

Published metadata for 2 models, each read from its own repository.
Field DeepSeek-R1-0528-Qwen3-8B
deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
Qwen3-8B
Qwen/Qwen3-8B
Publisher DeepSeek Qwen
Task Text generation Text generation
Modality Text Text
Parameters, as reported 8.2B parameters 8.2B parameters
Architecture Qwen3ForCausalLM Qwen3ForCausalLM
Library transformers transformers
Context length 131,072 tokens 40,960 tokens
Repository size 16.4 GB 16.4 GB
Artifact formats safetensors safetensors
License mit apache-2.0
Access Open weights, no gate Open weights, no gate
Memory at 16-bit (weights and margin) 19.7 GB 19.7 GB
Cheapest GPUs at 16-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Memory at 4-bit (weights and margin) 4.9 GB 4.9 GB
Cheapest GPUs at 4-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Revision viewed 6e8885a6ff5c b968826d9c46
Downloads reported by the hub 895.6k 13M
Last observed 2026-09-18 2026-09-18

An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.

Other Reported Results

These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.

Qwen3-8B

BenchmarkConditionsResultReported byRevisionDate
LiquidAI/ifstruct-v1.0 Task ifstruct_v1Metric ifstruct_v1Comparison conditions not established 79.75 Liquid AI — IFStruct v1.0 blog (Qwen3-8B)
Reported by a third party
Evaluated revision not stated 2026-06-30

SAVRN's Notes on DeepSeek-R1-0528-Qwen3-8B

At 16-bit precision this model asks for 19.7 GB of memory, which settles the hardware question early. DeepSeek built it on the Qwen3ForCausalLM architecture at 8.2 billion parameters for text generation, with a 131,072-token context window and weights that occupy 16.4 GB as safetensors. Drop to 8-bit and the memory need falls to 9.8 GB; at 4-bit it is 4.9 GB. The cheapest setup on our Index is a single MI300X with 192 GB at $1.85 per hour on demand, so one card holds it many times over.

The MIT license permits commercial use, modification and redistribution as long as the copyright and permission notices stay with the files. Before committing, confirm your workload needs the full 131,072-token context, read the describing paper arXiv:2501.12948, and note that our Index lists no per-token host prices for this model yet, so the hourly card rate is the only cost reference.

SAVRN's Notes on Qwen3-8B

At 16-bit the weights are 16.4 GB and the model needs 19.7 GB, which fits one MI300X with 192 GB, the Index's cheapest setup at $1.85 an hour on-demand. At 8-bit the need drops to 9.8 GB and at 4-bit to 4.9 GB, so the question is never which card but how many copies to stack on one. With 8.2 billion parameters, a 40,960-token context and a switch between thinking and non-thinking modes, this is a text model for everyday work.

The license is the easy part: Apache 2.0 permits commercial use, modification and redistribution, with notices and any NOTICE file kept and significant changes stated, plus an express patent grant. Two things to weigh: it derives from Qwen3-8B-Base, so decide whether you want this checkpoint or the base, and compare the hourly card against the one Index host price, Nscale at $0.07 in and $0.18 out per million tokens.

Questions

Which is larger, DeepSeek-R1-0528-Qwen3-8B or Qwen3-8B?

DeepSeek-R1-0528-Qwen3-8B (8.2B parameters) is larger than Qwen3-8B (8.2B parameters), by the parameter counts their publishers report.

Which is cheaper to run, DeepSeek-R1-0528-Qwen3-8B or Qwen3-8B?

At 4-bit, DeepSeek-R1-0528-Qwen3-8B fits on 1x MI300X from $1.85 an hour and Qwen3-8B on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-R1-0528-Qwen3-8B commercially?

Yes. DeepSeek-R1-0528-Qwen3-8B is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Can I use Qwen3-8B commercially?

Yes. Qwen3-8B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Related Comparisons