SAVRN
Search Contact SAVRN

SAVRN Model Hub · Comparisons

DeepSeek-R1-0528-Qwen3-8B vs Llama-3.1-8B-Instruct

DeepSeek-R1-0528-Qwen3-8B has 8.2B parameters and Llama-3.1-8B-Instruct has 8B parameters; DeepSeek-R1-0528-Qwen3-8B is released under MIT License and Llama-3.1-8B-Instruct under Meta Llama 3.1 Community License; at 16-bit, DeepSeek-R1-0528-Qwen3-8B needs about 19.7 GB (1x MI300X from $1.85 an hour) and Llama-3.1-8B-Instruct about 19.3 GB (1x MI300X from $1.85 an hour).

Published metadata for 2 models, each read from its own repository.
Field DeepSeek-R1-0528-Qwen3-8B
deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
Llama-3.1-8B-Instruct
meta-llama/Llama-3.1-8B-Instruct
Publisher DeepSeek Meta Llama
Task Text generation Text generation
Modality Text Text
Parameters, as reported 8.2B parameters 8B parameters
Architecture Qwen3ForCausalLM LlamaForCausalLM
Library transformers transformers
Context length 131,072 tokens Not stated
Repository size 16.4 GB 32.1 GB
Artifact formats safetensors safetensors, pytorch
License mit llama3.1
Access Open weights, no gate Access requested at publisher
Memory at 16-bit (weights and margin) 19.7 GB 19.3 GB
Cheapest GPUs at 16-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Memory at 4-bit (weights and margin) 4.9 GB 4.8 GB
Cheapest GPUs at 4-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Revision viewed 6e8885a6ff5c 0e9e39f249a1
Downloads reported by the hub 895.6k 5.9M
Last observed 2026-09-18 2026-09-18

An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.

Other Reported Results

These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.

Llama-3.1-8B-Instruct

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 30.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-01-27
LEXam-Benchmark/LEXam Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established 24.04 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
LEXam-Benchmark/LEXam Task open_questionMetric open_questionComparison conditions not established 10 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
openai/gsm8k Task gsm8kMetric gsm8kComparison conditions not established 84.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-03-23
thamilvendhan/signalbench Task access_denyMetric access_denySetup family=access_deny; n=12Comparison conditions not established 0.5 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task bot_policyMetric bot_policySetup family=bot_policy; n=12Comparison conditions not established 0.4167 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task injectionMetric injectionSetup family=injection; n=12Comparison conditions not established 0.75 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task memory_labelMetric memory_labelSetup family=memory_label; n=12Comparison conditions not established 0.6667 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task srcMetric srcSetup SRC overall; deterministic action-based grader, no LLM judge; seed 0, n=75Comparison conditions not established 0.6333 signalbench raw per-item responses
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task timeMetric timeSetup family=time; n=12Comparison conditions not established 0.8333 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08

SAVRN's Notes on DeepSeek-R1-0528-Qwen3-8B

At 16-bit precision this model asks for 19.7 GB of memory, which settles the hardware question early. DeepSeek built it on the Qwen3ForCausalLM architecture at 8.2 billion parameters for text generation, with a 131,072-token context window and weights that occupy 16.4 GB as safetensors. Drop to 8-bit and the memory need falls to 9.8 GB; at 4-bit it is 4.9 GB. The cheapest setup on our Index is a single MI300X with 192 GB at $1.85 per hour on demand, so one card holds it many times over.

The MIT license permits commercial use, modification and redistribution as long as the copyright and permission notices stay with the files. Before committing, confirm your workload needs the full 131,072-token context, read the describing paper arXiv:2501.12948, and note that our Index lists no per-token host prices for this model yet, so the hourly card rate is the only cost reference.

SAVRN's Notes on Llama-3.1-8B-Instruct

At 16-bit this checkpoint needs 19.3 GB of memory, and the cheapest SAVRN Index listing that covers it is a single 192 GB MI300X at $1.85 per hour on-demand. That card is the cheapest answer at 8-bit (9.6 GB) and 4-bit (4.8 GB) too, so quantizing does not buy a cheaper hour; it buys room for more concurrent multilingual dialogue sessions on one card.

Commercial use is allowed under the Meta Llama 3.1 Community License, with attribution and Meta's Acceptable Use Policy observed, unless your products had more than 700 million monthly active users on the release date; then you request a license from Meta. Gated access means approval precedes download. Two checks: the page lists no context length, and the $1.85 hour must beat the Index's hosted rates, $0.02 in and $0.05 out per million tokens at DeepInfra and Novita, $0.06 both ways at Nscale.

Questions

Which is larger, DeepSeek-R1-0528-Qwen3-8B or Llama-3.1-8B-Instruct?

DeepSeek-R1-0528-Qwen3-8B (8.2B parameters) is larger than Llama-3.1-8B-Instruct (8B parameters), by the parameter counts their publishers report.

Which is cheaper to run, DeepSeek-R1-0528-Qwen3-8B or Llama-3.1-8B-Instruct?

At 4-bit, DeepSeek-R1-0528-Qwen3-8B fits on 1x MI300X from $1.85 an hour and Llama-3.1-8B-Instruct on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-R1-0528-Qwen3-8B commercially?

Yes. DeepSeek-R1-0528-Qwen3-8B is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Can I use Llama-3.1-8B-Instruct commercially?

Yes, with conditions. Llama-3.1-8B-Instruct is released under Meta Llama 3.1 Community License. The Llama 3.1 Community License permits commercial use, except that a licensee whose products had more than 700 million monthly active users on the release date must request a license from Meta. It requires attribution as the license specifies and compliance with Meta's Acceptable Use Policy.

Related Comparisons