SAVRN Model Hub · Comparisons
Kimi-K3-DSpark vs Qwen3-1.7B-Base
Kimi-K3-DSpark has 2.2B parameters and Qwen3-1.7B-Base has 1.7B parameters; at 16-bit, Kimi-K3-DSpark needs about 5.4 GB (1x MI300X from $1.85 an hour) and Qwen3-1.7B-Base about 4.1 GB (1x MI300X from $1.85 an hour).
| Field | Kimi-K3-DSpark RadixArk/Kimi-K3-DSpark | Qwen3-1.7B-Base Qwen/Qwen3-1.7B-Base |
|---|---|---|
| Publisher | RadixArk | Qwen |
| Task | Text generation | Text generation |
| Modality | Text | Text |
| Parameters, as reported | 2.2B parameters | 1.7B parameters |
| Architecture | DSparkDraftModel | Qwen3ForCausalLM |
| Library | transformers | transformers |
| Context length | 1,048,576 tokens | 32,768 tokens |
| Repository size | 4.5 GB | 3.5 GB |
| Artifact formats | safetensors | safetensors |
| License | Not stated | apache-2.0 |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 5.4 GB | 4.1 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Memory at 4-bit (weights and margin) | 1.3 GB | 1 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | 3c5bac301d9c | ea980cb0a6c2 |
| Downloads reported by the hub | 3.5M | 1.6M |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
SAVRN's Notes on Kimi-K3-DSpark
Read this as a companion, not a model you serve on its own. RadixArk trained Kimi-K3-DSpark as a DSpark speculator for the Kimi K3 target, a 2.2 billion parameter draft built for faster inference through speculative decoding. Five layers and 4.5 GB of weights keep it light: 5.4 GB of memory at 16-bit, 2.7 GB at 8-bit, 1.3 GB at 4-bit, and one MI300X with 192 GB at $1.85 an hour on-demand is the cheapest listing we track. That covers the draft only; the Kimi K3 target sizes your box.
The license field is empty, so get written terms from RadixArk before deployment; open access is not permission to run it commercially. Match your stack too: the checkpoint was trained with SpecForge on hidden states from a live SGLang target engine, and the 1,048,576 token context is the draft's ceiling, so the target needs the same window.
SAVRN's Notes on Qwen3-1.7B-Base
At 4.1 GB of memory for 16-bit inference, 2.1 GB at 8-bit and 1.0 GB at 4-bit, this checkpoint fits on any accelerator we would rack; the cheapest setup on file, one MI300X with 192 GB at $1.85 per hour on-demand, could hold dozens of copies. A 1.7 billion parameter base model is raw material for post-training, and 32,768 tokens of context leave room to shape it for a narrow task.
Base in the title is the fact that matters: Qwen ships the pretraining checkpoint, so a deployment needs its own fine-tuning or prompting layer first. Apache 2.0 permits that and lets you redistribute the tuned result commercially, as long as notices stay intact and changes are stated. Read the Qwen3 technical report at arXiv:2505.09388 before committing; our record holds no reported evaluations and no SAVRN Index host prices yet, so the benchmark is yours.
Questions
Which is larger, Kimi-K3-DSpark or Qwen3-1.7B-Base?
Kimi-K3-DSpark (2.2B parameters) is larger than Qwen3-1.7B-Base (1.7B parameters), by the parameter counts their publishers report.
Which is cheaper to run, Kimi-K3-DSpark or Qwen3-1.7B-Base?
At 4-bit, Kimi-K3-DSpark fits on 1x MI300X from $1.85 an hour and Qwen3-1.7B-Base on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Qwen3-1.7B-Base commercially?
Yes. Qwen3-1.7B-Base is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.