Qwen3.5-2B-DSpark is an open-weight model from Yosef Worku Alemneh, released under Apache License 2.0. It has 452M parameters. At 16-bit it needs about 1.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 58 downloads a month.
A DSpark draft model for speculative decoding with Qwen/Qwen3.5-2B as the verifier, trained with speculators.
Runs On
What it takes to serve Qwen3.5-2B-DSpark (452M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.9 GB | 1.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.5 GB | 0.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.
Qwen3.5-2B-DSpark on every accelerator the SAVRN Index prices, at every precision
Model Card
By Yosef Worku Alemneh, published under apache-2.0, revision 1e39250b896f.
A DSpark draft model for speculative decoding with Qwen/Qwen3.5-2B as the verifier, trained with speculators. The drafter proposes 8 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone — a lossless speedup. Mean acceptance length is 2.50 tokens committed per verification step, up to 3.75 on math_reasoning.
Training code: rasyosef/train-dspark-draft-models.
Trained on 100,000 samples.
Usage
vLLM loads the verifier automatically from the config — don't pass it separately.
vllm serve rasyosef/Qwen3.5-2B-DSpark --port 8000 --gpu-memory-utilization 0.8
Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.
Details
5 Qwen3 layers (hidden size 2048, intermediate size 6144, 16 attention heads over 8 KV heads, head dim 128, all layers sliding-window attention with a 2048-token window), ~0.5B params, bfloat16. Block size 8, draft vocabulary reduced to 50,000 from the verifier's 248,320 (selected by token frequency over the training data), aux hidden-state layers 1/6/11/16/21, confidence head with Markov (vanilla, rank 256).
Configuration
- Architecture
- DSparkDraftModel
Identity and Version
- Repository
- rasyosef/Qwen3.5-2B-DSpark
- Publisher
- Yosef Worku Alemneh
- Task
- Not stated by the source
- Modality
- Other
- Library
- transformers
- Parameters
- 452M parameters
- Languages
- Not stated by the source
- Revision
- 1e39250b896fdc54d5121c1f6aea6ef6d3ff9dd2
- First published
- 2026-09-18
- Last updated
- 2026-09-29
Files and Weights
10 files, 1.8 GB in total. The weights are 3 files totalling 1.8 GB in pt, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 903.5 MB | 2d70b0eb2ef8 |
| optimizer_state_dict.pt | Weights | 850.9 MB | bd14bde98c9f |
| scheduler_state_dict.pt | Weights | 1.6 KB | 7f3f1221c4ac |
| config.json | Configuration | 2.0 KB | — |
| config.py | Configuration | 2.3 KB | — |
| training_state.json | Configuration | 51 B | — |
| val_metrics.json | Configuration | 812 B | — |
| README.md | Documentation | 5.2 KB | — |
| train_command.txt | Other | 778 B | — |
| .gitattributes | Repository | 1.5 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 1.8 GB
Released by Yosef Worku Alemneh through its official repository on Hugging Face. Read the license.
Built From
- Derived from Qwen/Qwen3.5-2B
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 1.8 GB |
| 16-bit | 0.9 GB |
| 8-bit | 0.5 GB |
| 4-bit | 0.2 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Qwen3.5-2B-DSpark
How much GPU memory does Qwen3.5-2B-DSpark need?
About 1.1 GB at 16-bit and 0.3 GB at 4-bit: the weights (452M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Qwen3.5-2B-DSpark on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Qwen3.5-2B-DSpark commercially?
Yes. Qwen3.5-2B-DSpark is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.