SAVRN
Search Contact SAVRN

Open-weight model

Qwen3.5-2B-DSpark

by Yosef Worku Alemneh rasyosef/Qwen3.5-2B-DSpark

Qwen3.5-2B-DSpark is an open-weight model from Yosef Worku Alemneh, released under Apache License 2.0. It has 452M parameters. At 16-bit it needs about 1.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 58 downloads a month.

A DSpark draft model for speculative decoding with Qwen/Qwen3.5-2B as the verifier, trained with speculators.

Parameters452M
Context—
Weights1.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads58

Runs On

What it takes to serve Qwen3.5-2B-DSpark (452M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.9 GB 1.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.5 GB 0.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

Qwen3.5-2B-DSpark on every accelerator the SAVRN Index prices, at every precision

Model Card

By Yosef Worku Alemneh, published under apache-2.0, revision 1e39250b896f.

A DSpark draft model for speculative decoding with Qwen/Qwen3.5-2B as the verifier, trained with speculators. The drafter proposes 8 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone — a lossless speedup. Mean acceptance length is 2.50 tokens committed per verification step, up to 3.75 on math_reasoning.

Training code: rasyosef/train-dspark-draft-models.

Trained on 100,000 samples.

Usage

vLLM loads the verifier automatically from the config — don't pass it separately.

vllm serve rasyosef/Qwen3.5-2B-DSpark --port 8000 --gpu-memory-utilization 0.8

Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.

Details

5 Qwen3 layers (hidden size 2048, intermediate size 6144, 16 attention heads over 8 KV heads, head dim 128, all layers sliding-window attention with a 2048-token window), ~0.5B params, bfloat16. Block size 8, draft vocabulary reduced to 50,000 from the verifier's 248,320 (selected by token frequency over the training data), aux hidden-state layers 1/6/11/16/21, confidence head with Markov (vanilla, rank 256).

Read the full model card (680 words)

Configuration

Architecture
DSparkDraftModel

Identity and Version

Repository
rasyosef/Qwen3.5-2B-DSpark
Publisher
Yosef Worku Alemneh
Task
Not stated by the source
Modality
Other
Library
transformers
Parameters
452M parameters
Languages
Not stated by the source
Revision
1e39250b896fdc54d5121c1f6aea6ef6d3ff9dd2
First published
2026-09-18
Last updated
2026-09-29

Files and Weights

10 files, 1.8 GB in total. The weights are 3 files totalling 1.8 GB in pt, safetensors.

Weights3 files · 1.8 GB
Configuration4 files · 5.2 KB
Documentation1 file · 5.2 KB
Other1 file · 778 B
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights903.5 MB 2d70b0eb2ef8
optimizer_state_dict.ptWeights850.9 MB bd14bde98c9f
scheduler_state_dict.ptWeights1.6 KB 7f3f1221c4ac
config.jsonConfiguration2.0 KB —
config.pyConfiguration2.3 KB —
training_state.jsonConfiguration51 B —
val_metrics.jsonConfiguration812 B —
README.mdDocumentation5.2 KB —
train_command.txtOther778 B —
.gitattributesRepository1.5 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.8 GB
Download from Yosef Worku Alemneh

Released by Yosef Worku Alemneh through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published1.8 GB
16-bit0.9 GB
8-bit0.5 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.5-2B-DSpark

How much GPU memory does Qwen3.5-2B-DSpark need?

About 1.1 GB at 16-bit and 0.3 GB at 4-bit: the weights (452M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.5-2B-DSpark on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.5-2B-DSpark commercially?

Yes. Qwen3.5-2B-DSpark is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.