# Qwen3.5-2B-DSpark by Yosef Worku Alemneh: Open-Weight Model
Source: https://savrn.com/models/qwen3-5-2b-dspark
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve Qwen3.5-2B-DSpark (452M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 0.9 GB | 1.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 0.5 GB | 0.5 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 0.2 GB | 0.3 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 8, 2026.

[Qwen3.5-2B-DSpark on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/qwen3-5-2b-dspark/gpus)

## Model Card

By Yosef Worku Alemneh, published under apache-2.0, revision 1e39250b896f.

A DSpark draft model for speculative decoding with [Qwen/Qwen3.5-2B](https://savrn.com/models/qwen3-5-2b) as the verifier, trained with [speculators](https://github.com/vllm-project/speculators). The drafter proposes 8 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone — a lossless speedup. Mean acceptance length is 2.50 tokens committed per verification step, up to 3.75 on math_reasoning.

Training code: [rasyosef/train-dspark-draft-models](https://github.com/rasyosef/train-dspark-draft-models).

Trained on 100,000 samples.

### Usage

vLLM loads the verifier automatically from the config — don't pass it separately.

```
vllm serve rasyosef/Qwen3.5-2B-DSpark --port 8000 --gpu-memory-utilization 0.8
```

Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.

### Details

5 Qwen3 layers (hidden size 2048, intermediate size 6144, 16 attention heads over 8 KV heads, head dim 128, all layers sliding-window attention with a 2048-token window), ~0.5B params, bfloat16. Block size 8, draft vocabulary reduced to 50,000 from the verifier's 248,320 (selected by token frequency over the training data), aux hidden-state layers 1/6/11/16/21, confidence head with Markov (vanilla, rank 256).

[Read the full model card (680 words)](https://savrn.com/models/qwen3-5-2b-dspark/card)

## Configuration

Architecture

DSparkDraftModel

## Identity and Version

Repository

rasyosef/Qwen3.5-2B-DSpark

Publisher

Yosef Worku Alemneh

Task

Not stated by the source

Modality

Other

Library

transformers

Parameters

452M parameters

Languages

Not stated by the source

Revision

1e39250b896fdc54d5121c1f6aea6ef6d3ff9dd2

First published

2026-09-18

Last updated

2026-09-29

## Files and Weights

10 files, 1.8 GB in total. The weights are 3 files totalling 1.8 GB in pt, safetensors.

Weights3 files · 1.8 GB

Configuration4 files · 5.2 KB

Documentation1 file · 5.2 KB

Other1 file · 778 B

Repository1 file · 1.5 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model.safetensors | Weights | 903.5 MB | 2d70b0eb2ef8 |
| optimizer_state_dict.pt | Weights | 850.9 MB | bd14bde98c9f |
| scheduler_state_dict.pt | Weights | 1.6 KB | 7f3f1221c4ac |
| config.json | Configuration | 2.0 KB | — |
| config.py | Configuration | 2.3 KB | — |
| training_state.json | Configuration | 51 B | — |
| val_metrics.json | Configuration | 812 B | — |
| README.md | Documentation | 5.2 KB | — |
| train_command.txt | Other | 778 B | — |
| .gitattributes | Repository | 1.5 KB | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

1.8 GB

[Download from Yosef Worku Alemneh](https://huggingface.co/rasyosef/Qwen3.5-2B-DSpark)

Released by Yosef Worku Alemneh through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from [Qwen/Qwen3.5-2B](https://savrn.com/models/qwen3-5-2b)

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 1.8 GB |
| 16-bit | 0.9 GB |
| 8-bit | 0.5 GB |
| 4-bit | 0.2 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About Qwen3.5-2B-DSpark

### How much GPU memory does Qwen3.5-2B-DSpark need?

About 1.1 GB at 16-bit and 0.3 GB at 4-bit: the weights (452M parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run Qwen3.5-2B-DSpark on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use Qwen3.5-2B-DSpark commercially?

Yes. Qwen3.5-2B-DSpark is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

## Yosef Worku Alemneh

[All models and datasets](https://savrn.com/model-publishers/rasyosef)

## Versions

- [1e39250b896f](https://savrn.com/models/qwen3-5-2b-dspark/versions/1e39250b896f) · current 2026-09-29

## Explore More

- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-09-29.
- [Hugging Face record](https://huggingface.co/rasyosef/Qwen3.5-2B-DSpark)
- [How the hub is built](https://savrn.com/model-hub/methodology)
