SAVRN
Search Contact SAVRN

Qwen3.5-2B-DSpark · Model Card

Qwen3.5-2B-DSpark: Model Card

Written by Yosef Worku Alemneh, published under apache-2.0, revision 1e39250b896f, read 2026-09-29. Shown as written; SAVRN's own facts about this model are on its page.

A DSpark draft model for speculative decoding with Qwen/Qwen3.5-2B as the verifier, trained with speculators. The drafter proposes 8 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone — a lossless speedup. Mean acceptance length is 2.50 tokens committed per verification step, up to 3.75 on math_reasoning.

Training code: rasyosef/train-dspark-draft-models.

Trained on 100,000 samples.

Usage

vLLM loads the verifier automatically from the config — don't pass it separately.

vllm serve rasyosef/Qwen3.5-2B-DSpark --port 8000 --gpu-memory-utilization 0.8

Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.

Details

5 Qwen3 layers (hidden size 2048, intermediate size 6144, 16 attention heads over 8 KV heads, head dim 128, all layers sliding-window attention with a 2048-token window), ~0.5B params, bfloat16. Block size 8, draft vocabulary reduced to 50,000 from the verifier's 248,320 (selected by token frequency over the training data), aux hidden-state layers 1/6/11/16/21, confidence head with Markov (vanilla, rank 256).

Trained for 4 epochs at lr 4e-4 (AdamW, cosine schedule with 4% warmup) on 100,000 filtered Open PerfectBlend prompts with responses regenerated by the verifier itself, split 96/4 into train and validation, with a {"ce": 0.1, "tv": 0.9} loss. Samples prepared at 2048 tokens; training sequence length 8192, up to 1024 anchors per sample. Verifier hidden states were pulled on demand from a running vLLM server during training and deleted after use rather than staged to disk up front.

Evaluation

evaluate.py throughput across the nine RedHatAI/speculator_benchmarks subsets. acceptance_length is mean tokens committed per verification step, including the bonus token — floor 1.0, ceiling 9.0 at block size 8.

subset acceptance_length pos_0 pos_1 pos_2 pos_3 pos_4 pos_5 pos_6 pos_7
math_reasoning 3.752 71.3% 53.4% 41.2% 32.6% 26.3% 20.8% 16.5% 13.0%
HumanEval 2.865 61.5% 40.6% 28.3% 20.0% 14.3% 10.1% 7.0% 4.7%
question 2.343 51.3% 29.1% 18.6% 12.8% 9.1% 6.2% 4.3% 3.1%
writing 2.322 51.9% 28.8% 18.1% 12.2% 8.6% 5.9% 4.0% 2.6%
tool_call 2.223 50.0% 28.0% 16.8% 10.9% 7.2% 4.7% 3.0% 1.9%
rag 2.219 51.8% 28.7% 17.5% 10.5% 6.5% 3.8% 2.0% 1.0%
translation 2.042 51.8% 28.0% 13.7% 6.2% 2.5% 1.3% 0.6% 0.1%
qa 1.858 43.0% 20.2% 10.4% 5.8% 3.3% 1.7% 0.9% 0.4%
summarization 1.852 44.5% 20.3% 10.1% 5.2% 2.9% 1.3% 0.6% 0.2%

Weighted across all subsets: 2.499 over 226,698 verification steps (up from 2.306 for the earlier 50,000-sample checkpoint, with every subset improving).

Acceptance is highest where the verifier's next token is most predictable. math_reasoning is the only subset above 3.0 and still accepts more than one draft in eight at pos_7 (13.0%); HumanEval follows at 2.86 and stays above 10% through pos_5. The next four — question, writing, tool_call and rag — sit in a tight 2.22–2.34 band, with tool_call (2.22) offering no real advantage over open-ended writing (2.32) or retrieval-grounded answering (2.22), so the structure of a tool call is apparently not buying much predictability here. translation (2.04) starts as strongly as that middle group at pos_0 (51.8%) but drops off sharply after pos_2; it is also the smallest subset (1,427 steps), so treat its number as noisy. qa and summarization are closest to the floor at about 1.85, accepting under 6% from pos_3 onward, so on that traffic the drafter contributes a little under one extra token per step.

First-position acceptance spans 43–71% across all nine subsets, and only math_reasoning and HumanEval remain above 10% at pos_4, so most of an 8-token block goes unused outside math and code.

Limitations

Works only with Qwen3.5-2B and is not usable as a standalone model. The verifier is vision-language, but the drafter was trained on text-only prompts and was evaluated on text-only benchmarks, so acceptance on image inputs is untested. Acceptance falls off steeply past the first few positions on most traffic, so a block size of 8 is largely wasted outside math_reasoning and code — a shorter block would capture most of the same gain at lower drafting cost. Real-world speedup depends on your traffic mix, and because verification is lossless, the verifier's own behavior and biases carry through unchanged.

License

Apache 2.0, inherited from the verifier. The speculators training code is Apache-2.0.