Qwen3.8-27B-TURBO-…-NM-DAU — mixed 4/8-bit, ≤22 GB, minimum perplexity
Weight-only quantization of
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU,
built to a hard 22 GB budget with the lowest perplexity achievable inside it.
This is not a uniform W4A16. Bit-width was allocated by measurement: every
candidate was quantized, served by vLLM, and scored on the same held-out corpus,
and the budget was spent where it bought the most.
- Size: 21.98 GB (base BF16 is 52 GB).
- Served by vLLM's Marlin kernels throughout — no fallback kernels.
- Text + MTP speculative decoding + vision all verified.
What is quantized
| module group |
scheme |
how the tensors were produced |
mlp.{gate,up,down}_proj (192) |
int4, group_size 64 |
round-to-nearest |
linear_attn.{in_proj_qkv,in_proj_z,out_proj} (144) |
int8, group_size 128 |
round-to-nearest |
self_attn.o_proj (16) |
BF16 (bit-identical to base) |
— |
self_attn.{q,k,v}_proj (48) |
int4, group_size 128 |
AutoRound (400 iters × 256 samples) |
lm_head |
int8, group_size 128 |
round-to-nearest |
embed_tokens, vision tower (333 tensors), MTP (15), all norms, linear_attn.in_proj_a/b, conv1d |
BF16 |
— |
Two things worth knowing before you use this
- The MLP, GatedDeltaNet and
lm_head tensors are round-to-nearest, not
AutoRound-tuned. Only q_proj/k_proj/v_proj carry AutoRound tuning (they
are inherited unchanged from an AutoRound int4 run). The measurements below
were taken on round-to-nearest tensors, so they describe exactly what is in
this repository — an AutoRound run at this same configuration would likely do
better, and that is the obvious next step.
mlp uses group_size: 64, not 128. Marlin supports group sizes
{32, 64, 128}; 64 was chosen because it measured better per byte. This makes
the MLP weights slightly larger than a group_size-128 build, and Marlin's
kernel does more scale loads per tile — expect a small throughput cost on
prefill relative to a pure g128 build.
Quality
Perplexity on 50 held-out texts from wikitext-103, 16,007 evaluated tokens —
deliberately not the calibration set. Same texts, same tokenization, same
method, both models served by vLLM. The BF16 reference is 8.0880, reproduced
identically across every run of this series, so the deltas are directly
comparable.
| version |
size |
perplexity |
delta |
| base BF16 |
52 GB |
8.0880 |
— |
| this build |
21.98 GB |
8.2371 |
+1.84% |
Against the other builds of this same model
| build |
size |
perplexity |
delta |
| W8A16, attention in BF16 |
33.3 GB |
8.1991 |
+1.37% |
| this build (mixed 4/8-bit) |
21.98 GB |
8.2371 |
+1.84% |
| W6A16, attention in BF16 |
27.6 GB |
8.2823 |
+2.40% |
| W4A16, attention + GDN in BF16 |
30.2 GB |
8.3888 |
+3.72% |
| W4A16, nothing excluded |
19.5 GB |
8.7945 |
+8.74% |
This build dominates the W6A16 and W4A16 builds: smaller and more faithful.
The W8A16 is still more faithful (+1.37%), at 52% more size.
How the budget was spent
Every row is a real vLLM measurement against the same fully-quantized 4-bit base
(8.7945), on the identical corpus. Gains are not additive — the attention
groups in particular are strongly sub-additive — so the final configuration was
measured, not computed.
Single groups
| change |
perplexity |
gain |
cost |
gain per GB |
| GatedDeltaNet group_size 128 → 32 |
8.6767 |
0.1163 |
0.260 GB |
0.447 |
o_proj → BF16 |
8.5724 |
0.2206 |
0.747 GB |
0.295 |
| whole attention → int8 |
8.5548 |
0.2382 |
0.839 GB |
0.284 |
qkv → int8 |
8.6522 |
0.1408 |
0.587 GB |
0.240 |
| whole attention → BF16 |
8.5524 |
0.2406 |
2.490 GB |
0.097 |
qkv → BF16 |
8.6509 |
0.1421 |
1.743 GB |
0.082 |
| GatedDeltaNet → int8 |
8.6286 |
0.1644 |
2.768 GB |
0.059 |
| GatedDeltaNet → int6 |
8.7218 |
0.0712 |
1.384 GB |
0.051 |
| MLP → int8 |
8.6117 |
0.1813 |
8.556 GB |
0.021 |
| MLP group_size 128 → 64 |
8.7711 |
0.0219 |
0.267 GB |
0.082 |
| MLP group_size 128 → 32 |
8.8337 |
−0.0407 |
0.802 GB |
— |
Combined configurations that fit the budget
| configuration |
perplexity |
delta |
size |
chosen: o_proj BF16 + GDN int8 + MLP g64 |
8.2371 |
+1.84% |
21.98 GB |
| attention BF16 + MLP g32 + GDN g32 |
8.2796 |
+2.37% |
21.75 GB |
attention int8 + o_proj BF16 + GDN g32 + MLP g32 |
8.2832 |
+2.41% |
20.60 GB |
o_proj BF16 + qkv int8 + GDN int6 + MLP g32 |
8.3157 |
+2.82% |
21.72 GB |
| attention int8 + GDN int8 |
8.3956 |
+3.80% |
21.81 GB |
attention int8 + o_proj BF16 + GDN g32 |
8.4731 |
+4.76% |
19.80 GB |
(over budget) o_proj BF16 + qkv int8 + GDN int8 + MLP g64 |
8.1930 |
+1.30% |
22.57 GB |
What the measurements showed
- int8 dominates BF16 in the attention.
attn int8 gets 99% of attn BF16's
gain for 34% of the bytes. Putting BF16 on qkv is the single worst use of
budget measured here.
o_proj carries ~92% of the attention damage. Alone (0.747 GB) it gains
0.2206, while all four attention projections in BF16 (2.490 GB) gain only
0.2406. This reproduces, in the 4-bit context, what the 8-bit sweep had hinted
at — and it is not portable between bit-widths in general.
- Interactions are strong in both directions. Adding
o_proj BF16 on top of
attention int8 gains only 0.0020 — the int8 already fixed it. Conversely, MLP
group_size 32 hurts when applied alone (−0.0407) but helps by 0.1755
once attention and GatedDeltaNet are corrected: with the larger errors removed,
the MLP's own error becomes the binding constraint.
- GatedDeltaNet int6 is dominated by int8 on both counts — it gains less per
byte (0.051 vs 0.059) and it drops off Marlin onto the JIT-compiled Humming
kernel. It is not used here.
Serving
At 21.98 GB this fits a single 24 GB card, but with very little left for KV cache. It is
sized to be served on 2× RTX 3090 (24 GB, NVLink) with tensor parallelism, which is the
configuration it was validated on: ~90-100 tok/s decode (with the pelican svg prompt) and a KV cache of ~490k tokens.
config.json declares four config_groups: two added for this build (MLP at group_size 64,
GatedDeltaNet at int8) and two inherited from the AutoRound base build (lm_head at int8, and the
generic Linear int4 fallback). vLLM resolves a scheme per layer by name, in config order — the
re:.*lm_head$ target is listed before the generic Linear target, and that order is
load-bearing: the generic group must stay last, as the fallback for everything not matched by a
specific rule.
lm_head is quantized, and the MTP draft head shares the lm_head.weight key, so both take the
same scheme. (See the MTP caveat below: with speculative decoding enabled the generated text is
coherent and the final answers are equivalent, but it does not reproduce the non-speculative
output token-for-token.)
Example command
vllm serve /root/models/DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound \
--served-model-name "qwen3.6-27b" \
--tensor-parallel-size 2 \
--pipeline-parallel-size 1 \
--max-model-len auto \
--override-generation-config '{"temperature": 0.7, "top_p": 0.8, "top_k": 20, "min_p": 0.0, "presence_penalty": 1.5, "repetition_penalty": 1.0}' \
--seed 1234 \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-batched-tokens 2048 \
--trust-remote-code \
--disable-custom-all-reduce \
--gpu-memory-utilization 0.94 \
--max-num-seqs 32 \
--host 0.0.0.0 \
--port 3434 \
--kv-cache-dtype fp8_e4m3 \
--default-chat-template-kwargs '{"enable_thinking": true, "reasoning_effort": "xhigh"}' \
--performance-mode balanced \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
Sizing the context
Only the 16 full-attention layers cache K and V. The other 48 are GatedDeltaNet, which carries a
constant-size recurrent state and grows no KV at all — so the cache is unusually cheap for a 27B
model:
16 layers × 4 KV heads × 256 head_dim × 2 (K and V) × 1 byte (fp8_e4m3) = 32 KiB per token
~490k tokens × 32 KiB ≈ 16 GB
That is why --max-num-seqs 32 with --gpu-memory-utilization 0.94 lands near half a million
tokens of context. Raise --max-num-seqs for more concurrency at the cost of per-sequence context,
and --max-model-len to cap it explicitly instead of auto.
--kv-cache-dtype fp8_e4m3 halves the cache versus bf16 and is what the numbers above assume;
dropping it doubles the per-token cost to 64 KiB and roughly halves the reachable context.
Caveats
- Round-to-nearest MLP/GatedDeltaNet/
lm_head tensors (see above) — measured as
shipped, but not optimal.
- The
+1.84% number is text-only perplexity. Vision and MTP were validated
functionally (a colour-identification prompt and speculative-decoding
acceptance) but not scored on a vision benchmark.
- Asymmetric weight quantization was not used: it was not implemented for this
build (it needs
weight_zero_point tensor support, and its value was never
measured). group_size 32/64 and int8 were used instead.
- MTP speculative decoding does not reproduce greedy output token-for-token on
this model + vLLM combination. Outputs are coherent in both modes but diverge
partway through long chain-of-thought generations. This is pre-existing and
not specific to this build: it reproduces on the unmodified 4-bit base and on
the previously published builds of this model (first divergence at 313 and 696
characters for the base, 1243 and 171 for the earlier W4A16 build, 131 and 437
here). Treat MTP here as a throughput feature, not a bit-exactness guarantee.
Tokenizer warning — safe to ignore, and do not "fix" it
Some transformers versions log this at load time:
The tokenizer you are loading from … with an incorrect regex pattern … This will lead to
incorrect tokenization. You should set the fix_mistral_regex=True flag …
It is a false positive, and following its advice would degrade this model. The tokenizer uses
the standard GPT-2/Qwen pre-tokenizer pattern — the one this model was trained with. The warning
comes from a heuristic that cannot tell which model family a tokenizer belongs to: it reads
transformers_version from config.json and, when that field is missing, assumes the buggy
Mistral pattern. This artifact now carries the field, so the warning no longer appears. (It only
ever fired for local-path loads, i.e. local_files_only=True; loading from the Hub never
reaches that code path for a non-Mistral model.)
fix_mistral_regex=True replaces the pattern with Mistral's, which splits camelCase apart.
Measured on this tokenizer:
| input |
default (correct for Qwen) |
with fix_mistral_regex=True |
iPhone macOS camelCaseWord |
iPhone · macOS · camel · Case · Word |
i · Phone · mac · OS · camel · Case · Word |
The difference is subtle — on ordinary prose, numbers, punctuation and precomposed accents the two
patterns agree exactly. It shows up on identifiers, brand names and camelCase, which is precisely
where you would not want it.
Reproducing
Built with graft_plan.py from a fully AutoRound-quantized int4 base, then two
groups re-quantized round-to-nearest and one restored to BF16 — each step
validated by dequantization round-trip (mean error ≈ quantisation step / 4):
--take oprop=bf16 --take gdn=int8 --take mlp=int4/g64
32 candidate configurations were served by vLLM and scored on the same corpus
before this one was chosen; the tables above are those measurements.