The int6-family corner of a fully measured frontier. MLP and GatedDeltaNet keep this repo's original int6 AutoRound weights bit-exact (native tuning context); attention q/k/v drop BF16 → int6/g128 RTN and lmhead BF16 → int8/g128 (−2.7 GB combined, cost inside the noise floor); oproj stays BF16. −9.8 % bytes at +0.0015 PPL — statistically the same model, 2.7 GB lighter. Kernel note: int6 is not a Marlin dtype on sm86 — mlp/gdn/qkv load through HummingLinearKernel (JIT; lmhead int8 via Marlin). Verified live on vLLM 0.30 / 2× RTX 3090 (text + MTP + vision; first boot pays JIT compile time). The finding this repo exists to show: plain RTN from BF16 with int6 GDN (row 7, 8.2475 at 20.6 GB — a…
Independent publisher
DoktorMincs
DoktorMincs
Models
The int8-family corner of a fully measured allocation frontier, rebuilt with the recipe that produced the sibling W4A16 mix — adapted to keep this repo's identity: MLP and GatedDeltaNet keep the original AutoRound int8 weights bit-exact (this repo is their native tuning context), attention q/k/v drop BF16 → int8/g128 (measured cost ≈ 0.002 PPL for −1.8 GB), lmhead BF16 → int8/g128 (the free trick: ≈ 0.002 for −1.25 GB), oproj stays BF16. −7.2 % bytes at +0.004 PPL vs the previous revision (combined qkv-int8 + lmhead-int8 cost measured at 0.0038 — statistically at the noise floor of this protocol, ±0.0011). Same 50-text wikitext-103 protocol everywhere; sorted by PPL. Mixed rows (int4 MLP)…
Model · Image and text to text
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound
Weight-only quantization of built to a hard 22 GB budget with the lowest perplexity achievable inside it. This is not a uniform W4A16. Bit-width was allocated by measurement: every candidate was quantized, served by vLLM, and scored on the same held-out corpus, and the budget was spent where it bought the most. - Served by vLLM's Marlin kernels throughout — no fallback kernels. - Text + MTP speculative decoding + vision all verified. 1. The MLP, GatedDeltaNet and lmhead tensors are round-to-nearest, not AutoRound-tuned. Only qproj/kproj/vproj carry AutoRound tuning (they are inherited unchanged from an AutoRound int4 run). The measurements below were taken on round-to-nearest tensors, so…