Open-weight model
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound
by DoktorMincs DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound is an open-weight model from DoktorMincs, released under Apache License 2.0. It has 27.8B parameters and a 262,144-token context. At 16-bit it needs about 66.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 58 downloads a month.
The int6-family corner of a fully measured frontier. MLP and GatedDeltaNet keep this repo's original int6 AutoRound weights bit-exact (native tuning context); attention q/k/v drop BF16 → int6/g128 RTN and lmhead BF16 → int8/g128 (−2.7 GB combined, cost inside…
Runs On
What it takes to serve Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound (27.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 55.6 GB | 66.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 27.8 GB | 33.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 13.9 GB | 16.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
Model Card
By DoktorMincs, published under apache-2.0, revision 3c96cecb31a1.
The int6-family corner of a fully measured frontier. MLP and GatedDeltaNet keep this repo's original int6 AutoRound weights bit-exact (native tuning context); attention q/k/v drop BF16 → int6/g128 RTN and lmhead BF16 → int8/g128 (−2.7 GB combined, cost inside the noise floor); oproj stays BF16. −9.8 % bytes at +0.0015 PPL — statistically the same model, 2.7 GB lighter. Kernel note: int6 is not a Marlin dtype on sm86 — mlp/gdn/qkv load through HummingLinearKernel (JIT; lmhead int8 via Marlin). Verified live on vLLM 0.30 / 2× RTX 3090 (text + MTP + vision; first boot pays JIT compile time). The finding this repo exists to show: plain RTN from BF16 with int6 GDN (row 7, 8.2475 at 20.6 GB — a…
Read DoktorMincs's full model card
Qwen3.8-27B TURBO — W6A16 (lean int6 family, rebuilt 2026-09-29)
The int6-family corner of a fully measured frontier. MLP and GatedDeltaNet keep this repo's
original int6 AutoRound weights bit-exact (native tuning context); attention q/k/v drop
BF16 → int6/g128 RTN and lm_head BF16 → int8/g128 (−2.7 GB combined, cost inside the noise
floor); o_proj stays BF16.
| this build | previous revision | BF16 base | |
|---|---|---|---|
| weights | 24.904 GB (24,904,015,344 B) | 27.605 GB | 52 GB |
| wikitext-103 PPL | 8.2838 (+2.42%) | 8.2823 (+2.40%) | 8.0880 |
−9.8 % bytes at +0.0015 PPL — statistically the same model, 2.7 GB lighter.
Per-group configuration
| group | scheme | source |
|---|---|---|
mlp.{gate,up,down}_proj (192) |
int6 g128 sym | AutoRound-tuned (original W6 run, grafted bit-exact) |
linear_attn.{in_proj_qkv,in_proj_z,out_proj} (144) |
int6 g128 sym | AutoRound-tuned (original W6 run, grafted bit-exact) |
self_attn.{q,k,v}_proj (48) |
int6 g128 sym | RTN from BF16 original |
self_attn.o_proj (16) |
BF16 | — |
lm_head |
int8 g128 sym | RTN from BF16 |
embed, norms, linear_attn.in_proj_a/b, vision (333), MTP (15) |
BF16 | — |
Kernel note: int6 is not a Marlin dtype on sm86 — mlp/gdn/qkv load through
HummingLinearKernel (JIT; lm_head int8 via Marlin). Verified live on vLLM 0.30 / 2× RTX 3090
(text + MTP + vision; first boot pays JIT compile time).
The complete measured frontier of this model (one protocol, 2026-09-29)
| build | weights | PPL | vs BF16 |
|---|---|---|---|
| BF16 | 51.95 GB | 8.0880 | — |
| mix + qkv BF16 (mixed family, documented) | 23.727 GB | 8.1922 | +1.29% |
| W4A16 mix + qkv int8 (mixed optimum) | 22.571 GB | 8.1944 | +1.32% |
| W8A16 previous revision | 33.268 GB | 8.1991 | +1.37% |
| W8A16 lean (sibling repo, int8 family) | 30.860 GB | 8.2029 | +1.42% |
| W4A16 mix (sibling repo, live) | 21.984 GB | 8.2371 | +1.84% |
| mix + qkv int8, GDN int8 AR(W8-context) | 22.571 GB | 8.2417 | +1.90% |
| mix + GDN int6 RTN (qkv int4) | 20.600 GB | 8.2475 | +1.97% |
| mix + qkv int8 AR donor | 22.571 GB | 8.2675 | +2.22% |
| W6A16 previous revision | 27.605 GB | 8.2823 | +2.40% |
| THIS BUILD — int6 family, lean | 24.904 GB | 8.2838 | +2.42% |
| mix + GDN int8 AR donor | 21.984 GB | 8.2853 | +2.44% |
| mix + GDN int6 AR + qkv int8 | 21.187 GB | 8.2886 | +2.48% |
| mix + GDN int6 AR | 20.600 GB | 8.3364 | +3.07% |
| W4A16 no exclusions | 19.453 GB | 8.7945 | +8.74% |
The finding this repo exists to show: plain RTN from BF16 with int6 GDN (row 7, 8.2475 at 20.6 GB — a mixed build) beats this family's own cross-context AutoRound int6 GDN by 0.089 PPL (row 14, same bytes). int6-tuned-in-int6-context wins only where the rest of the build shares that context — which is exactly the MLP/GDN pairing of this build. Cross-context grafting is a trap; measure before you inherit.
Serving
Validated on 2× RTX 3090 (TP2): vLLM ≥ 0.28, --kv-cache-dtype fp8_e4m3 --max-num-seqs 32,
MTP-3 speculative decoding works. At 12.5 GB/GPU this build leaves the most KV headroom of the
family trio on the same hardware.
Protocol & reproducibility
wikitext-103 test, 50 texts × ≤512 tokens (16,007 eval tokens), prompt_logprobs=0, greedy,
vLLM 0.30 TP2; BF16 anchor 8.0880 constant across the project; noise floor ±0.0011
(double-measured). Build: graft_plan.py --base <F1b = W4A16 mix + GDN int6 RTN> --orig <BF16>
--take mlp=int6@$DW6MLP --take gdn=int6@$DW6GDN --take qkv=int6; donors (13.1 + 4.2 GB)
range-read from this repo's previous revision via donor.py. Deterministic, no retune run.
Full F-grid and toolchain: sibling W4A16 repo.
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 64
- Hidden size
- 5,120
- Feed-forward size
- 17,408
- Attention heads
- 24
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Stored precision
- bfloat16
- Model type
- qwen3_5
- Quantization
- compressed-tensors
Identity and Version
- Repository
- DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound
- Publisher
- DoktorMincs
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- 27.8B parameters
- Languages
- en
- Revision
- 3c96cecb31a1fe222051a6606db6bbe7d896133b
- First published
- 2026-09-27
- Last updated
- 2026-09-30
Files and Weights
19 files, 24.9 GB in total. The weights are 5 files totalling 24.9 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00005.safetensors | Weights | 5.4 GB | f161591aacf6 |
| model-00002-of-00005.safetensors | Weights | 5.3 GB | e639f10c3003 |
| model-00003-of-00005.safetensors | Weights | 5.4 GB | f6441e372ebc |
| model-00004-of-00005.safetensors | Weights | 5.4 GB | 9139033d7f4b |
| model-00005-of-00005.safetensors | Weights | 3.5 GB | c0b5601b2f01 |
| config.json | Configuration | 12.2 KB | — |
| generation_config.json | Configuration | 213 B | — |
| model.safetensors.index.json | Configuration | 194.9 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| processor_config.json | Configuration | 1.2 KB | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| LICENSE | Documentation | 11.4 KB | — |
| NOTICE | Documentation | 1.9 KB | — |
| README.md | Documentation | 3.9 KB | — |
| chat_template.jinja | Other | 9.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 87a7830d63fc |
| tokenizer_config.json | Tokenizer | 16.4 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 24.9 GB
Released by DoktorMincs through its official repository on Hugging Face. Read the license.
Built From
- Derived from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
- Quantized from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 24.9 GB |
| 16-bit | 55.6 GB |
| 8-bit | 27.8 GB |
| 4-bit | 13.9 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound
How much GPU memory does Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound need?
About 66.7 GB at 16-bit and 16.7 GB at 4-bit: the weights (27.8B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound commercially?
Yes. Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W6A16-AutoRound's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.