Open-weight model
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound
by DoktorMincs DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound is an open-weight model from DoktorMincs, released under Apache License 2.0. It has 27.8B parameters and a 262,144-token context. At 16-bit it needs about 66.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 41 downloads a month.
The int8-family corner of a fully measured allocation frontier, rebuilt with the recipe that produced the sibling W4A16 mix — adapted to keep this repo's identity: MLP and GatedDeltaNet keep the original AutoRound int8 weights bit-exact (this repo is their…
Runs On
What it takes to serve Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound (27.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 55.6 GB | 66.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 27.8 GB | 33.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 13.9 GB | 16.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
Model Card
By DoktorMincs, published under apache-2.0, revision 7663e92725b8.
The int8-family corner of a fully measured allocation frontier, rebuilt with the recipe that produced the sibling W4A16 mix — adapted to keep this repo's identity: MLP and GatedDeltaNet keep the original AutoRound int8 weights bit-exact (this repo is their native tuning context), attention q/k/v drop BF16 → int8/g128 (measured cost ≈ 0.002 PPL for −1.8 GB), lmhead BF16 → int8/g128 (the free trick: ≈ 0.002 for −1.25 GB), oproj stays BF16. −7.2 % bytes at +0.004 PPL vs the previous revision (combined qkv-int8 + lmhead-int8 cost measured at 0.0038 — statistically at the noise floor of this protocol, ±0.0011). Same 50-text wikitext-103 protocol everywhere; sorted by PPL. Mixed rows (int4 MLP)…
Read DoktorMincs's full model card
Qwen3.8-27B TURBO — W8A16 (lean int8 family, rebuilt 2026-09-29)
The int8-family corner of a fully measured allocation frontier, rebuilt with the recipe that
produced the sibling W4A16 mix — adapted to keep this repo's identity: MLP and
GatedDeltaNet keep the original AutoRound int8 weights bit-exact (this repo is their native
tuning context), attention q/k/v drop BF16 → int8/g128 (measured cost ≈ 0.002 PPL for −1.8 GB),
lm_head BF16 → int8/g128 (the free trick: ≈ 0.002 for −1.25 GB), o_proj stays BF16.
| this build | previous revision | BF16 base | |
|---|---|---|---|
| weights | 30.860 GB (30,859,927,024 B) | 33.268 GB | 52 GB |
| wikitext-103 PPL | 8.2029 (+1.42%) | 8.1991 (+1.37%) | 8.0880 |
−7.2 % bytes at +0.004 PPL vs the previous revision (combined qkv-int8 + lm_head-int8 cost measured at 0.0038 — statistically at the noise floor of this protocol, ±0.0011).
Per-group configuration (100% Marlin uint8b128 on sm86, verified live in boot log)
| group | scheme | source |
|---|---|---|
mlp.{gate,up,down}_proj (192) |
int8 g128 sym | AutoRound-tuned (original W8 run, grafted bit-exact) |
linear_attn.{in_proj_qkv,in_proj_z,out_proj} (144) |
int8 g128 sym | AutoRound-tuned (original W8 run, grafted bit-exact) |
self_attn.{q,k,v}_proj (48) |
int8 g128 sym | RTN from BF16 original |
self_attn.o_proj (16) |
BF16 | — |
lm_head |
int8 g128 sym | RTN from BF16 |
embed, norms, linear_attn.in_proj_a/b, vision (333), MTP (15) |
BF16 | — |
The complete measured frontier of this model (one protocol, 2026-09-29)
Same 50-text wikitext-103 protocol everywhere; sorted by PPL. Mixed rows (int4 MLP) belong to the sibling W4A16 family and are published there or documented only — they dominate this repo on both axes by design choice: this repo's job is the pure int8 family.
| build | weights | PPL | vs BF16 |
|---|---|---|---|
| BF16 | 51.95 GB | 8.0880 | — |
| mix + qkv BF16 (mixed family, documented) | 23.727 GB | 8.1922 | +1.29% |
| W4A16 mix + qkv int8 (mixed optimum) | 22.571 GB | 8.1944 | +1.32% |
| W8A16 previous revision | 33.268 GB | 8.1991 | +1.37% |
| THIS BUILD — int8 family, lean | 30.860 GB | 8.2029 | +1.42% |
| mix + qkv int8, GDN int8 AR(W8-context) | 22.571 GB | 8.2417 | +1.90% |
| mix + GDN int6 RTN (qkv int4) | 20.600 GB | 8.2475 | +1.97% |
| mix + qkv int8 AR donor | 22.571 GB | 8.2675 | +2.22% |
| W6A16 previous revision | 27.605 GB | 8.2823 | +2.40% |
| W6A16 lean (sibling repo, int6 family) | 24.904 GB | 8.2838 | +2.42% |
| mix + GDN int8 AR donor | 21.984 GB | 8.2853 | +2.44% |
| mix + GDN int6 AR + qkv int8 | 21.187 GB | 8.2886 | +2.48% |
| W4A16 mix (sibling repo, live) | 21.984 GB | 8.2371 | +1.84% |
| mix + GDN int6 AR | 20.600 GB | 8.3364 | +3.07% |
| W4A16 no exclusions | 19.453 GB | 8.7945 | +8.74% |
Two findings worth more than the build itself: 1. Cross-context AutoRound is a trap. Grafting tensors tuned in a different width context (this run's GDN/qkv were tuned with MLP int8) into a build whose other groups differ costs MORE than plain RTN from BF16 at identical bytes: GDN int8 AR +0.047, qkv int8 AR +0.026, GDN int6 AR +0.089 vs RTN. Grafting wins only in the native context — which is why MLP/GDN int8 here come from THIS repo's own run. 2. q/k/v precision above int8 is worth ~nothing on this model (BF16 → int8: +0.002, at any MLP width tested). The previous revision paid 2.3 GB of BF16 qkv for a margin inside the noise floor.
Serving (validated 2× RTX 3090 TP2, vLLM ≥ 0.28)
15.4 GB/GPU of weights; validated with --tensor-parallel-size 2 --kv-cache-dtype fp8_e4m3
--max-num-seqs 32 --gpu-memory-utilization 0.95, MTP-3 speculative decoding on. For a second
3090 pair this is the largest honest int8 fit; for more KV headroom use the sibling W4A16 repo.
Protocol & reproducibility
wikitext-103 test, 50 texts × ≤512 tokens (16,007 eval tokens), prompt_logprobs=0, greedy,
vLLM 0.30 TP2. BF16 anchor 8.0880 identical in every run of this project; protocol noise floor
±0.0011 (double-measured). Build: graft_plan.py --base <W4A16 mix checkpoint> --orig <BF16>
--take mlp=int8@$DW8MLP --take gdn=int8@$DW8GDN, donors range-read from this repo's previous
revision (donor.py, 17.4 GB); deterministic, no retune run needed. Full F-grid and toolchain:
sibling W4A16 repo.
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 64
- Hidden size
- 5,120
- Feed-forward size
- 17,408
- Attention heads
- 24
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Stored precision
- bfloat16
- Model type
- qwen3_5
- Quantization
- compressed-tensors
Identity and Version
- Repository
- DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound
- Publisher
- DoktorMincs
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- 27.8B parameters
- Languages
- en
- Revision
- 7663e92725b85d39701882f994729cb7ab55705b
- First published
- 2026-09-27
- Last updated
- 2026-09-30
Files and Weights
20 files, 30.9 GB in total. The weights are 6 files totalling 30.9 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00006.safetensors | Weights | 5.3 GB | e830ee9b5497 |
| model-00002-of-00006.safetensors | Weights | 5.3 GB | be370dacf6d2 |
| model-00003-of-00006.safetensors | Weights | 5.3 GB | 7eb7d34aea1c |
| model-00004-of-00006.safetensors | Weights | 5.4 GB | 5f72f29d0d97 |
| model-00005-of-00006.safetensors | Weights | 5.4 GB | fe05e165a27f |
| model-00006-of-00006.safetensors | Weights | 4.2 GB | bac390ba3d58 |
| config.json | Configuration | 12.8 KB | — |
| generation_config.json | Configuration | 213 B | — |
| model.safetensors.index.json | Configuration | 194.9 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| processor_config.json | Configuration | 1.2 KB | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| LICENSE | Documentation | 11.4 KB | — |
| NOTICE | Documentation | 1.9 KB | — |
| README.md | Documentation | 4.6 KB | — |
| chat_template.jinja | Other | 9.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 87a7830d63fc |
| tokenizer_config.json | Tokenizer | 16.4 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 30.9 GB
Released by DoktorMincs through its official repository on Hugging Face. Read the license.
Built From
- Derived from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
- Quantized from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 30.9 GB |
| 16-bit | 55.6 GB |
| 8-bit | 27.8 GB |
| 4-bit | 13.9 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound
How much GPU memory does Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound need?
About 66.7 GB at 16-bit and 16.7 GB at 4-bit: the weights (27.8B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound commercially?
Yes. Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.