SAVRN
Search Contact SAVRN

Open-weight model

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound

by DoktorMincs DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound is an open-weight model from DoktorMincs, released under Apache License 2.0. It has 27.8B parameters and a 262,144-token context. At 16-bit it needs about 66.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 41 downloads a month.

The int8-family corner of a fully measured allocation frontier, rebuilt with the recipe that produced the sibling W4A16 mix — adapted to keep this repo's identity: MLP and GatedDeltaNet keep the original AutoRound int8 weights bit-exact (this repo is their…

Parameters27.8B
Context262,144
Weights30.9 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads41

Runs On

What it takes to serve Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound (27.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 55.6 GB 66.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.8 GB 33.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.9 GB 16.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound on every accelerator the SAVRN Index prices, at every precision

Model Card

By DoktorMincs, published under apache-2.0, revision 7663e92725b8.

The int8-family corner of a fully measured allocation frontier, rebuilt with the recipe that produced the sibling W4A16 mix — adapted to keep this repo's identity: MLP and GatedDeltaNet keep the original AutoRound int8 weights bit-exact (this repo is their native tuning context), attention q/k/v drop BF16 → int8/g128 (measured cost ≈ 0.002 PPL for −1.8 GB), lmhead BF16 → int8/g128 (the free trick: ≈ 0.002 for −1.25 GB), oproj stays BF16. −7.2 % bytes at +0.004 PPL vs the previous revision (combined qkv-int8 + lmhead-int8 cost measured at 0.0038 — statistically at the noise floor of this protocol, ±0.0011). Same 50-text wikitext-103 protocol everywhere; sorted by PPL. Mixed rows (int4 MLP)…

Read DoktorMincs's full model card

Qwen3.8-27B TURBO — W8A16 (lean int8 family, rebuilt 2026-09-29)

The int8-family corner of a fully measured allocation frontier, rebuilt with the recipe that produced the sibling W4A16 mix — adapted to keep this repo's identity: MLP and GatedDeltaNet keep the original AutoRound int8 weights bit-exact (this repo is their native tuning context), attention q/k/v drop BF16 → int8/g128 (measured cost ≈ 0.002 PPL for −1.8 GB), lm_head BF16 → int8/g128 (the free trick: ≈ 0.002 for −1.25 GB), o_proj stays BF16.

this build previous revision BF16 base
weights 30.860 GB (30,859,927,024 B) 33.268 GB 52 GB
wikitext-103 PPL 8.2029 (+1.42%) 8.1991 (+1.37%) 8.0880

−7.2 % bytes at +0.004 PPL vs the previous revision (combined qkv-int8 + lm_head-int8 cost measured at 0.0038 — statistically at the noise floor of this protocol, ±0.0011).

Per-group configuration (100% Marlin uint8b128 on sm86, verified live in boot log)

group scheme source
mlp.{gate,up,down}_proj (192) int8 g128 sym AutoRound-tuned (original W8 run, grafted bit-exact)
linear_attn.{in_proj_qkv,in_proj_z,out_proj} (144) int8 g128 sym AutoRound-tuned (original W8 run, grafted bit-exact)
self_attn.{q,k,v}_proj (48) int8 g128 sym RTN from BF16 original
self_attn.o_proj (16) BF16 —
lm_head int8 g128 sym RTN from BF16
embed, norms, linear_attn.in_proj_a/b, vision (333), MTP (15) BF16 —

The complete measured frontier of this model (one protocol, 2026-09-29)

Same 50-text wikitext-103 protocol everywhere; sorted by PPL. Mixed rows (int4 MLP) belong to the sibling W4A16 family and are published there or documented only — they dominate this repo on both axes by design choice: this repo's job is the pure int8 family.

build weights PPL vs BF16
BF16 51.95 GB 8.0880 —
mix + qkv BF16 (mixed family, documented) 23.727 GB 8.1922 +1.29%
W4A16 mix + qkv int8 (mixed optimum) 22.571 GB 8.1944 +1.32%
W8A16 previous revision 33.268 GB 8.1991 +1.37%
THIS BUILD — int8 family, lean 30.860 GB 8.2029 +1.42%
mix + qkv int8, GDN int8 AR(W8-context) 22.571 GB 8.2417 +1.90%
mix + GDN int6 RTN (qkv int4) 20.600 GB 8.2475 +1.97%
mix + qkv int8 AR donor 22.571 GB 8.2675 +2.22%
W6A16 previous revision 27.605 GB 8.2823 +2.40%
W6A16 lean (sibling repo, int6 family) 24.904 GB 8.2838 +2.42%
mix + GDN int8 AR donor 21.984 GB 8.2853 +2.44%
mix + GDN int6 AR + qkv int8 21.187 GB 8.2886 +2.48%
W4A16 mix (sibling repo, live) 21.984 GB 8.2371 +1.84%
mix + GDN int6 AR 20.600 GB 8.3364 +3.07%
W4A16 no exclusions 19.453 GB 8.7945 +8.74%

Two findings worth more than the build itself: 1. Cross-context AutoRound is a trap. Grafting tensors tuned in a different width context (this run's GDN/qkv were tuned with MLP int8) into a build whose other groups differ costs MORE than plain RTN from BF16 at identical bytes: GDN int8 AR +0.047, qkv int8 AR +0.026, GDN int6 AR +0.089 vs RTN. Grafting wins only in the native context — which is why MLP/GDN int8 here come from THIS repo's own run. 2. q/k/v precision above int8 is worth ~nothing on this model (BF16 → int8: +0.002, at any MLP width tested). The previous revision paid 2.3 GB of BF16 qkv for a margin inside the noise floor.

Serving (validated 2× RTX 3090 TP2, vLLM ≥ 0.28)

15.4 GB/GPU of weights; validated with --tensor-parallel-size 2 --kv-cache-dtype fp8_e4m3 --max-num-seqs 32 --gpu-memory-utilization 0.95, MTP-3 speculative decoding on. For a second 3090 pair this is the largest honest int8 fit; for more KV headroom use the sibling W4A16 repo.

Protocol & reproducibility

wikitext-103 test, 50 texts × ≤512 tokens (16,007 eval tokens), prompt_logprobs=0, greedy, vLLM 0.30 TP2. BF16 anchor 8.0880 identical in every run of this project; protocol noise floor ±0.0011 (double-measured). Build: graft_plan.py --base <W4A16 mix checkpoint> --orig <BF16> --take mlp=int8@$DW8MLP --take gdn=int8@$DW8GDN, donors range-read from this repo's previous revision (donor.py, 17.4 GB); deterministic, no retune run needed. Full F-grid and toolchain: sibling W4A16 repo.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Stored precision
bfloat16
Model type
qwen3_5
Quantization
compressed-tensors

Identity and Version

Repository
DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound
Publisher
DoktorMincs
Task
Not stated by the source
Modality
Other
Library
Not stated by the source
Parameters
27.8B parameters
Languages
en
Revision
7663e92725b85d39701882f994729cb7ab55705b
First published
2026-09-27
Last updated
2026-09-30

Files and Weights

20 files, 30.9 GB in total. The weights are 6 files totalling 30.9 GB in safetensors.

Weights6 files · 30.9 GB
Configuration6 files · 209.9 KB
Tokenizer3 files · 26.7 MB
Documentation3 files · 17.8 KB
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00006.safetensorsWeights5.3 GB e830ee9b5497
model-00002-of-00006.safetensorsWeights5.3 GB be370dacf6d2
model-00003-of-00006.safetensorsWeights5.3 GB 7eb7d34aea1c
model-00004-of-00006.safetensorsWeights5.4 GB 5f72f29d0d97
model-00005-of-00006.safetensorsWeights5.4 GB fe05e165a27f
model-00006-of-00006.safetensorsWeights4.2 GB bac390ba3d58
config.jsonConfiguration12.8 KB —
generation_config.jsonConfiguration213 B —
model.safetensors.index.jsonConfiguration194.9 KB —
preprocessor_config.jsonConfiguration390 B —
processor_config.jsonConfiguration1.2 KB —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation11.4 KB —
NOTICEDocumentation1.9 KB —
README.mdDocumentation4.6 KB —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 87a7830d63fc
tokenizer_config.jsonTokenizer16.4 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
30.9 GB
Download from DoktorMincs

Released by DoktorMincs through its official repository on Hugging Face. Read the license.

Built From

  • Derived from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
  • Quantized from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU

Memory Requirements

PrecisionWeights in memory
As published30.9 GB
16-bit55.6 GB
8-bit27.8 GB
4-bit13.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound

How much GPU memory does Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound need?

About 66.7 GB at 16-bit and 16.7 GB at 4-bit: the weights (27.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound commercially?

Yes. Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W8A16-AutoRound's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.