SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound

by DoktorMincs DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound is an open-weight model for image and text to text from DoktorMincs, released under Apache License 2.0. It has 33.3B parameters and a 262,144-token context. At 16-bit it needs about 80 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Weight-only quantization of built to a hard 22 GB budget with the lowest perplexity achievable inside it. This is not a uniform W4A16.

Parameters33.3B
Context262,144
Weights22.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound (33.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 66.6 GB 80.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 33.3 GB 40.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 16.7 GB 20.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound on every accelerator the SAVRN Index prices, at every precision

Model Card

By DoktorMincs, published under apache-2.0, revision 93b7ff2da1ea.

Weight-only quantization of built to a hard 22 GB budget with the lowest perplexity achievable inside it. This is not a uniform W4A16. Bit-width was allocated by measurement: every candidate was quantized, served by vLLM, and scored on the same held-out corpus, and the budget was spent where it bought the most. - Served by vLLM's Marlin kernels throughout — no fallback kernels. - Text + MTP speculative decoding + vision all verified. 1. The MLP, GatedDeltaNet and lmhead tensors are round-to-nearest, not AutoRound-tuned. Only qproj/kproj/vproj carry AutoRound tuning (they are inherited unchanged from an AutoRound int4 run). The measurements below were taken on round-to-nearest tensors, so…

Read DoktorMincs's full model card

Qwen3.8-27B-TURBO-…-NM-DAU — mixed 4/8-bit, ≤22 GB, minimum perplexity

Weight-only quantization of DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU, built to a hard 22 GB budget with the lowest perplexity achievable inside it.

This is not a uniform W4A16. Bit-width was allocated by measurement: every candidate was quantized, served by vLLM, and scored on the same held-out corpus, and the budget was spent where it bought the most.

  • Size: 21.98 GB (base BF16 is 52 GB).
  • Served by vLLM's Marlin kernels throughout — no fallback kernels.
  • Text + MTP speculative decoding + vision all verified.

What is quantized

module group scheme how the tensors were produced
mlp.{gate,up,down}_proj (192) int4, group_size 64 round-to-nearest
linear_attn.{in_proj_qkv,in_proj_z,out_proj} (144) int8, group_size 128 round-to-nearest
self_attn.o_proj (16) BF16 (bit-identical to base) —
self_attn.{q,k,v}_proj (48) int4, group_size 128 AutoRound (400 iters × 256 samples)
lm_head int8, group_size 128 round-to-nearest
embed_tokens, vision tower (333 tensors), MTP (15), all norms, linear_attn.in_proj_a/b, conv1d BF16 —

Two things worth knowing before you use this

  1. The MLP, GatedDeltaNet and lm_head tensors are round-to-nearest, not AutoRound-tuned. Only q_proj/k_proj/v_proj carry AutoRound tuning (they are inherited unchanged from an AutoRound int4 run). The measurements below were taken on round-to-nearest tensors, so they describe exactly what is in this repository — an AutoRound run at this same configuration would likely do better, and that is the obvious next step.
  2. mlp uses group_size: 64, not 128. Marlin supports group sizes {32, 64, 128}; 64 was chosen because it measured better per byte. This makes the MLP weights slightly larger than a group_size-128 build, and Marlin's kernel does more scale loads per tile — expect a small throughput cost on prefill relative to a pure g128 build.

Quality

Perplexity on 50 held-out texts from wikitext-103, 16,007 evaluated tokens — deliberately not the calibration set. Same texts, same tokenization, same method, both models served by vLLM. The BF16 reference is 8.0880, reproduced identically across every run of this series, so the deltas are directly comparable.

version size perplexity delta
base BF16 52 GB 8.0880 —
this build 21.98 GB 8.2371 +1.84%

Against the other builds of this same model

build size perplexity delta
W8A16, attention in BF16 33.3 GB 8.1991 +1.37%
this build (mixed 4/8-bit) 21.98 GB 8.2371 +1.84%
W6A16, attention in BF16 27.6 GB 8.2823 +2.40%
W4A16, attention + GDN in BF16 30.2 GB 8.3888 +3.72%
W4A16, nothing excluded 19.5 GB 8.7945 +8.74%

This build dominates the W6A16 and W4A16 builds: smaller and more faithful. The W8A16 is still more faithful (+1.37%), at 52% more size.

How the budget was spent

Every row is a real vLLM measurement against the same fully-quantized 4-bit base (8.7945), on the identical corpus. Gains are not additive — the attention groups in particular are strongly sub-additive — so the final configuration was measured, not computed.

Single groups

change perplexity gain cost gain per GB
GatedDeltaNet group_size 128 → 32 8.6767 0.1163 0.260 GB 0.447
o_proj → BF16 8.5724 0.2206 0.747 GB 0.295
whole attention → int8 8.5548 0.2382 0.839 GB 0.284
qkv → int8 8.6522 0.1408 0.587 GB 0.240
whole attention → BF16 8.5524 0.2406 2.490 GB 0.097
qkv → BF16 8.6509 0.1421 1.743 GB 0.082
GatedDeltaNet → int8 8.6286 0.1644 2.768 GB 0.059
GatedDeltaNet → int6 8.7218 0.0712 1.384 GB 0.051
MLP → int8 8.6117 0.1813 8.556 GB 0.021
MLP group_size 128 → 64 8.7711 0.0219 0.267 GB 0.082
MLP group_size 128 → 32 8.8337 −0.0407 0.802 GB —

Combined configurations that fit the budget

configuration perplexity delta size
chosen: o_proj BF16 + GDN int8 + MLP g64 8.2371 +1.84% 21.98 GB
attention BF16 + MLP g32 + GDN g32 8.2796 +2.37% 21.75 GB
attention int8 + o_proj BF16 + GDN g32 + MLP g32 8.2832 +2.41% 20.60 GB
o_proj BF16 + qkv int8 + GDN int6 + MLP g32 8.3157 +2.82% 21.72 GB
attention int8 + GDN int8 8.3956 +3.80% 21.81 GB
attention int8 + o_proj BF16 + GDN g32 8.4731 +4.76% 19.80 GB
(over budget) o_proj BF16 + qkv int8 + GDN int8 + MLP g64 8.1930 +1.30% 22.57 GB

What the measurements showed

  • int8 dominates BF16 in the attention. attn int8 gets 99% of attn BF16's gain for 34% of the bytes. Putting BF16 on qkv is the single worst use of budget measured here.
  • o_proj carries ~92% of the attention damage. Alone (0.747 GB) it gains 0.2206, while all four attention projections in BF16 (2.490 GB) gain only 0.2406. This reproduces, in the 4-bit context, what the 8-bit sweep had hinted at — and it is not portable between bit-widths in general.
  • Interactions are strong in both directions. Adding o_proj BF16 on top of attention int8 gains only 0.0020 — the int8 already fixed it. Conversely, MLP group_size 32 hurts when applied alone (−0.0407) but helps by 0.1755 once attention and GatedDeltaNet are corrected: with the larger errors removed, the MLP's own error becomes the binding constraint.
  • GatedDeltaNet int6 is dominated by int8 on both counts — it gains less per byte (0.051 vs 0.059) and it drops off Marlin onto the JIT-compiled Humming kernel. It is not used here.

Serving

At 21.98 GB this fits a single 24 GB card, but with very little left for KV cache. It is sized to be served on 2× RTX 3090 (24 GB, NVLink) with tensor parallelism, which is the configuration it was validated on: ~90-100 tok/s decode (with the pelican svg prompt) and a KV cache of ~490k tokens.

config.json declares four config_groups: two added for this build (MLP at group_size 64, GatedDeltaNet at int8) and two inherited from the AutoRound base build (lm_head at int8, and the generic Linear int4 fallback). vLLM resolves a scheme per layer by name, in config order — the re:.*lm_head$ target is listed before the generic Linear target, and that order is load-bearing: the generic group must stay last, as the fallback for everything not matched by a specific rule.

lm_head is quantized, and the MTP draft head shares the lm_head.weight key, so both take the same scheme. (See the MTP caveat below: with speculative decoding enabled the generated text is coherent and the final answers are equivalent, but it does not reproduce the non-speculative output token-for-token.)

Example command

vllm serve /root/models/DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound \
    --served-model-name "qwen3.6-27b" \
    --tensor-parallel-size 2 \
    --pipeline-parallel-size 1 \
    --max-model-len auto \
    --override-generation-config '{"temperature": 0.7, "top_p": 0.8, "top_k": 20, "min_p": 0.0, "presence_penalty": 1.5, "repetition_penalty": 1.0}' \
    --seed 1234 \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --max-num-batched-tokens 2048 \
    --trust-remote-code \
    --disable-custom-all-reduce \
    --gpu-memory-utilization 0.94 \
    --max-num-seqs 32 \
    --host 0.0.0.0 \
    --port 3434 \
    --kv-cache-dtype fp8_e4m3 \
    --default-chat-template-kwargs '{"enable_thinking": true, "reasoning_effort": "xhigh"}' \
    --performance-mode balanced \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Sizing the context

Only the 16 full-attention layers cache K and V. The other 48 are GatedDeltaNet, which carries a constant-size recurrent state and grows no KV at all — so the cache is unusually cheap for a 27B model:

16 layers × 4 KV heads × 256 head_dim × 2 (K and V) × 1 byte (fp8_e4m3) = 32 KiB per token
~490k tokens × 32 KiB ≈ 16 GB

That is why --max-num-seqs 32 with --gpu-memory-utilization 0.94 lands near half a million tokens of context. Raise --max-num-seqs for more concurrency at the cost of per-sequence context, and --max-model-len to cap it explicitly instead of auto.

--kv-cache-dtype fp8_e4m3 halves the cache versus bf16 and is what the numbers above assume; dropping it doubles the per-token cost to 64 KiB and roughly halves the reachable context.

Caveats

  • Round-to-nearest MLP/GatedDeltaNet/lm_head tensors (see above) — measured as shipped, but not optimal.
  • The +1.84% number is text-only perplexity. Vision and MTP were validated functionally (a colour-identification prompt and speculative-decoding acceptance) but not scored on a vision benchmark.
  • Asymmetric weight quantization was not used: it was not implemented for this build (it needs weight_zero_point tensor support, and its value was never measured). group_size 32/64 and int8 were used instead.
  • MTP speculative decoding does not reproduce greedy output token-for-token on this model + vLLM combination. Outputs are coherent in both modes but diverge partway through long chain-of-thought generations. This is pre-existing and not specific to this build: it reproduces on the unmodified 4-bit base and on the previously published builds of this model (first divergence at 313 and 696 characters for the base, 1243 and 171 for the earlier W4A16 build, 131 and 437 here). Treat MTP here as a throughput feature, not a bit-exactness guarantee.

Tokenizer warning — safe to ignore, and do not "fix" it

Some transformers versions log this at load time:

The tokenizer you are loading from … with an incorrect regex pattern … This will lead to incorrect tokenization. You should set the fix_mistral_regex=True flag …

It is a false positive, and following its advice would degrade this model. The tokenizer uses the standard GPT-2/Qwen pre-tokenizer pattern — the one this model was trained with. The warning comes from a heuristic that cannot tell which model family a tokenizer belongs to: it reads transformers_version from config.json and, when that field is missing, assumes the buggy Mistral pattern. This artifact now carries the field, so the warning no longer appears. (It only ever fired for local-path loads, i.e. local_files_only=True; loading from the Hub never reaches that code path for a non-Mistral model.)

fix_mistral_regex=True replaces the pattern with Mistral's, which splits camelCase apart. Measured on this tokenizer:

input default (correct for Qwen) with fix_mistral_regex=True
iPhone macOS camelCaseWord iPhone · macOS · camel · Case · Word i · Phone · mac · OS · camel · Case · Word

The difference is subtle — on ordinary prose, numbers, punctuation and precomposed accents the two patterns agree exactly. It shows up on identifiers, brand names and camelCase, which is precisely where you would not want it.

Reproducing

Built with graft_plan.py from a fully AutoRound-quantized int4 base, then two groups re-quantized round-to-nearest and one restored to BF16 — each step validated by dequantization round-trip (mean error ≈ quantisation step / 4):

--take oprop=bf16 --take gdn=int8 --take mlp=int4/g64

32 candidate configurations were served by vLLM and scored on the same corpus before this one was chosen; the tables above are those measurements.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Stored precision
bfloat16
Model type
qwen3_5
Quantization
compressed-tensors

Identity and Version

Repository
DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound
Publisher
DoktorMincs
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
33.3B parameters
Languages
mtp
Revision
93b7ff2da1ea7931b74b25fa3b2b9eb80fa964fb
First published
2026-09-28
Last updated
2026-09-29

Files and Weights

19 files, 22.0 GB in total. The weights are 5 files totalling 22.0 GB in safetensors.

Weights5 files · 22.0 GB
Configuration6 files · 209.8 KB
Tokenizer3 files · 26.7 MB
Documentation3 files · 25.8 KB
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00005.safetensorsWeights5.4 GB 763de0747e1c
model-00002-of-00005.safetensorsWeights5.4 GB c915ad3eec47
model-00003-of-00005.safetensorsWeights5.4 GB 36a3d40427d0
model-00004-of-00005.safetensorsWeights5.4 GB b021b08e0043
model-00005-of-00005.safetensorsWeights565.7 MB 4a20b4e7dea4
config.jsonConfiguration12.7 KB —
generation_config.jsonConfiguration213 B —
model.safetensors.index.jsonConfiguration194.9 KB —
preprocessor_config.jsonConfiguration390 B —
processor_config.jsonConfiguration1.2 KB —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation11.4 KB —
NOTICEDocumentation1.9 KB —
README.mdDocumentation12.6 KB —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 87a7830d63fc
tokenizer_config.jsonTokenizer16.4 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
22.0 GB
Download from DoktorMincs

Released by DoktorMincs through its official repository on Hugging Face. Read the license.

Built From

  • Derived from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
  • Quantized from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU

Memory Requirements

PrecisionWeights in memory
As published22.0 GB
16-bit66.6 GB
8-bit33.3 GB
4-bit16.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound

How much GPU memory does Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound need?

About 80 GB at 16-bit and 20 GB at 4-bit: the weights (33.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound commercially?

Yes. Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

HyperCLOVA X KETI-HAECHI-32B is a multimodal model derived from It was developed for two primary purposes: improving Korean cultural-heritage understanding and Korean OCR, and improving tool calling and multi-step, stateful agent execution. Alongside these goals, the model retains broad multimodal, language, and coding capabilities from the base model. heritage objects, answers questions grounded in heritage images, and reads Korean text from signs, scenes, rendered text, and public documents. - Tool calling and long-horizon task execution: selects and calls tools, carries information across multiple turns, tracks changing state, and works toward an end-to-end goal over several steps.…

Open weights other 33.3B parameters 131,072 tokens transformers

Model · Image and text to text

Qwen2.5-VL-32B-Instruct-AWQ

Qwen

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…

Open weights apache-2.0 33.5B parameters 128,000 tokens transformers

Model · Image and text to text

Qwen2.5-VL-32B-Instruct

Qwen

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…

Open weights apache-2.0 33.5B parameters 128,000 tokens transformers

Model · Image and text to text

gemma-4-31B-it

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 31.3B parameters 262,144 tokens transformers