Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw · Model Card
Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw: Model Card
Written by Lee, published under apache-2.0, revision 64371e7543b2, read 2026-09-20. Shown as written; SAVRN's own facts about this model are on its page.
This repository contains an EXL3 3.0 bpw quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated.
- Base model: Qwen/Qwen3.8-27B
- Modified model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
- Quantization: EXL3 3.0 bpw
- Quantization by:
grimlee - Recommended runtime: ExLlamaV3
- License: Apache-2.0
The original model and abliteration work are attributed to huihui-ai; this
repository contains the EXL3 quantization produced by grimlee.
The source model is an abliterated / uncensored derivative of Qwen3.8-27B. For details about the original model modification and its behavior, refer to the source model card.
The checkpoint retains the Qwen3.8 multimodal model structure, including the vision component. The optional SM89 runtime work documented below does not change the model format or weights.
Quantization details
The checkpoint uses the standard EXL3 format with a target body bitrate of 3.0 bits per weight.
The target bitrate applies to the quantized body; not every tensor in the checkpoint is stored at exactly 3 bits.
- EXL3 target body bitrate:
3.0 bpw lm_head:6-bit- Vision component:
6-bit - MTP component:
4-bit - Calibration rows:
250 - Calibration context:
2048 - Codebook:
mul1 - Output scales:
always - Shard size setting:
4096 MB - Final output: four
safetensorsshards
The four shards are the expected result of the 4096 MB shard-size setting
and do not indicate a missing or incomplete model.
The checkpoint metadata reports:
quant_method=exl3version=1.5.0bits=3.0head_bits=6vision_bits=6mtp_bits=4
The conversion configuration was approximately:
python convert.py \
-i "$SRC" \
-o "$OUT" \
-w "$WORK" \
-b 3.0 \
-hb 6 \
-mb 4 \
-vb 6 \
-cr 250 \
-cc 2048 \
-cb mul1 \
--out_scales always \
-ss 4096 \
-cpi 120 \
-d 0
Runtime compatibility
This is a standard EXL3 checkpoint.
The model itself does not require the SM89 fork described below. The fork is an optional runtime optimization for Ada SM89 GPUs and does not modify the EXL3 checkpoint format.
-
Upstream ExLlamaV3:
https://github.com/turboderp-org/exllamav3 -
Experimental Ada / SM89 fork:
https://github.com/grimlee/exllamav3 -
SM89 optimization branch:
https://github.com/grimlee/exllamav3/tree/sm89-ada-optimizations
Optional SM89 / Ada runtime optimization
I profiled ExLlamaV3 1.5.0 prefill performance on an NVIDIA RTX 4060 Ti 16 GB (Ada SM89) and tested two small SM89-specific runtime specializations.
These results are specific to the tested GPU, model shapes and runtime configuration. Other Ada / SM89 GPUs and model shapes have not yet been characterized to the same extent.
F16ACC GEMM specialization
The tested SM89 F16ACC layout is:
BM=128
BN=64
BK=64
4 warps (2x2)
STAGES=2
GROUP_M=16
PAD=0
Matched cold-prefill server A/B results:
| Context | Control | SM89 F16ACC | Uplift |
|---|---|---|---|
| 16K | 727.38 tok/s | 761.64 tok/s | +4.710% |
| 32K | 652.37 tok/s | 679.31 tok/s | +4.129% |
| 64K | 535.73 tok/s | 554.60 tok/s | +3.522% |
Correctness was checked on all 16 F16ACC shapes observed in the real workload using three deterministic seeds per shape.
Results:
max_abs = 0mean_abs = 0max_relative = 0- no NaN
- no Inf
A per-shape oracle / dispatcher was also evaluated. It projected only about
1.8 tok/s beyond the simpler global SM89 configuration, so the global
configuration was retained.
The oracle number is a projection, not an additional end-to-end server measurement.
Long-query paged attention
The second specialization targets the staged/cache-backed FP16 long-query prefill path.
The direct packed quantized-cache path keeps the upstream behavior.
Tested configuration:
Control:
BM64 / BN32 / W8 / S2
SM89 candidate:
BM64 / BN32 / W4 / S2
The key change is therefore:
8 warps -> 4 warps
The F16ACC specialization above was enabled in both A/B arms, so the following experiment isolates the paged-attention change.
Formal cold-request server results:
| Context | Control | SM89 optimized | Uplift |
|---|---|---|---|
| 16K | 763.530 tok/s | 787.156 tok/s | +3.094% |
| 32K | 680.912 tok/s | 723.184 tok/s | +6.208% |
| 64K | 555.397 tok/s | 613.711 tok/s | +10.500% |
The formal requests used:
cached_tokens=0
cached_pages=0
so these are cold-prefill measurements rather than prefix-cache reuse results.
Correctness gates at 16K, 32K and 64K all passed with:
max_abs=0
No measured VRAM increase was observed.
Peak VRAM:
| Context | Peak VRAM |
|---|---|
| 16K | 14,488 MiB |
| 32K | 14,488 MiB |
| 64K | 14,608 MiB |
During the formal A/B runs there was no observed:
- OOM
- CUDA illegal-access error
- NVIDIA Xid
- server crash
Benchmark configuration
The reported runtime measurements used:
| Setting | Value |
|---|---|
| GPU | NVIDIA RTX 4060 Ti 16 GB |
| Architecture | Ada SM89 / compute capability 8.9 |
| ExLlamaV3 | 1.5.0 |
| Upstream base commit | 02aef45cd681b960a00afcd0749a4ab99e6c1bfe |
| Model | Qwen3.8-27B EXL3 3.0 bpw |
| KV cache | int4 |
| QC staging | EXL3_QC_STAGING=1 |
| Maximum prefill chunk | 4096 tokens |
| MTP | enabled |
| Vision | disabled for the performance A/B |
| Request type | cold requests |
| Prefix cache | cached_tokens=0 |
Prefill throughput in the tables is the server's measured prefill timing, rather than a client-side prompt-token / TTFT approximation.
These measurements should not be interpreted as a guarantee of identical performance gains across every Ada GPU, model shape, batch size, cache mode or context length.
The checkpoint retains the multimodal vision structure and its quantization
metadata reports vision_bits=6. Vision was disabled during these performance
benchmarks, so the results above should not be interpreted as a complete
vision-runtime compatibility test.
Runtime source and validation
-
Fork:
https://github.com/grimlee/exllamav3 -
SM89 branch:
https://github.com/grimlee/exllamav3/tree/sm89-ada-optimizations -
Detailed benchmark notes:
https://github.com/grimlee/exllamav3/blob/sm89-ada-optimizations/doc/sm89_ada_optimizations.md -
Full GitHub Actions build matrix:
https://github.com/grimlee/exllamav3/actions/runs/35473907070 -
Upstream discussion:
https://github.com/turboderp-org/exllamav3/issues/384
The referenced GitHub Actions run completed successfully with:
target=all
release=0
This validates the source and wheel build matrix. It does not replace the GPU correctness testing and server A/B measurements described above.
EXL3 runtime usage
For normal use, upstream ExLlamaV3 can be used directly:
https://github.com/turboderp-org/exllamav3
For the optional RTX 4060 Ti / SM89 optimized runtime:
git clone https://github.com/grimlee/exllamav3.git
cd exllamav3
git checkout sm89-ada-optimizations
The SM89 branch is an experimental runtime optimization.
It is not required to load this model and does not introduce a different model format.
Source model and usage considerations
This repository only provides an EXL3 quantization and the runtime benchmark work documented above.
For information about:
- the original Qwen3.8 model,
- the abliteration procedure,
- model behavior,
- safety characteristics,
- previous source-model revisions,
- and source-model-specific usage instructions,
please refer to:
huihui-ai/Huihui-Qwen3.8-27B-abliterated
https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated
This quantization does not introduce separate safety tuning beyond the source model.