SAVRN
Search Contact SAVRN

Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw · Model Card

Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw: Model Card

Written by Lee, published under apache-2.0, revision 64371e7543b2, read 2026-09-20. Shown as written; SAVRN's own facts about this model are on its page.

This repository contains an EXL3 3.0 bpw quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated.

The original model and abliteration work are attributed to huihui-ai; this repository contains the EXL3 quantization produced by grimlee.

The source model is an abliterated / uncensored derivative of Qwen3.8-27B. For details about the original model modification and its behavior, refer to the source model card.

The checkpoint retains the Qwen3.8 multimodal model structure, including the vision component. The optional SM89 runtime work documented below does not change the model format or weights.

Quantization details

The checkpoint uses the standard EXL3 format with a target body bitrate of 3.0 bits per weight.

The target bitrate applies to the quantized body; not every tensor in the checkpoint is stored at exactly 3 bits.

  • EXL3 target body bitrate: 3.0 bpw
  • lm_head: 6-bit
  • Vision component: 6-bit
  • MTP component: 4-bit
  • Calibration rows: 250
  • Calibration context: 2048
  • Codebook: mul1
  • Output scales: always
  • Shard size setting: 4096 MB
  • Final output: four safetensors shards

The four shards are the expected result of the 4096 MB shard-size setting and do not indicate a missing or incomplete model.

The checkpoint metadata reports:

  • quant_method=exl3
  • version=1.5.0
  • bits=3.0
  • head_bits=6
  • vision_bits=6
  • mtp_bits=4

The conversion configuration was approximately:

python convert.py \
  -i "$SRC" \
  -o "$OUT" \
  -w "$WORK" \
  -b 3.0 \
  -hb 6 \
  -mb 4 \
  -vb 6 \
  -cr 250 \
  -cc 2048 \
  -cb mul1 \
  --out_scales always \
  -ss 4096 \
  -cpi 120 \
  -d 0

Runtime compatibility

This is a standard EXL3 checkpoint.

The model itself does not require the SM89 fork described below. The fork is an optional runtime optimization for Ada SM89 GPUs and does not modify the EXL3 checkpoint format.

  • Upstream ExLlamaV3:
    https://github.com/turboderp-org/exllamav3

  • Experimental Ada / SM89 fork:
    https://github.com/grimlee/exllamav3

  • SM89 optimization branch:
    https://github.com/grimlee/exllamav3/tree/sm89-ada-optimizations

Optional SM89 / Ada runtime optimization

I profiled ExLlamaV3 1.5.0 prefill performance on an NVIDIA RTX 4060 Ti 16 GB (Ada SM89) and tested two small SM89-specific runtime specializations.

These results are specific to the tested GPU, model shapes and runtime configuration. Other Ada / SM89 GPUs and model shapes have not yet been characterized to the same extent.

F16ACC GEMM specialization

The tested SM89 F16ACC layout is:

BM=128
BN=64
BK=64
4 warps (2x2)
STAGES=2
GROUP_M=16
PAD=0

Matched cold-prefill server A/B results:

Context Control SM89 F16ACC Uplift
16K 727.38 tok/s 761.64 tok/s +4.710%
32K 652.37 tok/s 679.31 tok/s +4.129%
64K 535.73 tok/s 554.60 tok/s +3.522%

Correctness was checked on all 16 F16ACC shapes observed in the real workload using three deterministic seeds per shape.

Results:

  • max_abs = 0
  • mean_abs = 0
  • max_relative = 0
  • no NaN
  • no Inf

A per-shape oracle / dispatcher was also evaluated. It projected only about 1.8 tok/s beyond the simpler global SM89 configuration, so the global configuration was retained.

The oracle number is a projection, not an additional end-to-end server measurement.

Long-query paged attention

The second specialization targets the staged/cache-backed FP16 long-query prefill path.

The direct packed quantized-cache path keeps the upstream behavior.

Tested configuration:

Control:
BM64 / BN32 / W8 / S2

SM89 candidate:
BM64 / BN32 / W4 / S2

The key change is therefore:

8 warps -> 4 warps

The F16ACC specialization above was enabled in both A/B arms, so the following experiment isolates the paged-attention change.

Formal cold-request server results:

Context Control SM89 optimized Uplift
16K 763.530 tok/s 787.156 tok/s +3.094%
32K 680.912 tok/s 723.184 tok/s +6.208%
64K 555.397 tok/s 613.711 tok/s +10.500%

The formal requests used:

cached_tokens=0
cached_pages=0

so these are cold-prefill measurements rather than prefix-cache reuse results.

Correctness gates at 16K, 32K and 64K all passed with:

max_abs=0

No measured VRAM increase was observed.

Peak VRAM:

Context Peak VRAM
16K 14,488 MiB
32K 14,488 MiB
64K 14,608 MiB

During the formal A/B runs there was no observed:

  • OOM
  • CUDA illegal-access error
  • NVIDIA Xid
  • server crash

Benchmark configuration

The reported runtime measurements used:

Setting Value
GPU NVIDIA RTX 4060 Ti 16 GB
Architecture Ada SM89 / compute capability 8.9
ExLlamaV3 1.5.0
Upstream base commit 02aef45cd681b960a00afcd0749a4ab99e6c1bfe
Model Qwen3.8-27B EXL3 3.0 bpw
KV cache int4
QC staging EXL3_QC_STAGING=1
Maximum prefill chunk 4096 tokens
MTP enabled
Vision disabled for the performance A/B
Request type cold requests
Prefix cache cached_tokens=0

Prefill throughput in the tables is the server's measured prefill timing, rather than a client-side prompt-token / TTFT approximation.

These measurements should not be interpreted as a guarantee of identical performance gains across every Ada GPU, model shape, batch size, cache mode or context length.

The checkpoint retains the multimodal vision structure and its quantization metadata reports vision_bits=6. Vision was disabled during these performance benchmarks, so the results above should not be interpreted as a complete vision-runtime compatibility test.

Runtime source and validation

  • Fork:
    https://github.com/grimlee/exllamav3

  • SM89 branch:
    https://github.com/grimlee/exllamav3/tree/sm89-ada-optimizations

  • Detailed benchmark notes:
    https://github.com/grimlee/exllamav3/blob/sm89-ada-optimizations/doc/sm89_ada_optimizations.md

  • Full GitHub Actions build matrix:
    https://github.com/grimlee/exllamav3/actions/runs/35473907070

  • Upstream discussion:
    https://github.com/turboderp-org/exllamav3/issues/384

The referenced GitHub Actions run completed successfully with:

target=all
release=0

This validates the source and wheel build matrix. It does not replace the GPU correctness testing and server A/B measurements described above.

EXL3 runtime usage

For normal use, upstream ExLlamaV3 can be used directly:

https://github.com/turboderp-org/exllamav3

For the optional RTX 4060 Ti / SM89 optimized runtime:

git clone https://github.com/grimlee/exllamav3.git
cd exllamav3
git checkout sm89-ada-optimizations

The SM89 branch is an experimental runtime optimization.

It is not required to load this model and does not introduce a different model format.

Source model and usage considerations

This repository only provides an EXL3 quantization and the runtime benchmark work documented above.

For information about:

  • the original Qwen3.8 model,
  • the abliteration procedure,
  • model behavior,
  • safety characteristics,
  • previous source-model revisions,
  • and source-model-specific usage instructions,

please refer to:

huihui-ai/Huihui-Qwen3.8-27B-abliterated
https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated

This quantization does not introduce separate safety tuning beyond the source model.