# Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw by Lee
Source: https://savrn.com/models/huihui-qwen3-8-27b-abliterated-exl3-3-0bpw
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw (6.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 13.5 GB | 16.2 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 6.7 GB | 8.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 3.4 GB | 4.0 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/huihui-qwen3-8-27b-abliterated-exl3-3-0bpw/gpus)

## Model Card

By Lee, published under apache-2.0, revision 64371e7543b2.

This repository contains an EXL3 3.0 bpw quantization of [huihui-ai/Huihui-Qwen3.8-27B-abliterated](https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated).

- Base model: [Qwen/Qwen3.8-27B](https://savrn.com/models/qwen3-8-27b)
- Modified model: [huihui-ai/Huihui-Qwen3.8-27B-abliterated](https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated)
- Quantization: EXL3 3.0 bpw
- Quantization by: grimlee
- Recommended runtime: ExLlamaV3
- License: Apache-2.0

The original model and abliteration work are attributed to huihui-ai; this repository contains the EXL3 quantization produced by grimlee.

The source model is an abliterated / uncensored derivative of Qwen3.8-27B. For details about the original model modification and its behavior, refer to the [source model card](https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated).

The checkpoint retains the Qwen3.8 multimodal model structure, including the vision component. The optional SM89 runtime work documented below does not change the model format or weights.

### Quantization details

The checkpoint uses the standard EXL3 format with a target body bitrate of 3.0 bits per weight.

The target bitrate applies to the quantized body; not every tensor in the checkpoint is stored at exactly 3 bits.

[Read the full model card (955 words)](https://savrn.com/models/huihui-qwen3-8-27b-abliterated-exl3-3-0bpw/card)

## Configuration

Architecture

Qwen3_5ForConditionalGeneration

Context length (tokens)

262,144

Layers

64

Hidden size

5,120

Feed-forward size

17,408

Attention heads

24

Key/value heads

4

Head dimension

256

Vocabulary size

248,320

Model type

qwen3_5

Quantization

exl3

## Identity and Version

Repository

grimlee/Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw

Publisher

Lee

Task

Image and text to text

Modality

Image and text

Library

transformers

Parameters

6.7B parameters

Languages

Not stated by the source

Revision

64371e7543b20705b491b4a1ab4b2ba42722b3a8

First published

2026-09-19

Last updated

2026-09-20

## Files and Weights

19 files, 13.5 GB in total. The weights are 4 files totalling 13.5 GB in safetensors.

Weights4 files · 13.5 GB

Configuration6 files · 930.3 KB

Tokenizer4 files · 22.9 MB

Documentation2 files · 19.9 KB

Other2 files · 9.2 KB

Repository1 file · 1.6 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model-00001-of-00004.safetensors | Weights | 4.3 GB | 234b08e7a8d1 |
| model-00002-of-00004.safetensors | Weights | 4.2 GB | 8c1d6c456318 |
| model-00003-of-00004.safetensors | Weights | 4.3 GB | 5acac8bc06f6 |
| model-00004-of-00004.safetensors | Weights | 756.6 MB | de640403d5aa |
| config.json | Configuration | 4.6 KB | — |
| generation_config.json | Configuration | 202 B | — |
| model.safetensors.index.json | Configuration | 295.6 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| quantization_config.json | Configuration | 629.1 KB | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| LICENSE | Documentation | 11.5 KB | — |
| README.md | Documentation | 8.4 KB | — |
| chat_template.jinja | Other | 9.0 KB | — |
| crc32.txt | Other | 238 B | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 12.8 MB | 0997f410c57a |
| tokenizer_config.json | Tokenizer | 17.9 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

13.5 GB

[Download from Lee](https://huggingface.co/grimlee/Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw)

Released by Lee through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from huihui-ai/Huihui-Qwen3.8-27B-abliterated
- Quantized from huihui-ai/Huihui-Qwen3.8-27B-abliterated

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 13.5 GB |
| 16-bit | 13.5 GB |
| 8-bit | 6.7 GB |
| 4-bit | 3.4 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw

### How much GPU memory does Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw need?

About 16.2 GB at 16-bit and 4 GB at 4-bit: the weights (6.7B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw commercially?

Yes. Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

### What is Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Image and text to text

### [llava-1.5-7b-hf](https://savrn.com/models/llava-1-5-7b-hf)

[Llava Hugging Face](https://savrn.com/model-publishers/llava-hf)

Below is the model card of Llava model 7b, which is copied from the original Llava model card that you can find here. Check out also the Google Colab demo to run Llava on a free-tier Google Colab instance: Or check out our Spaces demo! LLaVA is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal instruction-following data. It is an auto-regressive language model, based on the transformer architecture. LLaVA-v1.5-7B was trained in September 2023. Paper or resources for more information: https://llava-vl.github.io/ First, make sure to have transformers >= 4.35.3. The model supports multi-image and multi-prompt generation. Meaning that you can pass multiple…

Open weights llama2 7.1B parameters 4,096 tokens transformers

[View model](https://savrn.com/models/llava-1-5-7b-hf)

Model · Image and text to text

### [Huihui-Qwen3.6-27B-abliterated-AWQ-MTP](https://savrn.com/models/huihui-qwen3-6-27b-abliterated-awq-mtp)

[Shawn Wei](https://savrn.com/model-publishers/shawnw3i)

This is an uncensored version of Qwen/Qwen3.6-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. - AWQ Marlin kernel supported (auto-converted by vLLM at runtime) - MTP speculative decoding supported out of the box - 110+ tok/s on a single A800 80GB (vLLM 0.21.0, MTP enabled, fp8 KV cache) - Risk of Sensitive or Controversial Outputs: This model’s safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated…

Open weights apache-2.0 6.3B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/huihui-qwen3-6-27b-abliterated-awq-mtp)

Model · Image and text to text

### [Swift-1.5-Qwen3.8-27b-heretic-W4A16](https://savrn.com/models/swift-1-5-qwen3-8-27b-heretic-w4a16)

[Akumaburn](https://savrn.com/model-publishers/akumaburn)

A 18G INT4 build of for single-user decoding and low-VRAM deployment. (Marlin on Ampere). - Unrotated, so the DFlash2 drafter works. GatedDeltaNet gates, all norms. 4-bit costs 4.5× the rotated W8A8's quantization divergence (0.0264) and 1.5× the unrotated W8A8's (0.0791). Choose this build when the 18G footprint or single-user decode speed matters more than fidelity. Produced with Heretic v2.0.0.dev0: 260 TPE trials, seed 42, directional ablation on attn.oproj (16 modules), attn.outproj (48) and mlp.downproj (64). Baseline refusal score before abliteration: 98/100. Trial 140 was selected — Pareto index 1, not index 0. Index 0 (trial 161) scored 22/100 keywords at KL 0.1091; trial 140…

Open weights other 6.3B parameters 262,144 tokens vllm

[View model](https://savrn.com/models/swift-1-5-qwen3-8-27b-heretic-w4a16)

Model · Image and text to text

### [openvla-7b-finetuned-libero-spatial](https://savrn.com/models/openvla-7b-finetuned-libero-spatial)

[OpenVLA Collaboration](https://savrn.com/model-publishers/openvla)

This model was produced by fine-tuning the OpenVLA 7B model via LoRA (r=32) on the LIBERO-Spatial dataset from the LIBERO simulation benchmark. We made a few modifications to the training dataset to improve final performance (see the OpenVLA paper for details). Below are the hyperparameters we used for all LIBERO experiments: - No gradient accumulation (i.e. gradaccumulationsteps == 1) - shufflebuffersize == 100000 See the OpenVLA GitHub README for instructions on how to run and evaluate this model in the LIBERO simulator.

Open weights mit 7.5B parameters transformers

[View model](https://savrn.com/models/openvla-7b-finetuned-libero-spatial)

Model · Image and text to text

### [openvla-7b-finetuned-libero-object](https://savrn.com/models/openvla-7b-finetuned-libero-object)

[OpenVLA Collaboration](https://savrn.com/model-publishers/openvla)

This model was produced by fine-tuning the OpenVLA 7B model via LoRA (r=32) on the LIBERO-Object dataset from the LIBERO simulation benchmark. We made a few modifications to the training dataset to improve final performance (see the OpenVLA paper for details). Below are the hyperparameters we used for all LIBERO experiments: - No gradient accumulation (i.e. gradaccumulationsteps == 1) - shufflebuffersize == 100000 See the OpenVLA GitHub README for instructions on how to run and evaluate this model in the LIBERO simulator.

Open weights mit 7.5B parameters transformers

[View model](https://savrn.com/models/openvla-7b-finetuned-libero-object)

Model · Image and text to text

### [openvla-7b-finetuned-libero-10](https://savrn.com/models/openvla-7b-finetuned-libero-10)

[OpenVLA Collaboration](https://savrn.com/model-publishers/openvla)

This model was produced by fine-tuning the OpenVLA 7B model via LoRA (r=32) on the LIBERO-10 (LIBERO-Long) dataset from the LIBERO simulation benchmark. We made a few modifications to the training dataset to improve final performance (see the OpenVLA paper for details). Below are the hyperparameters we used for all LIBERO experiments: - No gradient accumulation (i.e. gradaccumulationsteps == 1) - shufflebuffersize == 100000 See the OpenVLA GitHub README for instructions on how to run and evaluate this model in the LIBERO simulator.

Open weights mit 7.5B parameters transformers

[View model](https://savrn.com/models/openvla-7b-finetuned-libero-10)

## Lee

[All models and datasets](https://savrn.com/model-publishers/grimlee)

## Versions

- [64371e7543b2](https://savrn.com/models/huihui-qwen3-8-27b-abliterated-exl3-3-0bpw/versions/64371e7543b2) · current 2026-09-20

## Explore More

- [All image and text to text models](https://savrn.com/models/tasks/image-and-text-to-text)
- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-09-20.
- [Hugging Face record](https://huggingface.co/grimlee/Huihui-Qwen3.8-27B-abliterated-EXL3-3.0bpw)
- [How the hub is built](https://savrn.com/model-hub/methodology)
