Open-weight model
Qwen-Image-2.1-Text-Encoder-Heretic
by Pottokao pottokao/Qwen-Image-2.1-Text-Encoder-Heretic
Qwen-Image-2.1-Text-Encoder-Heretic is an open-weight model from Pottokao, released under Apache License 2.0. It has 8.8B parameters and a 262,144-token context. At 16-bit it needs about 21 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 799 downloads a month.
The text encoder of Qwen/Qwen-Image-2.1 (a Qwen3-VL-8B-Instruct) with refusal behaviour removed via Heretic directional ablation. Drop-in replacement for the stock text encoder. Weights are bf16, same shapes, same parameter count — nothing else was changed.
Runs On
What it takes to serve Qwen-Image-2.1-Text-Encoder-Heretic (8.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 17.5 GB | 21.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 8.8 GB | 10.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 4.4 GB | 5.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 23, 2026.
Qwen-Image-2.1-Text-Encoder-Heretic on every accelerator the SAVRN Index prices, at every precision
Model Card
By Pottokao, published under apache-2.0, revision 106be9554084.
Qwen-Image-2.1 Text Encoder — Heretic (Abliterated)
Not affiliated with, or endorsed by, Alibaba / Qwen. Community derivative (refusal-ablated) of
Qwen/Qwen3-VL-8B-Instruct— the model Qwen-Image-2.1 uses, unmodified, as its text encoder. Qwen releases that model under Apache-2.0, so this derivative is redistributed under Apache-2.0 (seeLICENSEandNOTICE).
The text encoder of Qwen/Qwen-Image-2.1
(a Qwen3-VL-8B-Instruct) with refusal behaviour removed via
Heretic directional ablation.
Drop-in replacement for the stock text encoder. Weights are bf16, same shapes, same parameter count — nothing else was changed.
Results
| Refusals | KL divergence | |
|---|---|---|
| Original (measured baseline) | 100/100 | 0 (by definition) |
| This model | 5/100 | 0.0220 |
Measured by Heretic on mlabonne/harmful_behaviors (test split) for refusals and
mlabonne/harmless_alpaca for KL divergence — i.e. lower refusals and lower
distribution shift on benign inputs.
Independent verification
Refusal rate and general capability were re-checked with a separate script (different code, different refusal keyword set) rather than trusting the optimiser's own numbers:
Configuration
- Architecture
- Qwen3VLForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 36
- Hidden size
- 4,096
- Feed-forward size
- 12,288
- Attention heads
- 32
- Key/value heads
- 8
- Head dimension
- 128
- Vocabulary size
- 151,936
- Model type
- qwen3_vl
Identity and Version
- Repository
- pottokao/Qwen-Image-2.1-Text-Encoder-Heretic
- Publisher
- Pottokao
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- 8.8B parameters
- Languages
- Not stated by the source
- Revision
- 106be95540846f5905485637472144091b83e524
- First published
- 2026-09-20
- Last updated
- 2026-09-23
Files and Weights
16 files, 35.1 GB in total. The weights are 5 files totalling 35.1 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00004.safetensors | Weights | 4.9 GB | 6e217f5e222d |
| model-00002-of-00004.safetensors | Weights | 4.9 GB | 0226ae289505 |
| model-00003-of-00004.safetensors | Weights | 5.0 GB | 9c62f604a704 |
| model-00004-of-00004.safetensors | Weights | 2.7 GB | 8516dd606abf |
| qwen3vl_8b_bf16_heretic.safetensors | Weights | 17.5 GB | b1f17ffe6e04 |
| config.json | Configuration | 1.6 KB | — |
| generation_config.json | Configuration | 213 B | — |
| model.safetensors.index.json | Configuration | 67.8 KB | — |
| processor_config.json | Configuration | 1.4 KB | — |
| LICENSE | Documentation | 11.4 KB | — |
| NOTICE | Documentation | 812 B | — |
| README.md | Documentation | 6.8 KB | — |
| chat_template.jinja | Other | 5.3 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 11.4 MB | be75606093db |
| tokenizer_config.json | Tokenizer | 418 B | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 35.1 GB
Released by Pottokao through its official repository on Hugging Face. Read the license.
Built From
- Derived from Qwen/Qwen3-VL-8B-Instruct
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 35.1 GB |
| 16-bit | 17.5 GB |
| 8-bit | 8.8 GB |
| 4-bit | 4.4 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Built on This Model
- Quantized fromQwen-Image-2.1-Text-Encoder-Heretic-GGUF
- Derived fromQwen-Image-2.1-Text-Encoder-Heretic-GGUF
- Derived fromQwen-Image-2.1-Text-Encoder-Heretic-NVFP4
- Derived fromQwen-Image-2.1-Text-Encoder-Heretic-int8-convrot
- Derived fromQwen-Image-2.1-Text-Encoder-Heretic-W4A8
Questions About Qwen-Image-2.1-Text-Encoder-Heretic
How much GPU memory does Qwen-Image-2.1-Text-Encoder-Heretic need?
About 21 GB at 16-bit and 5.3 GB at 4-bit: the weights (8.8B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Qwen-Image-2.1-Text-Encoder-Heretic on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Qwen-Image-2.1-Text-Encoder-Heretic commercially?
Yes. Qwen-Image-2.1-Text-Encoder-Heretic is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is Qwen-Image-2.1-Text-Encoder-Heretic's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.