This is an uncensored version of Qwen/Qwen3.6-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. - AWQ Marlin kernel supported (auto-converted by vLLM at runtime) - MTP speculative decoding supported out of the box - 110+ tok/s on a single A800 80GB (vLLM 0.21.0, MTP enabled, fp8 KV cache) - Risk of Sensitive or Controversial Outputs: This model’s safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated…
Open-weight model · Image and text to text
Swift-1.5-Qwen3.8-27b-heretic-W4A16
by Akumaburn akumaburn/Swift-1.5-Qwen3.8-27b-heretic-W4A16
Swift-1.5-Qwen3.8-27b-heretic-W4A16 is an open-weight model for image and text to text from Akumaburn, released under other. It has 6.3B parameters and a 262,144-token context. At 16-bit it needs about 15 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
A 18G INT4 build of for single-user decoding and low-VRAM deployment. (Marlin on Ampere). - Unrotated, so the DFlash2 drafter works. GatedDeltaNet gates, all norms.
Runs On
What it takes to serve Swift-1.5-Qwen3.8-27b-heretic-W4A16 (6.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 12.5 GB | 15.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 6.3 GB | 7.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 3.1 GB | 3.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
Swift-1.5-Qwen3.8-27b-heretic-W4A16 on every accelerator the SAVRN Index prices, at every precision
Model Card
A 18G INT4 build of for single-user decoding and low-VRAM deployment. (Marlin on Ampere). - Unrotated, so the DFlash2 drafter works. GatedDeltaNet gates, all norms. 4-bit costs 4.5× the rotated W8A8's quantization divergence (0.0264) and 1.5× the unrotated W8A8's (0.0791). Choose this build when the 18G footprint or single-user decode speed matters more than fidelity. Produced with Heretic v2.0.0.dev0: 260 TPE trials, seed 42, directional ablation on attn.oproj (16 modules), attn.outproj (48) and mlp.downproj (64). Baseline refusal score before abliteration: 98/100. Trial 140 was selected — Pareto index 1, not index 0. Index 0 (trial 161) scored 22/100 keywords at KL 0.1091; trial 140…
Excerpt from the card by Akumaburn, licensed other.
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 64
- Hidden size
- 5,120
- Feed-forward size
- 17,408
- Attention heads
- 24
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Model type
- qwen3_5
- Quantization
- compressed-tensors
Identity and Version
- Repository
- akumaburn/Swift-1.5-Qwen3.8-27b-heretic-W4A16
- Publisher
- Akumaburn
- Task
- Image and text to text
- Modality
- Image and text
- Library
- vllm
- Parameters
- 6.3B parameters
- Languages
- mtp
- Revision
- b4b5f07b113c10141d7195884dbf00680054497a
- First published
- 2026-09-27
- Last updated
- 2026-09-27
Files and Weights
26 files, 18.6 GB in total. The weights are 9 files totalling 18.6 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00009.safetensors | Weights | 2.5 GB | a24be664519c |
| model-00002-of-00009.safetensors | Weights | 2.5 GB | cec21a535170 |
| model-00003-of-00009.safetensors | Weights | 2.0 GB | cb189f2cbb0d |
| model-00004-of-00009.safetensors | Weights | 2.0 GB | 6a187022c9ba |
| model-00005-of-00009.safetensors | Weights | 2.0 GB | a244f1181854 |
| model-00006-of-00009.safetensors | Weights | 2.0 GB | 5164fdaeda9c |
| model-00007-of-00009.safetensors | Weights | 2.0 GB | 8114a3dc2d7e |
| model-00008-of-00009.safetensors | Weights | 2.0 GB | 1e3783f46165 |
| model-00009-of-00009.safetensors | Weights | 1.6 GB | 8ba3fc17da35 |
| config.json | Configuration | 20.6 KB | — |
| generation_config.json | Configuration | 214 B | — |
| model.safetensors.index.json | Configuration | 197.1 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| processor_config.json | Configuration | 1.2 KB | — |
| recipe.yaml | Configuration | 405 B | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| LICENSE | Documentation | 13.3 KB | — |
| LICENSE-APACHE-2.0 | Documentation | 11.5 KB | — |
| NOTICE | Documentation | 2.0 KB | — |
| README.md | Documentation | 9.0 KB | — |
| chat_template.jinja | Other | 9.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 06b9509352d2 |
| tokenizer_config.json | Tokenizer | 1.2 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- other
- Access
- Open weights, no gate
- Download size
- 18.6 GB
Released by Akumaburn through its official repository on Hugging Face. Read the license.
Built From
- Derived from akumaburn/Swift-1.5-Qwen3.8-27b-heretic
- Quantized from akumaburn/Swift-1.5-Qwen3.8-27b-heretic
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 18.6 GB |
| 16-bit | 12.5 GB |
| 8-bit | 6.3 GB |
| 4-bit | 3.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Swift-1.5-Qwen3.8-27b-heretic-W4A16
How much GPU memory does Swift-1.5-Qwen3.8-27b-heretic-W4A16 need?
About 15 GB at 16-bit and 3.8 GB at 4-bit: the weights (6.3B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Swift-1.5-Qwen3.8-27b-heretic-W4A16 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What license is Swift-1.5-Qwen3.8-27b-heretic-W4A16 released under?
other, as its publisher declares it. Read the license text before commercial use.
What is Swift-1.5-Qwen3.8-27b-heretic-W4A16's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
This repository contains an EXL3 3.0 bpw quantization of The original model and abliteration work are attributed to huihui-ai; this repository contains the EXL3 quantization produced by grimlee. The source model is an abliterated / uncensored derivative of Qwen3.8-27B. For details about the original model modification and its behavior, refer to the source model card. The checkpoint retains the Qwen3.8 multimodal model structure, including the vision component. The optional SM89 runtime work documented below does not change the model format or weights. The checkpoint uses the standard EXL3 format with a target body bitrate of 3.0 bits per weight. The target bitrate applies to the quantized…
Papers: https://arxiv.org/abs/2609.19745 (Vision-RL²) · https://arxiv.org/abs/2509.16944 (SD-RPN) SD-RPN stage-1 checkpoint: a self-distilled RoI predictor twig (K = 21, T = 3) trained on a frozen Qwen/Qwen3.5-4B. This is the initialisation of the Vision-RL² RL run YuhengSSS/VisionRL2-Qwen3.5-4B. The backbone weights are unchanged from the base model; only the three attached twig blocks are trained, from self-distilled attention pseudo-labels (no human RoI annotation). These weights need the modeling code in YuHengsss/VisionRL2. They are not loadable for RoI inference through a plain AutoModel / AutoModelForCausalLM call: the RoI gating path (heatmap head, peak-relative gate…
Below is the model card of Llava model 7b, which is copied from the original Llava model card that you can find here. Check out also the Google Colab demo to run Llava on a free-tier Google Colab instance: Or check out our Spaces demo! LLaVA is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal instruction-following data. It is an auto-regressive language model, based on the transformer architecture. LLaVA-v1.5-7B was trained in September 2023. Paper or resources for more information: https://llava-vl.github.io/ First, make sure to have transformers >= 4.35.3. The model supports multi-image and multi-prompt generation. Meaning that you can pass multiple…
Expert-paged build of Vontra/Qwen3.8-Flash-Next-MLX-4bit. The weights that are read a fraction at a time live in their own containers, so a machine loads what it needs rather than all Total 105.46 GiB. Of that, 103.94 GiB is the source build, whose bytes moved into containers rather than being copied, and 1.52 GiB is the draft head, which no published build of this model carries. Where the weights fit they are filled from experts.bin and the model runs the stock path at stock speed; where they do not, they stream from disk. Reading the machine decides that, not a flag. To override that: GBXPAGING=off holds the experts resident, GBXPLE=off holds the n-gram table resident. Checked at build…
Chandra 2 is a state of the art OCR model from Datalab that outputs markdown, HTML, and JSON. It is highly accurate at extracting text from images and PDFs, while preserving layout information. Try Chandra in the free playground, or use the hosted API for higher accuracy and speed. - 85.8% olmocr bench score (sota), 77.8% multilingual bench score (12% improvement over Chandra 1) - Significant improvements to math, tables, complex layouts - 90+ language support with major accuracy gains - Convert documents to markdown, HTML, or JSON with detailed layout information - Reconstructs forms accurately, including checkboxes - Strong performance with tables, math, and complex layouts - Extracts…