SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Swift-1.5-Qwen3.8-27b-heretic

by Akumaburn akumaburn/Swift-1.5-Qwen3.8-27b-heretic

Swift-1.5-Qwen3.8-27b-heretic is an open-weight model for image and text to text from Akumaburn, released under other. It has 27.4B parameters and a 262,144-token context. At 16-bit it needs about 65.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

An abliterated build of ukisai/Swift-1.5-Qwen3.8-27b, UkisAI's reasoning-efficient derivative of Qwen3.8-27B. Refusal directions are removed by directional ablation; nothing else is retrained.

Parameters27.4B
Context262,144
Weights55.6 GB
Licenseother
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Swift-1.5-Qwen3.8-27b-heretic (27.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 54.7 GB 65.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.4 GB 32.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.7 GB 16.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Swift-1.5-Qwen3.8-27b-heretic on every accelerator the SAVRN Index prices, at every precision

Model Card

An abliterated build of ukisai/Swift-1.5-Qwen3.8-27b, UkisAI's reasoning-efficient derivative of Qwen3.8-27B. Refusal directions are removed by directional ablation; nothing else is retrained. identical to the source's), lmhead, embedtokens, all norms, and every GatedDeltaNet recurrent gate. - Size 52G, 1,199 tensors (333 vision, 15 MTP). Produced with Heretic v2.0.0.dev0: 260 TPE trials, seed 42, directional ablation on attn.oproj (16 modules), attn.outproj (48) and mlp.downproj (64). Baseline refusal score before abliteration: 98/100. Trial 140 was selected — Pareto index 1, not index 0. Index 0 (trial 161) scored 22/100 keywords at KL 0.1091; trial 140 scores 23/100 at KL 0.0602. One…

Excerpt from the card by Akumaburn, licensed other.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
akumaburn/Swift-1.5-Qwen3.8-27b-heretic
Publisher
Akumaburn
Task
Image and text to text
Modality
Image and text
Library
vllm
Parameters
27.4B parameters
Languages
mtp
Revision
529f2bef62bbc2f757b8f2fa58d7e3dcb5c4da1c
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

29 files, 55.6 GB in total. The weights are 13 files totalling 55.6 GB in safetensors.

Weights13 files · 55.6 GB
Configuration6 files · 118.0 KB
Tokenizer4 files · 30.1 MB
Documentation4 files · 35.5 KB
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00012.safetensorsWeights2.5 GB a24be664519c
model-00002-of-00012.safetensorsWeights4.8 GB a32722bef609
model-00003-of-00012.safetensorsWeights5.0 GB 7e0cd0fed8eb
model-00004-of-00012.safetensorsWeights4.9 GB f82be4c6fdfa
model-00005-of-00012.safetensorsWeights5.0 GB 5ac19103095e
model-00006-of-00012.safetensorsWeights4.9 GB 4ae4206c1428
model-00007-of-00012.safetensorsWeights4.9 GB 59a953a59654
model-00008-of-00012.safetensorsWeights5.0 GB 89fe42fddc34
model-00009-of-00012.safetensorsWeights5.0 GB 02eb020cfd4c
model-00010-of-00012.safetensorsWeights4.9 GB 7588f985344c
model-00011-of-00012.safetensorsWeights5.0 GB c2183b788129
model-00012-of-00012.safetensorsWeights2.8 GB 659154f40a09
model-mtp.safetensorsWeights849.4 MB 90fa0e3eed5a
config.jsonConfiguration3.7 KB —
generation_config.jsonConfiguration213 B —
model.safetensors.index.jsonConfiguration112.1 KB —
preprocessor_config.jsonConfiguration390 B —
processor_config.jsonConfiguration1.2 KB —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation13.3 KB —
LICENSE-APACHE-2.0Documentation11.5 KB —
NOTICEDocumentation2.0 KB —
README.mdDocumentation8.6 KB —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.2 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
other
Access
Open weights, no gate
Download size
55.6 GB
Download from Akumaburn

Released by Akumaburn through its official repository on Hugging Face. Read the license.

Built From

  • Derived from ukisai/Swift-1.5-Qwen3.8-27b

Memory Requirements

PrecisionWeights in memory
As published55.6 GB
16-bit54.7 GB
8-bit27.4 GB
4-bit13.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About Swift-1.5-Qwen3.8-27b-heretic

How much GPU memory does Swift-1.5-Qwen3.8-27b-heretic need?

About 65.7 GB at 16-bit and 16.4 GB at 4-bit: the weights (27.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Swift-1.5-Qwen3.8-27b-heretic on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Swift-1.5-Qwen3.8-27b-heretic released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Swift-1.5-Qwen3.8-27b-heretic's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B-MLX-4bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-8bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-6bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-5bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

openthai2.0-qwen3.8-27b-MLX-4bit

iApp Technology

MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory. pip install mlx-vlm python -m mlxvlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192 Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50…

Open weights apache-2.0 27.4B parameters 262,144 tokens mlx

Qwen3.5-KETI-HAECHI-27B is a multimodal model derived from Qwen/Qwen3.5-27B. It was developed Korean OCR, and improving tool calling and multi-step, stateful agent execution. Alongside these goals, the model preserves broad multimodal, language, and coding capabilities from the base model. heritage objects, answers questions grounded in heritage images, and reads Korean text from signs, scenes, rendered text, and public documents. - Tool calling and long-horizon task execution: selects and calls tools, carries information across multiple turns, tracks changing state, and works toward an end-to-end goal over several steps. follows Korean and English instructions, and performs visual…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers