SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8

by Akumaburn akumaburn/Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8

Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 is an open-weight model for image and text to text from Akumaburn, released under other. It has 27.4B parameters and a 262,144-token context. At 16-bit it needs about 65.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

An INT8 W8A8 quantization of (the abliterated BF16 build of ukisai/Swift-1.5-Qwen3.8-27b), for fast serving on GPUs where native FP8 is unavailable or undesirable. native CUTLASS INT8 tensor-core path.

Parameters27.4B
Context262,144
Weights31.2 GB
Licenseother
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 (27.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 54.7 GB 65.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.4 GB 32.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.7 GB 16.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 on every accelerator the SAVRN Index prices, at every precision

Model Card

An INT8 W8A8 quantization of (the abliterated BF16 build of ukisai/Swift-1.5-Qwen3.8-27b), for fast serving on GPUs where native FP8 is unavailable or undesirable. native CUTLASS INT8 tensor-core path. No activation scales are stored on disk; they are computed per token at runtime. vision tower (333 tensors), GatedDeltaNet recurrent gates, all norms. These are FP16, not the source's BF16 — the 16-bit residual precision of this build is float16 throughout (config.json: "dtype": "float16"). with the MTP drafter re-expressed and ablated in the rotated basis. The rotation was verified before calibration by an fp64 whole-graph equivalence check (all 19 paths within 1.6–2.5e-4; every tensor…

Excerpt from the card by Akumaburn, licensed other.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5
Quantization
compressed-tensors

Identity and Version

Repository
akumaburn/Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8
Publisher
Akumaburn
Task
Image and text to text
Modality
Image and text
Library
vllm
Parameters
27.4B parameters
Languages
mtp
Revision
717ac9c071aaebdf7dc393c68bbb6c87e73bab32
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

33 files, 31.3 GB in total. The weights are 16 files totalling 31.2 GB in safetensors.

Weights16 files · 31.2 GB
Configuration7 files · 177.3 KB
Tokenizer4 files · 30.1 MB
Documentation4 files · 37.4 KB
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00015.safetensorsWeights2.5 GB 789da826ed3e
model-00002-of-00015.safetensorsWeights2.5 GB 4cad0336d17a
model-00003-of-00015.safetensorsWeights2.0 GB d2f03c3b947f
model-00004-of-00015.safetensorsWeights1.9 GB 45f091f433af
model-00005-of-00015.safetensorsWeights2.0 GB c7f672269c8b
model-00006-of-00015.safetensorsWeights1.9 GB 9af5209a9452
model-00007-of-00015.safetensorsWeights2.0 GB 3285e2fdd1bf
model-00008-of-00015.safetensorsWeights2.0 GB bc7ef1116b2c
model-00009-of-00015.safetensorsWeights2.0 GB 3828cb8d8127
model-00010-of-00015.safetensorsWeights1.9 GB bd583e9f70bf
model-00011-of-00015.safetensorsWeights2.0 GB 9d2fc983b903
model-00012-of-00015.safetensorsWeights2.0 GB 4cb89797f0ae
model-00013-of-00015.safetensorsWeights1.9 GB 960b1187e179
model-00014-of-00015.safetensorsWeights2.0 GB 7ea1c7a4bc6c
model-00015-of-00015.safetensorsWeights1.7 GB cb4dbfe0e0f9
model-mtp.safetensorsWeights849.4 MB 33a21e714b1d
config.jsonConfiguration21.0 KB —
generation_config.jsonConfiguration214 B —
model.safetensors.index.jsonConfiguration153.7 KB —
preprocessor_config.jsonConfiguration390 B —
processor_config.jsonConfiguration1.2 KB —
recipe.yamlConfiguration427 B —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation13.3 KB —
LICENSE-APACHE-2.0Documentation11.5 KB —
NOTICEDocumentation2.1 KB —
README.mdDocumentation10.4 KB —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.2 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
other
Access
Open weights, no gate
Download size
31.2 GB
Download from Akumaburn

Released by Akumaburn through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published31.2 GB
16-bit54.7 GB
8-bit27.4 GB
4-bit13.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8

How much GPU memory does Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 need?

About 65.7 GB at 16-bit and 16.4 GB at 4-bit: the weights (27.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

An INT8 W8A8 quantization of (the abliterated BF16 build of a merged LoRA finetune of Qwen/Qwen3.8-27B), for fast serving on GPUs where native FP8 is unavailable or undesirable. int-quantized, native CUTLASS INT8 tensor-core path. embedtokens, the vision tower, GatedDeltaNet recurrent gates, all norms. Note these are FP16, not the source's BF16 — the 16-bit residual precision of this build is float16 throughout (config.json: "dtype": "float16"). (llm-compressor), with the MTP drafter ablated in the rotated basis. The MTP head is unquantized but is not byte-identical to the BF16 build's: it carries the same residual-stream rotation and RMSNorm fold as the rest of the model, so the two heads…

Open weights other 27.4B parameters 262,144 tokens vllm

Model · Image and text to text

Qwen3.8-27B-MLX-4bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-8bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-6bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-5bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

openthai2.0-qwen3.8-27b-MLX-4bit

iApp Technology

MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory. pip install mlx-vlm python -m mlxvlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192 Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50…

Open weights apache-2.0 27.4B parameters 262,144 tokens mlx