SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8

by Akumaburn akumaburn/Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8

An INT8 W8A8 quantization of (the abliterated BF16 build of a merged LoRA finetune of Qwen/Qwen3.8-27B), for fast serving on GPUs where native FP8 is unavailable or undesirable. int-quantized, native CUTLASS INT8 tensor-core path.

Parameters27.4B
Context262,144
Weights31.2 GB
Licenseother
AccessOpen weights
Monthly Downloads

Runs On

What it takes to serve Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 (27.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 54.7 GB 65.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.4 GB 32.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.7 GB 16.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

An INT8 W8A8 quantization of (the abliterated BF16 build of a merged LoRA finetune of Qwen/Qwen3.8-27B), for fast serving on GPUs where native FP8 is unavailable or undesirable. int-quantized, native CUTLASS INT8 tensor-core path. embedtokens, the vision tower, GatedDeltaNet recurrent gates, all norms. Note these are FP16, not the source's BF16 — the 16-bit residual precision of this build is float16 throughout (config.json: "dtype": "float16"). (llm-compressor), with the MTP drafter ablated in the rotated basis. The MTP head is unquantized but is not byte-identical to the BF16 build's: it carries the same residual-stream rotation and RMSNorm fold as the rest of the model, so the two heads…

Excerpt from the card by Akumaburn, licensed other.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5
Quantization
compressed-tensors

Identity and Version

Repository
akumaburn/Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8
Publisher
Akumaburn
Task
Image and text to text
Modality
Image and text
Library
vllm
Parameters
27.4B parameters
Languages
mtp
Revision
88a0ef3430e878a2d66441d0198f90f458db48c5
First published
2026-09-18
Last updated
2026-09-18

Files and Weights

32 files, 31.3 GB in total. The weights are 16 files totalling 31.2 GB in safetensors.

Weights16 files · 31.2 GB
Configuration6 files · 176.2 KB
Tokenizer4 files · 30.1 MB
Documentation4 files · 33.5 KB
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00015.safetensorsWeights2.5 GB 987e5c0b3b41
model-00002-of-00015.safetensorsWeights2.5 GB 39fda25ce17c
model-00003-of-00015.safetensorsWeights2.0 GB fb517edf3330
model-00004-of-00015.safetensorsWeights1.9 GB 8370550a063b
model-00005-of-00015.safetensorsWeights2.0 GB 4439dd32d5c0
model-00006-of-00015.safetensorsWeights1.9 GB ae718174ec27
model-00007-of-00015.safetensorsWeights2.0 GB 101811723f5d
model-00008-of-00015.safetensorsWeights2.0 GB 6adc6408451b
model-00009-of-00015.safetensorsWeights2.0 GB 6814d6e0258a
model-00010-of-00015.safetensorsWeights1.9 GB d17d6864ac29
model-00011-of-00015.safetensorsWeights2.0 GB 85748547c086
model-00012-of-00015.safetensorsWeights2.0 GB 5e002d07c4ef
model-00013-of-00015.safetensorsWeights1.9 GB 0b53d65e7d47
model-00014-of-00015.safetensorsWeights2.0 GB 7ce457f95118
model-00015-of-00015.safetensorsWeights1.7 GB 6cb95db79b72
model-mtp.safetensorsWeights849.4 MB 550a9f03dcd6
config.jsonConfiguration21.0 KB
generation_config.jsonConfiguration257 B
model.safetensors.index.jsonConfiguration153.7 KB
preprocessor_config.jsonConfiguration390 B
recipe.yamlConfiguration427 B
video_preprocessor_config.jsonConfiguration385 B
LICENSEDocumentation13.3 KB
LICENSE-APACHE-2.0Documentation11.5 KB
NOTICEDocumentation1.0 KB
README.mdDocumentation7.6 KB
chat_template.jinjaOther9.0 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer3.4 MB
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.2 KB
vocab.jsonTokenizer6.7 MB

License and Download

License
other
Access
Open weights, no gate
Download size
31.2 GB
Download from Akumaburn

Released by Akumaburn through its official repository on Hugging Face.

Built From

  • Derived from akumaburn/Swift-Qwen3.8-27b-heretic
  • Quantized from akumaburn/Swift-Qwen3.8-27b-heretic

Memory Requirements

PrecisionWeights in memory
As published31.2 GB
16-bit54.7 GB
8-bit27.4 GB
4-bit13.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8

How much GPU memory does Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 need?

About 65.7 GB at 16-bit and 16.4 GB at 4-bit: the weights (27.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Swift-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B-MLX-4bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-8bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-6bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-5bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-Continuum-mxfp4-mlx

Gheorghe Chesler

(He slides three glasses across the bar—wine for Shakespeare, water for Data, and a mysterious blue liquid for Spock) This model is a merge of: - nightmedia/Qwen3.8-27B-Brainwaves - migtissera/Synthia-4-27B Brainwaves nightmedia/Qwen3.8-27B-Brainwaves migtissera/Synthia-4-27B Late-Night Architectural Takeaways The Quantization Shield (mxfp8 Master Pass): Hitting 0.735 ARC-C on the 8-bit layout proves that your Cold-Fusion flagship anchors and Migel Tissera's Synthia agent paths reached absolute geometric equilibrium. Instead of losing performance to the 0.596 Heretic collapse, the curved hypersphere calculation completely shielded the model's core intelligence. The Perplexity Sweet Spot…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

gemma-3-27b-it-int4-awq

Thien Tran

This is the QAT INT4 Flax checkpoint (from Kaggle) converted to HF+AWQ format for ease of use. AWQ was NOT used for quantization. You can find the conversion script convertflax.py in this model repo. NOTE: this is NOT the same as the official QAT INT4 GGUFs released here https://huggingface.co/collections/google/gemma-3-qat-67ee61ccacbf2be4195c265b Below is the original Model card from https://huggingface.co/google/gemma-3-27b-it [Gemma 3 Technical Report][g3-tech-report] [Responsible Generative AI Toolkit][rai-toolkit] [Gemma on Kaggle][kaggle-gemma] [Gemma on Vertex Model Garden][vertex-mg-gemma3] Summary description and brief definition of inputs and outputs. Gemma is a family of…

Open weights gemma 27.4B parameters transformers