An INT8 W8A8 quantization of (the abliterated BF16 build of a merged LoRA finetune of Qwen/Qwen3.8-27B), for fast serving on GPUs where native FP8 is unavailable or undesirable. int-quantized, native CUTLASS INT8 tensor-core path. embedtokens, the vision tower, GatedDeltaNet recurrent gates, all norms. Note these are FP16, not the source's BF16 — the 16-bit residual precision of this build is float16 throughout (config.json: "dtype": "float16"). (llm-compressor), with the MTP drafter ablated in the rotated basis. The MTP head is unquantized but is not byte-identical to the BF16 build's: it carries the same residual-stream rotation and RMSNorm fold as the rest of the model, so the two heads…
Open-weight model · Image and text to text
Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8
by Akumaburn akumaburn/Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8
Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 is an open-weight model for image and text to text from Akumaburn, released under other. It has 27.4B parameters and a 262,144-token context. At 16-bit it needs about 65.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
An INT8 W8A8 quantization of (the abliterated BF16 build of ukisai/Swift-1.5-Qwen3.8-27b), for fast serving on GPUs where native FP8 is unavailable or undesirable. native CUTLASS INT8 tensor-core path.
Runs On
What it takes to serve Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 (27.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 54.7 GB | 65.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 27.4 GB | 32.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 13.7 GB | 16.4 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
Model Card
An INT8 W8A8 quantization of (the abliterated BF16 build of ukisai/Swift-1.5-Qwen3.8-27b), for fast serving on GPUs where native FP8 is unavailable or undesirable. native CUTLASS INT8 tensor-core path. No activation scales are stored on disk; they are computed per token at runtime. vision tower (333 tensors), GatedDeltaNet recurrent gates, all norms. These are FP16, not the source's BF16 — the 16-bit residual precision of this build is float16 throughout (config.json: "dtype": "float16"). with the MTP drafter re-expressed and ablated in the rotated basis. The rotation was verified before calibration by an fp64 whole-graph equivalence check (all 19 paths within 1.6–2.5e-4; every tensor…
Excerpt from the card by Akumaburn, licensed other.
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 64
- Hidden size
- 5,120
- Feed-forward size
- 17,408
- Attention heads
- 24
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Model type
- qwen3_5
- Quantization
- compressed-tensors
Identity and Version
- Repository
- akumaburn/Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8
- Publisher
- Akumaburn
- Task
- Image and text to text
- Modality
- Image and text
- Library
- vllm
- Parameters
- 27.4B parameters
- Languages
- mtp
- Revision
- 717ac9c071aaebdf7dc393c68bbb6c87e73bab32
- First published
- 2026-09-27
- Last updated
- 2026-09-27
Files and Weights
33 files, 31.3 GB in total. The weights are 16 files totalling 31.2 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00015.safetensors | Weights | 2.5 GB | 789da826ed3e |
| model-00002-of-00015.safetensors | Weights | 2.5 GB | 4cad0336d17a |
| model-00003-of-00015.safetensors | Weights | 2.0 GB | d2f03c3b947f |
| model-00004-of-00015.safetensors | Weights | 1.9 GB | 45f091f433af |
| model-00005-of-00015.safetensors | Weights | 2.0 GB | c7f672269c8b |
| model-00006-of-00015.safetensors | Weights | 1.9 GB | 9af5209a9452 |
| model-00007-of-00015.safetensors | Weights | 2.0 GB | 3285e2fdd1bf |
| model-00008-of-00015.safetensors | Weights | 2.0 GB | bc7ef1116b2c |
| model-00009-of-00015.safetensors | Weights | 2.0 GB | 3828cb8d8127 |
| model-00010-of-00015.safetensors | Weights | 1.9 GB | bd583e9f70bf |
| model-00011-of-00015.safetensors | Weights | 2.0 GB | 9d2fc983b903 |
| model-00012-of-00015.safetensors | Weights | 2.0 GB | 4cb89797f0ae |
| model-00013-of-00015.safetensors | Weights | 1.9 GB | 960b1187e179 |
| model-00014-of-00015.safetensors | Weights | 2.0 GB | 7ea1c7a4bc6c |
| model-00015-of-00015.safetensors | Weights | 1.7 GB | cb4dbfe0e0f9 |
| model-mtp.safetensors | Weights | 849.4 MB | 33a21e714b1d |
| config.json | Configuration | 21.0 KB | — |
| generation_config.json | Configuration | 214 B | — |
| model.safetensors.index.json | Configuration | 153.7 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| processor_config.json | Configuration | 1.2 KB | — |
| recipe.yaml | Configuration | 427 B | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| LICENSE | Documentation | 13.3 KB | — |
| LICENSE-APACHE-2.0 | Documentation | 11.5 KB | — |
| NOTICE | Documentation | 2.1 KB | — |
| README.md | Documentation | 10.4 KB | — |
| chat_template.jinja | Other | 9.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 06b9509352d2 |
| tokenizer_config.json | Tokenizer | 1.2 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- other
- Access
- Open weights, no gate
- Download size
- 31.2 GB
Released by Akumaburn through its official repository on Hugging Face. Read the license.
Built From
- Derived from akumaburn/Swift-1.5-Qwen3.8-27b-heretic
- Quantized from akumaburn/Swift-1.5-Qwen3.8-27b-heretic
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 31.2 GB |
| 16-bit | 54.7 GB |
| 8-bit | 27.4 GB |
| 4-bit | 13.7 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8
How much GPU memory does Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 need?
About 65.7 GB at 16-bit and 16.4 GB at 4-bit: the weights (27.4B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What license is Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8 released under?
other, as its publisher declares it. Read the license text before commercial use.
What is Swift-1.5-Qwen3.8-27b-heretic-SmoothQuant-W8A8-INT8's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory. pip install mlx-vlm python -m mlxvlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192 Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50…