FP8-dynamic variant of gemma-4-26B-A4B-it.
Search public pages, research tools, and SAVRN solutions.
Open-weight model · Image and text to text
by Vy Thông Nguyễn nvythong/Qwen3.8-27B-mlx-6Bit
Qwen3.8-27B-mlx-6Bit is an open-weight model for image and text to text from Vy Thông Nguyễn, released under Apache License 2.0. It has 26.9B parameters and a 262,144-token context. At 16-bit it needs about 64.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
The Model nvythong/Qwen3.8-27B-mlx-6Bit was converted to MLX format from Qwen/Qwen3.8-27B using mlx-lm version 0.31.2.
What it takes to serve Qwen3.8-27B-mlx-6Bit (26.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 53.8 GB | 64.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 26.9 GB | 32.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 13.4 GB | 16.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
Qwen3.8-27B-mlx-6Bit on every accelerator the SAVRN Index prices, at every precision
By Vy Thông Nguyễn, published under apache-2.0, revision a0b3c7ce31f2.
The Model nvythong/Qwen3.8-27B-mlx-6Bit was converted to MLX format from Qwen/Qwen3.8-27B using mlx-lm version 0.31.2.
The Model nvythong/Qwen3.8-27B-mlx-6Bit was converted to MLX format from Qwen/Qwen3.8-27B using mlx-lm version 0.31.2.
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("nvythong/Qwen3.8-27B-mlx-6Bit")
prompt="hello"
if hasattr(tokenizer, "apply_chat_template") and tokenizer.chat_template is not None:
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
response = generate(model, tokenizer, prompt=prompt, verbose=True)
13 files, 21.9 GB in total. The weights are 5 files totalling 21.9 GB in safetensors.
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00005.safetensors | Weights | 5.4 GB | 9df2771ada97 |
| model-00002-of-00005.safetensors | Weights | 5.3 GB | 2d90d2b7bc01 |
| model-00003-of-00005.safetensors | Weights | 5.3 GB | bfda8a00c8b4 |
| model-00004-of-00005.safetensors | Weights | 4.8 GB | 6f7b36885f53 |
| model-00005-of-00005.safetensors | Weights | 1.0 GB | 88c510a06aa9 |
| config.json | Configuration | 4.1 KB | — |
| generation_config.json | Configuration | 202 B | — |
| model.safetensors.index.json | Configuration | 189.8 KB | — |
| README.md | Documentation | 884 B | — |
| chat_template.jinja | Other | 9.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 87a7830d63fc |
| tokenizer_config.json | Tokenizer | 1.1 KB | — |
Released by Vy Thông Nguyễn through its official repository on Hugging Face. Read the license.
| Precision | Weights in memory |
|---|---|
| As published | 21.9 GB |
| 16-bit | 53.8 GB |
| 8-bit | 26.9 GB |
| 4-bit | 13.4 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
About 64.6 GB at 16-bit and 16.1 GB at 4-bit: the weights (26.9B parameters) plus a working margin. A long context needs more.
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Yes. Qwen3.8-27B-mlx-6Bit is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
262,144 tokens, from the maximum position embeddings in its published configuration.
FP8-dynamic variant of gemma-4-26B-A4B-it.
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory. pip install mlx-vlm python -m mlxvlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192 Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50…