SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.8-27B-mlx-6Bit

by Vy Thông Nguyễn nvythong/Qwen3.8-27B-mlx-6Bit

Qwen3.8-27B-mlx-6Bit is an open-weight model for image and text to text from Vy Thông Nguyễn, released under Apache License 2.0. It has 26.9B parameters and a 262,144-token context. At 16-bit it needs about 64.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

The Model nvythong/Qwen3.8-27B-mlx-6Bit was converted to MLX format from Qwen/Qwen3.8-27B using mlx-lm version 0.31.2.

Parameters26.9B
Context262,144
Weights21.9 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Qwen3.8-27B-mlx-6Bit (26.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 53.8 GB 64.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 26.9 GB 32.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.4 GB 16.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Qwen3.8-27B-mlx-6Bit on every accelerator the SAVRN Index prices, at every precision

Model Card

By Vy Thông Nguyễn, published under apache-2.0, revision a0b3c7ce31f2.

The Model nvythong/Qwen3.8-27B-mlx-6Bit was converted to MLX format from Qwen/Qwen3.8-27B using mlx-lm version 0.31.2.

Read Vy Thông Nguyễn's full model card

The Model nvythong/Qwen3.8-27B-mlx-6Bit was converted to MLX format from Qwen/Qwen3.8-27B using mlx-lm version 0.31.2.

Use with mlx

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("nvythong/Qwen3.8-27B-mlx-6Bit")

prompt="hello"

if hasattr(tokenizer, "apply_chat_template") and tokenizer.chat_template is not None:
    messages = [{"role": "user", "content": prompt}]
    prompt = tokenizer.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True
    )

response = generate(model, tokenizer, prompt=prompt, verbose=True)

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
nvythong/Qwen3.8-27B-mlx-6Bit
Publisher
Vy Thông Nguyễn
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
26.9B parameters
Languages
mlx
Revision
a0b3c7ce31f243f94ba628f4feb4d8acfe1f5d01
First published
2026-10-03
Last updated
2026-10-03

Files and Weights

13 files, 21.9 GB in total. The weights are 5 files totalling 21.9 GB in safetensors.

Weights5 files · 21.9 GB
Configuration3 files · 194.1 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 884 B
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00005.safetensorsWeights5.4 GB 9df2771ada97
model-00002-of-00005.safetensorsWeights5.3 GB 2d90d2b7bc01
model-00003-of-00005.safetensorsWeights5.3 GB bfda8a00c8b4
model-00004-of-00005.safetensorsWeights4.8 GB 6f7b36885f53
model-00005-of-00005.safetensorsWeights1.0 GB 88c510a06aa9
config.jsonConfiguration4.1 KB —
generation_config.jsonConfiguration202 B —
model.safetensors.index.jsonConfiguration189.8 KB —
README.mdDocumentation884 B —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 87a7830d63fc
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
21.9 GB
Download from Vy Thông Nguyễn

Released by Vy Thông Nguyễn through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published21.9 GB
16-bit53.8 GB
8-bit26.9 GB
4-bit13.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.8-27B-mlx-6Bit

How much GPU memory does Qwen3.8-27B-mlx-6Bit need?

About 64.6 GB at 16-bit and 16.1 GB at 4-bit: the weights (26.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.8-27B-mlx-6Bit on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.8-27B-mlx-6Bit commercially?

Yes. Qwen3.8-27B-mlx-6Bit is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.8-27B-mlx-6Bit's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B-MLX-4bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-8bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-6bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-5bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

openthai2.0-qwen3.8-27b-MLX-4bit

iApp Technology

MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory. pip install mlx-vlm python -m mlxvlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192 Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50…

Open weights apache-2.0 27.4B parameters 262,144 tokens mlx