SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

openthai2.0-qwen3.8-27b-MLX-4bit

by iApp Technology iapp/openthai2.0-qwen3.8-27b-MLX-4bit

openthai2.0-qwen3.8-27b-MLX-4bit is an open-weight model for image and text to text from iApp Technology, released under Apache License 2.0. It has 27.4B parameters and a 262,144-token context. At 16-bit it needs about 65.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 187 downloads a month.

MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory.

Parameters27.4B
Context262,144
Weights17.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads187

Runs On

What it takes to serve openthai2.0-qwen3.8-27b-MLX-4bit (27.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 54.7 GB 65.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.4 GB 32.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.7 GB 16.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

openthai2.0-qwen3.8-27b-MLX-4bit on every accelerator the SAVRN Index prices, at every precision

Model Card

By iApp Technology, published under apache-2.0, revision 50280b0926a1.

MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory. pip install mlx-vlm python -m mlxvlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192 Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50…

Read iApp Technology's full model card

OpenThai2.0 - Opensource Thai Knowledge, Document, and Agentic AI (MLX 4-bit)

MLX 4-bit quantization (mlx-vlm) of iapp/openthai2.0-qwen3.8-27b for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory.

v2.0.3 — rebuilt from the v2.0.3 weights (Thai knowledge-recall fix). The v2.0.0 launch build stays available at revision tag v2.0.0.

Run

pip install mlx-vlm
python -m mlx_vlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \
  --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192

The model reasons before it answers — leave a large generation budget (8k+), or replies may come back empty.

Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter instead. Sanity-verified on-device: Thai factual prompts answered correctly.

Run with TensorFold (Mac or NVIDIA)

TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50, Hopper and Blackwell, including the DGX Spark. The commands below pin TensorFold v0.6.5, because the project changes quickly.

TensorFold speeds up decoding with draft tokens from z-lab/Qwen3.8-27B-DFlash2 (Apache-2.0) and from copies of your own context. A draft is kept only when it equals the token the model would produce on its own, so replies are the same with or without drafts. DFlash2 was trained on base Qwen3.8-27B, not this Thai fine-tune, so expect fewer accepted drafts — and a smaller speedup — on Thai text than TensorFold reports for the base model.

DGX Spark and other NVIDIA GPUs

docker run -it --gpus all --ipc=host --network host nvcr.io/nvidia/pytorch:26.07-py3
# inside the container:
python -m pip install 'tensorfold[vision] @ git+https://github.com/ashhart/[email protected]'
tensorfold pull iapp/openthai2.0-qwen3.8-27b-MLX-4bit z-lab/Qwen3.8-27B-DFlash2
tensorfold serve iapp/openthai2.0-qwen3.8-27b-MLX-4bit --name openthai2.0 \
  --vision --max-tokens 8192 --host 0.0.0.0

On NVIDIA, TensorFold needs DFlash2 pulled before serving. To serve without it, add --no-drafts.

Mac (Apple silicon, 48 GB+)

python -m pip install 'tensorfold[vision] @ git+https://github.com/ashhart/[email protected]'
tensorfold pull iapp/openthai2.0-qwen3.8-27b-MLX-4bit z-lab/Qwen3.8-27B-DFlash2
tensorfold serve iapp/openthai2.0-qwen3.8-27b-MLX-4bit --name openthai2.0 \
  --vision --max-tokens 8192

DFlash2 is optional on a Mac: skip pulling it, or pass --drafter none, and TensorFold drafts from context copies only. On a 32 GB Mac, TensorFold's default memory budget (70% of RAM) is too small for this model.

--max-tokens 8192 raises TensorFold's default reply limit of 4,096 tokens, because the model reasons before it answers.

Read a Thai document

pip install openai
import base64
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="unused")
image = base64.b64encode(open("document.jpg", "rb").read()).decode()

reply = client.chat.completions.create(
    model="openthai2.0",
    messages=[{"role": "user", "content": [
        {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
        {"type": "text", "text": "อ่านข้อความในเอกสารนี้ทั้งหมด"},
    ]}],
)
print(reply.choices[0].message.content)

For text-only questions, send a plain string as content. Local images go in as data URLs; remote image URLs are refused unless the server starts with --vision-urls.

Full benchmarks and model card: see the main bf16 repo.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
iapp/openthai2.0-qwen3.8-27b-MLX-4bit
Publisher
iApp Technology
Task
Image and text to text
Modality
Image and text
Library
mlx
Parameters
27.4B parameters
Languages
th, en
Revision
50280b0926a19ac7df24bd16b6029e05d98ab63e
First published
2026-08-25
Last updated
2026-10-04

Files and Weights

21 files, 17.0 GB in total. The weights are 4 files totalling 17.0 GB in safetensors.

Weights4 files · 17.0 GB
Configuration12 files · 323.1 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 4.5 KB
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
adapter/adapter_model.safetensorsWeights934.0 MB 25988c4cc05a
model-00001-of-00003.safetensorsWeights5.3 GB e03c3d5ec46b
model-00002-of-00003.safetensorsWeights5.4 GB 6ff313adf88d
model-00003-of-00003.safetensorsWeights5.4 GB 002c079a31ae
adapter/adapter_config.jsonConfiguration1.1 KB —
adapter/additional_config.jsonConfiguration67 B —
adapter/args.jsonConfiguration20.2 KB —
adapter/trainer_state.jsonConfiguration21.3 KB —
adapter/zero_to_fp32.pyConfiguration34.9 KB —
args.jsonConfiguration20.2 KB —
config.jsonConfiguration5.1 KB —
generation_config.jsonConfiguration214 B —
model.safetensors.index.jsonConfiguration218.3 KB —
preprocessor_config.jsonConfiguration390 B —
processor_config.jsonConfiguration991 B —
video_preprocessor_config.jsonConfiguration385 B —
README.mdDocumentation4.5 KB —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.2 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
17.0 GB
Download from iApp Technology

Released by iApp Technology through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published17.0 GB
16-bit54.7 GB
8-bit27.4 GB
4-bit13.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About openthai2.0-qwen3.8-27b-MLX-4bit

How much GPU memory does openthai2.0-qwen3.8-27b-MLX-4bit need?

About 65.7 GB at 16-bit and 16.4 GB at 4-bit: the weights (27.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run openthai2.0-qwen3.8-27b-MLX-4bit on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use openthai2.0-qwen3.8-27b-MLX-4bit commercially?

Yes. openthai2.0-qwen3.8-27b-MLX-4bit is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is openthai2.0-qwen3.8-27b-MLX-4bit's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B-MLX-4bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-8bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-6bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-5bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Qwen3.5-KETI-HAECHI-27B is a multimodal model derived from Qwen/Qwen3.5-27B. It was developed Korean OCR, and improving tool calling and multi-step, stateful agent execution. Alongside these goals, the model preserves broad multimodal, language, and coding capabilities from the base model. heritage objects, answers questions grounded in heritage images, and reads Korean text from signs, scenes, rendered text, and public documents. - Tool calling and long-horizon task execution: selects and calls tools, carries information across multiple turns, tracks changing state, and works toward an end-to-end goal over several steps. follows Korean and English instructions, and performs visual…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-Continuum-mxfp4-mlx

Gheorghe Chesler

(He slides three glasses across the bar—wine for Shakespeare, water for Data, and a mysterious blue liquid for Spock) This model is a merge of: - nightmedia/Qwen3.8-27B-Brainwaves - migtissera/Synthia-4-27B Brainwaves nightmedia/Qwen3.8-27B-Brainwaves migtissera/Synthia-4-27B Late-Night Architectural Takeaways The Quantization Shield (mxfp8 Master Pass): Hitting 0.735 ARC-C on the 8-bit layout proves that your Cold-Fusion flagship anchors and Migel Tissera's Synthia agent paths reached absolute geometric equilibrium. Instead of losing performance to the 0.596 Heretic collapse, the curved hypersphere calculation completely shielded the model's core intelligence. The Perplexity Sweet Spot…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers