OpenThai2.0 - Opensource Thai Knowledge, Document, and Agentic AI (MLX 4-bit)
MLX 4-bit quantization (mlx-vlm) of
iapp/openthai2.0-qwen3.8-27b
for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory.
v2.0.3 — rebuilt from the v2.0.3 weights (Thai knowledge-recall fix). The v2.0.0
launch build stays available at revision tag v2.0.0.
Run
pip install mlx-vlm
python -m mlx_vlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \
--image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192
The model reasons before it answers — leave a large generation budget (8k+),
or replies may come back empty.
Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this
architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter
instead. Sanity-verified on-device: Thai factual prompts answered correctly.
Run with TensorFold (Mac or NVIDIA)
TensorFold serves this checkpoint as published,
images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses,
/v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+:
RTX 40/50, Hopper and Blackwell, including the DGX Spark. The commands below pin TensorFold
v0.6.5, because the project changes quickly.
TensorFold speeds up decoding with draft tokens from
z-lab/Qwen3.8-27B-DFlash2 (Apache-2.0)
and from copies of your own context. A draft is kept only when it equals the token the model
would produce on its own, so replies are the same with or without drafts. DFlash2 was trained
on base Qwen3.8-27B, not this Thai fine-tune, so expect fewer accepted drafts — and a smaller
speedup — on Thai text than TensorFold reports for the base model.
DGX Spark and other NVIDIA GPUs
docker run -it --gpus all --ipc=host --network host nvcr.io/nvidia/pytorch:26.07-py3
# inside the container:
python -m pip install 'tensorfold[vision] @ git+https://github.com/ashhart/[email protected]'
tensorfold pull iapp/openthai2.0-qwen3.8-27b-MLX-4bit z-lab/Qwen3.8-27B-DFlash2
tensorfold serve iapp/openthai2.0-qwen3.8-27b-MLX-4bit --name openthai2.0 \
--vision --max-tokens 8192 --host 0.0.0.0
On NVIDIA, TensorFold needs DFlash2 pulled before serving. To serve without it, add --no-drafts.
Mac (Apple silicon, 48 GB+)
python -m pip install 'tensorfold[vision] @ git+https://github.com/ashhart/[email protected]'
tensorfold pull iapp/openthai2.0-qwen3.8-27b-MLX-4bit z-lab/Qwen3.8-27B-DFlash2
tensorfold serve iapp/openthai2.0-qwen3.8-27b-MLX-4bit --name openthai2.0 \
--vision --max-tokens 8192
DFlash2 is optional on a Mac: skip pulling it, or pass --drafter none, and TensorFold drafts
from context copies only. On a 32 GB Mac, TensorFold's default memory budget (70% of RAM) is
too small for this model.
--max-tokens 8192 raises TensorFold's default reply limit of 4,096 tokens, because the model
reasons before it answers.
Read a Thai document
pip install openai
import base64
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="unused")
image = base64.b64encode(open("document.jpg", "rb").read()).decode()
reply = client.chat.completions.create(
model="openthai2.0",
messages=[{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
{"type": "text", "text": "อ่านข้อความในเอกสารนี้ทั้งหมด"},
]}],
)
print(reply.choices[0].message.content)
For text-only questions, send a plain string as content. Local images go in as data URLs;
remote image URLs are refused unless the server starts with --vision-urls.
Full benchmarks and model card: see the main bf16 repo.