An open Thai + English "System One" decision model. It does not generate text. Given a state (any text or JSON) and typed questions, it returns calibrated probabilities over the options in one forward pass: The request/response contract mirrors TypeSafe's POST /v1/systemone so code written for the TypeSafe SDK can be pointed at this model unchanged. Typical uses: ticket routing, moderation, intent detection, RAG relevance judging, LLM-output verification, and computer-use / browser-agent action selection (which element to click, which tool to call). Gated-DeltaNet / attention, 262k context), vision encoder removed, then continued-pretrained on ~5B tokens of Thai (web, Wikipedia, parallel…
Open weights
apache-2.0
753M parameters
262,144 tokens
iApp Technology ร่วมกับ สมาคมผู้ประกอบการปัญญาประดิษฐ์ประเทศไทย (AIEAT) · Apache 2.0 answers with Thai knowledge in natural, explanatory Thai, leads its base model and Typhoon on agentic tool use (BFCL) — and keeps its base model's general intelligence. Or serve the full bf16 weights with vLLM: One 80 GB GPU (~56 GB bf16). The LoRA adapter alone (7 GB, rank 64) is in adapter/ for serving on top of Qwen/Qwen3.8-27B with dynamic LoRA. The checkpoint ships the Qwen3.8 MTP draft head (mtp. tensors, textconfig.mtpnumhiddenlayers: 1), so vLLM can run self-speculative decoding: Speculative decoding is verified token-by-token by the main model, so outputs are identical with or without it — the…
Open weights
apache-2.0
27.8B parameters
262,144 tokens
MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory. pip install mlx-vlm python -m mlxvlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192 Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50…
Open weights
apache-2.0
27.4B parameters
262,144 tokens
mlx
Paper: Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation, AACL-IJCNLP 2026 Main Conference ChindaMT-4B is an open-weight Thai-English machine translation model fine-tuned from Qwen/Qwen3.5-4B on Grounded, a 1.97M-record dataset built by Reference-Grounded Data Curation (RGDC). It translates in both directions and follows auxiliary rules given in the prompt, such as terminology, register, length, and output format. It is one of three sizes in the ChindaMT family (4B, 2B, 0.8B). Plain translation. Same template for both directions; swap the language line and the source tag: With rules. Add a Rules: block between the language line and the source line.…
Open weights
apache-2.0
5.2B parameters
262,144 tokens
transformers
OpenThai-SystemOne is an open Thai + English System One decision model (0.8B, Apache-2.0). It does not generate text: given a state (text or JSON) and typed questions it returns probabilities: choice between named options, noul (yes/no) and score on an ordered scale. This repo is its Ollama build for Ollama's System One API (POST /v1/systemone, Ollama ≥ 0.35), the same API Ollama serves Nimble and Tev1 with. Everything runs on your machine; no API key. The curl example below uses the hf.co name; with the ollama.com pull, use "model": "iapp/openthai-systemone". (Probabilities rounded. The whole request is part of every question's prompt, so the same question can score slightly differently…
Open weights
apache-2.0
gguf