SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.5-KETI-HAECHI-27B

by Korea Electronics Technology Institute NLP KETI-NLP/Qwen3.5-KETI-HAECHI-27B

Qwen3.5-KETI-HAECHI-27B is an open-weight model for image and text to text from Korea Electronics Technology Institute NLP, released under Apache License 2.0. It has 27.4B parameters and a 262,144-token context. At 16-bit it needs about 65.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 67 downloads a month.

Qwen3.5-KETI-HAECHI-27B is a multimodal model derived from Qwen/Qwen3.5-27B. It was developed Korean OCR, and improving tool calling and multi-step, stateful agent execution.

Parameters27.4B
Context262,144
Weights54.7 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads67

Runs On

What it takes to serve Qwen3.5-KETI-HAECHI-27B (27.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 54.7 GB 65.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.4 GB 32.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.7 GB 16.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Qwen3.5-KETI-HAECHI-27B on every accelerator the SAVRN Index prices, at every precision

Model Card

By Korea Electronics Technology Institute NLP, published under apache-2.0, revision 259830921a2f.

Qwen3.5-KETI-HAECHI-27B is a multimodal model derived from Qwen/Qwen3.5-27B. It was developed Korean OCR, and improving tool calling and multi-step, stateful agent execution. Alongside these goals, the model preserves broad multimodal, language, and coding capabilities from the base model. heritage objects, answers questions grounded in heritage images, and reads Korean text from signs, scenes, rendered text, and public documents. - Tool calling and long-horizon task execution: selects and calls tools, carries information across multiple turns, tracks changing state, and works toward an end-to-end goal over several steps. follows Korean and English instructions, and performs visual…

Read Korea Electronics Technology Institute NLP's full model card

Qwen3.5-KETI-HAECHI-27B is a multimodal model derived from Qwen/Qwen3.5-27B. It was developed for two primary purposes: improving Korean cultural-heritage understanding and Korean OCR, and improving tool calling and multi-step, stateful agent execution. Alongside these goals, the model preserves broad multimodal, language, and coding capabilities from the base model.

Core capability profile

  • Korean cultural heritage and OCR: identifies the official names of heritage objects, answers questions grounded in heritage images, and reads Korean text from signs, scenes, rendered text, and public documents.
  • Tool calling and long-horizon task execution: selects and calls tools, carries information across multiple turns, tracks changing state, and works toward an end-to-end goal over several steps.
  • General multimodal understanding: interprets images and text together, follows Korean and English instructions, and performs visual reasoning beyond the specialized heritage domain.

Performance overview

The chart highlights representative metrics from each major capability area. Detailed benchmark tables appear in the Evaluation section.

Quickstart

Qwen3.5 support and AutoModelForMultimodalLM may be unavailable in older Transformers releases. The commands below reproduce the clean environment used to verify this repository. torchvision is required when AutoProcessor initializes the bundled video processor, including for image-only inference.

pip install "torch==2.9.1" "torchvision==0.24.1" accelerate pillow safetensors
pip install "transformers @ git+https://github.com/huggingface/transformers.git@c93057d4835cd31752bb56f59989dd27696eb45b"

When the model files are stored at the root of a Hugging Face repository:

import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "KETI-AIR/Qwen3.5-KETI-HAECHI-27B"
image_url = (
    f"https://huggingface.co/{model_id}/resolve/main/"
    "samples/qualitative/06_cheomseongdae.jpg"
)

processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
).eval()

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "url": image_url,
            },
            {
                "type": "text",
                "text": "사진 속 국가유산의 정확한 공식 명칭만 답하세요. 설명은 쓰지 마세요.",
            },
        ],
    }
]

inputs = processor.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    enable_thinking=False,
).to(model.device)

with torch.inference_mode():
    output_ids = model.generate(
        **inputs,
        max_new_tokens=32,
        do_sample=False,
    )

generated_ids = output_ids[:, inputs["input_ids"].shape[1]:]
answer = processor.batch_decode(generated_ids, skip_special_tokens=True)[0].strip()
print(answer)

# Example output from this checkpoint:
# 경주 첨성대

For the packaged directory in this release, use model_id = "./model" and image_url = "samples/qualitative/06_cheomseongdae.jpg". For a text-only request, omit the image item and retain only a {"type": "text", "text": ...} item in the message content.

Thinking mode is enabled by default in the bundled chat template. The example disables it for concise answers. Set enable_thinking=True for tasks that benefit from explicit reasoning, and adjust the generation budget accordingly.

Evaluation

All reported deltas are calculated as Qwen3.5-KETI-HAECHI-27B minus the untouched official Qwen/Qwen3.5-27B checkpoint. Percentage deltas are absolute percentage points. Both models used matched API, chat-template, thinking, stop-token, parser, and dataset settings.

Qualitative comparison with the base model

Korean cultural-heritage recognition
Target (reference) Qwen3.5-27B base Qwen3.5-KETI-HAECHI-27B

첨성대 (경주 첨성대)
No경주 대릉원 석물 Yes경주 첨성대

금동연가7년명여래입상
No금동약사여래입상 Yes금동연가7년명여래입상

백자 철화포도원숭이문 항아리
No분청사기철화포도문호 Yes백자 철화포도원숭이문 항아리

무령왕 금제 관식
No금동관식 Yes무령왕 금제 관식
Korean OCR
Font
Target (reference) Qwen3.5-27B base Qwen3.5-KETI-HAECHI-27B

비서
No日1人→ Yes비서

팔월
No파일 Yes팔월

떠들다
No따라들다 Yes떠들다

벌금
No별금 Yes벌금
Outdoor
Target (reference) Qwen3.5-27B base Qwen3.5-KETI-HAECHI-27B

아이원 아동발달상담센터
No아아원
아동발달상담센터
Yes아이원 아동발달상담센터

필 노래타운
No꿀 노래타운 Yes필 노래타운

써니네호프
No서니네호프 Yes써니네호프

못된
No몬된 Yes못된
Public exec
Target (reference) Qwen3.5-27B base Qwen3.5-KETI-HAECHI-27B

종합토지세
No홍합회포지
홍합회포지
홍합회포지
Yes종합토지세

영락공원묘지
No영락공원토지 Yes영락공원묘지

김해시장(전산정보과장)
No김해시청(전산정보과장) Yes김해시장(전산정보과장)

Quantitative benchmarks

Korean cultural-heritage performance
Benchmark Official base Qwen3.5-KETI-HAECHI-27B Delta
H400 direct exact 0.92% 38.07% +37.16 pp
H400 hard choice 26.45% 80.43% +53.98 pp
H400 knowledge image 11.45% 25.30% +13.85 pp
H400 knowledge text 5.97% 6.28% +0.31 pp

H400 is an internal evaluation set built around 400 Korean cultural-heritage items. It evaluates canonical-name recognition, fine-grained discrimination among visually similar heritage items, and heritage knowledge grounded in either images or text. H400 direct exact measures open-ended canonical-name retrieval, whereas H400 hard is closed-set selection. Their large gap indicates that recognition and discrimination are substantially stronger than exact free-form naming.

Korean (Hangul) OCR performance
Benchmark Official base Qwen3.5-KETI-HAECHI-27B Delta
ocr_font exact 19.66% 19.66% +0.00 pp
ocr_outdoor exact 36.13% 41.41% +5.27 pp
ocr_public_exec exact 4.17% 5.21% +1.04 pp

These Mammoth OCR results use strict exact match, where additional prose or differences in spacing and normalization can cause an otherwise useful answer to be marked incorrect.

Tool calling and long-horizon task performance
Tau2
Benchmark Official base Qwen3.5-KETI-HAECHI-27B Delta
Tau2 airline (n=32) 65.62% 71.88% +6.25 pp
Tau2 retail (n=32) 56.25% 65.62% +9.38 pp
Tau2 telecom (n=32) 87.50% 78.12% -9.38 pp
Tau2 weighted overall (n=96) 69.79% 71.88% +2.08 pp

Tau2 measures end-to-end success across multi-step airline, retail, and telecom workflows. It improves overall, especially in airline and retail, although the telecom domain declines.

BFCL V4
Benchmark Official base Qwen3.5-KETI-HAECHI-27B Delta
Overall accuracy 32.77% 32.11% -0.66 pp
Non-Live AST accuracy 89.60% 89.44% -0.16 pp
Non-Live Simple AST 79.42% 79.25% -0.17 pp
Non-Live Multiple AST 95.50% 95.50% +0.00 pp
Non-Live Parallel AST 91.50% 91.00% -0.50 pp
Non-Live Parallel Multiple AST 92.00% 92.00% +0.00 pp
Multi-turn accuracy 66.25% 64.50% -1.75 pp
Multi-turn base 76.50% 76.00% -0.50 pp
Multi-turn missing function 66.00% 64.00% -2.00 pp
Multi-turn missing parameter 53.00% 52.00% -1.00 pp
Multi-turn long context 69.50% 66.00% -3.50 pp

BFCL V4 measures structured function-call generation and multi-turn recovery from missing functions or parameters. The small overall regression indicates that tool calling and long-horizon performance remain sensitive to the environment and task protocol.

General multimodal and language retention
Benchmark Official base Qwen3.5-KETI-HAECHI-27B Delta
General-VL macro (6) 45.44% 77.60% +32.16 pp
MMBench DEV EN v1.1 30.42% 90.63% +60.22 pp
MMStar 39.40% 77.33% +37.93 pp
MMStar-KO 45.20% 72.33% +27.13 pp
KRETA 48.54% 86.15% +37.60 pp
MMMU-Pro 10c 44.28% 61.85% +17.57 pp
HallusionBench aAcc 64.77% 77.29% +12.51 pp
OpenCompass Core 63.49% 64.80% +1.31 pp
OpenCompass Extra 57.26% 57.61% +0.35 pp
OpenCompass Korean 17.18% 16.77% -0.41 pp

See the detailed score report for all 148 metrics, the benchmark guide for definitions and caveats, and the machine-readable results for downstream analysis.

Intended use

The model is intended for research and prototyping involving:

  • Korean cultural-heritage image identification and visual question answering;
  • Korean scene, sign, font, and public-document OCR;
  • image-grounded heritage knowledge retrieval;
  • multi-step, stateful tool-use workflows with explicit monitoring and recovery;
  • general image understanding and text generation; and
  • analysis of domain specialization and capability retention in multimodal models.

Responsible use

Verify cultural-property names and factual claims against authoritative catalogs or domain experts. Preserve human review for public descriptions, education, archival metadata, and research outputs. Do not treat model output as evidence of authenticity or provenance. Avoid submitting sensitive or personal documents for OCR unless the deployment provides appropriate privacy controls.

Limitations

  • OCR outputs may omit text, normalize spelling incorrectly, or add unsupported text.
  • The model can hallucinate names, dates, designations, provenance, and historical claims. Similar-looking artifacts and uncommon viewpoints are especially challenging.
  • General multilingual, video, robustness, demographic-bias, privacy, and safety behavior were not comprehensively evaluated for this release.
  • As with the base model, generated content may be inaccurate, biased, unsafe, or unsuitable for the user's context.

Core contributors

  • San Kim ([email protected]) — Led long-context agent task performance and development/management of the main training loop.
  • Byunggill Joe ([email protected]) — Secured GPU compute resources through the support program and led Korean cultural-heritage recognition/OCR performance.

Acknowledgements

Qwen3.5-KETI-HAECHI-27B was developed using compute resources provided through the 첨단 GPU 활용 지원 사업 of the National IT Industry Promotion Agency (NIPA; 정보통신산업진흥원).

License and attribution

This derivative checkpoint is released under the Apache License 2.0, following the upstream Qwen/Qwen3.5-27B license. Use of third-party data and generated outputs may be subject to additional terms. Please acknowledge the Qwen team and cite the upstream model when using this checkpoint in published work.

Citation

@misc{qwen35_keti_haechi_27b,
  title        = {Qwen3.5-KETI-HAECHI-27B},
  author       = {Kim, San and Joe, Byunggill},
  year         = {2026},
  howpublished = {Hugging Face model repository},
  url          = {https://huggingface.co/KETI-NLP/Qwen3.5-KETI-HAECHI-27B}
}

한국어 요약

이 모델은 Qwen/Qwen3.5-27B를 기반으로 두 가지 목적을 위해 개발한 멀티모달 모델입니다. 첫 번째 목적은 한국 문화유산의 정확한 명칭 식별, 이미지 기반 문화유산 질의응답, 한글 OCR 능력을 향상하는 것입니다. 두 번째 목적은 적절한 도구를 호출하고 여러 단계에 걸쳐 정보를 기억하며 변화하는 상태를 추적하는 장기작업 수행 능력을 향상하는 것입니다.

공식 베이스 모델과 동일한 조건의 내부 평가에서 H400 직접 식별 정확도는 0.92%에서 38.07%로, H400 유사 문화재 선택 정확도는 26.45%에서 80.43%로 향상되었습니다. 장기작업 평가인 Tau2 가중 점수는 69.79%에서 71.88%로 향상되었습니다. 다만 세부 환경에 따라 성능 차이가 있으므로, 실제 에이전트 작업에서는 단계별 상태 확인과 실패 복구 절차를 함께 적용하는 것을 권장합니다.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
KETI-NLP/Qwen3.5-KETI-HAECHI-27B
Publisher
Korea Electronics Technology Institute NLP
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
27.4B parameters
Languages
ko, en
Revision
259830921a2f45322018c496b123ce075511286e
First published
2026-09-04
Last updated
2026-10-05

Files and Weights

62 files, 54.7 GB in total. The weights are 1 file totalling 54.7 GB in safetensors.

Weights1 file · 54.7 GB
Configuration6 files · 37.6 KB
Tokenizer4 files · 22.9 MB
Documentation4 files · 46.1 KB
Other46 files · 12.4 MB
Repository1 file · 3.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights54.7 GB 6f4b2da2ba93
PACKAGE_MANIFEST.jsonConfiguration654 B —
config.jsonConfiguration3.6 KB —
generation_config.jsonConfiguration148 B —
preprocessor_config.jsonConfiguration390 B —
reports/EVALUATION_RESULTS.jsonConfiguration32.4 KB —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation11.3 KB —
README.mdDocumentation16.0 KB —
reports/BENCHMARK_GUIDE.mdDocumentation4.6 KB —
reports/DETAILED_SCORE_REPORT.mdDocumentation14.1 KB —
SHA256SUMSOther6.2 KB —
chat_template.jinjaOther7.8 KB —
reports/EVALUATION_RESULTS.csvOther12.4 KB —
samples/images/01_mammoth_and_ocr.jpgOther183.8 KB 4db5b69e2cf8
samples/images/02_reverse_candidates.jpgOther268.7 KB cd9f799b8a2e
samples/images/03_cultural_focus.jpgOther140.0 KB 072527d21c25
samples/images/04_h400_hs100.jpgOther473.9 KB efb5c63cb739
samples/images/cultural_focus_caption.jpgOther279.5 KB 562f08978d92
samples/images/cultural_focus_identity.jpgOther264.5 KB e9a3e13625ea
samples/images/h400_direct.jpgOther325.0 KB f2e47757ec4e
samples/images/h400_hard.jpgOther264.5 KB a50ae1be0768
samples/images/h400_knowledge_image.jpgOther401.2 KB c2ce57715548
samples/images/haechi.pngOther1.6 MB 5763215253dc
samples/images/haechi_keti.pngOther1.6 MB 5763215253dc
samples/images/haechi_personalization.pngOther1.8 MB da6a719c4fd6
samples/images/hs100_train.jpgOther329.3 KB eaa821540d2e
samples/images/hs100_unseen.jpgOther393.7 KB bde6f39c1ac8
samples/images/key_performance_overview.pngOther158.4 KB f7a079200a70
samples/images/mammoth_heritage_multi.jpgOther222.3 KB daf02ac937a0
samples/images/mammoth_heritage_reverse_A.jpgOther60.8 KB —
samples/images/mammoth_heritage_reverse_B.jpgOther44.8 KB —
samples/images/mammoth_heritage_reverse_C.jpgOther58.7 KB —
samples/images/mammoth_heritage_reverse_D.jpgOther61.0 KB —
samples/images/mammoth_heritage_simple.jpgOther80.8 KB —
samples/images/national_treasure.jpgOther72.8 KB —
samples/images/ocr_font.jpgOther17.0 KB —
samples/images/ocr_outdoor.jpgOther77.7 KB —
samples/images/ocr_public.jpgOther19.2 KB —
samples/qualitative/01_geumdong_yeonga_buddha.jpgOther198.0 KB 83e8f789aa29
samples/qualitative/02_white_porcelain_grape_monkey_jar.jpgOther151.7 KB eea45f66c3c6
samples/qualitative/03_muryeong_crown_ornament.jpgOther221.0 KB 123a1e1d627f
samples/qualitative/04_seokguram.jpgOther260.9 KB 7feba042ac5f
samples/qualitative/05_baekje_incense_burner.jpgOther215.9 KB aae0bf5501d8
samples/qualitative/06_cheomseongdae.jpgOther293.8 KB bcf6e809411f
samples/qualitative/07_janggyeong_panjeon.jpgOther276.1 KB 95cb0c5126f7
samples/qualitative/08_ocr_font_biseo.pngOther12.0 KB —
samples/qualitative/09_ocr_font_palwol.pngOther34.0 KB —
samples/qualitative/10_ocr_outdoor_aiwon.pngOther317.0 KB 5a8f93701a74
samples/qualitative/11_ocr_outdoor_pil.pngOther378.5 KB 7090be0b2816
samples/qualitative/12_ocr_public_landtax.pngOther86.4 KB —
samples/qualitative/13_ocr_font_tteodeulda.pngOther21.7 KB —
samples/qualitative/14_ocr_font_beolgeum.pngOther25.9 KB —
samples/qualitative/15_ocr_outdoor_sunny_ne_hope.pngOther283.8 KB bac586418440
samples/qualitative/16_ocr_outdoor_mot_doen.pngOther262.6 KB c157a57a3a8c
samples/qualitative/17_ocr_public_yeongnak_cemetery.pngOther80.3 KB —
samples/qualitative/18_ocr_public_gimhae_market.pngOther39.5 KB —
.gitattributesRepository3.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 5f9e4d4901a9
tokenizer_config.jsonTokenizer16.7 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
54.7 GB
Download from Korea Electronics Technology Institute NLP

Released by Korea Electronics Technology Institute NLP through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published54.7 GB
16-bit54.7 GB
8-bit27.4 GB
4-bit13.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.5-KETI-HAECHI-27B

How much GPU memory does Qwen3.5-KETI-HAECHI-27B need?

About 65.7 GB at 16-bit and 16.4 GB at 4-bit: the weights (27.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.5-KETI-HAECHI-27B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.5-KETI-HAECHI-27B commercially?

Yes. Qwen3.5-KETI-HAECHI-27B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.5-KETI-HAECHI-27B's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B-MLX-4bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-8bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-6bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-5bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

openthai2.0-qwen3.8-27b-MLX-4bit

iApp Technology

MLX 4-bit quantization (mlx-vlm) of for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory. pip install mlx-vlm python -m mlxvlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192 Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Under TensorFold (below), drafting comes from z-lab's DFlash2 drafter TensorFold serves this checkpoint as published, images included, through an OpenAI-compatible API (/v1/chat/completions, /v1/responses, /v1/messages). It runs on Apple silicon and on NVIDIA GPUs with compute capability 8.9+: RTX 40/50…

Open weights apache-2.0 27.4B parameters 262,144 tokens mlx

Model · Image and text to text

Qwen3.8-27B-Continuum-mxfp4-mlx

Gheorghe Chesler

(He slides three glasses across the bar—wine for Shakespeare, water for Data, and a mysterious blue liquid for Spock) This model is a merge of: - nightmedia/Qwen3.8-27B-Brainwaves - migtissera/Synthia-4-27B Brainwaves nightmedia/Qwen3.8-27B-Brainwaves migtissera/Synthia-4-27B Late-Night Architectural Takeaways The Quantization Shield (mxfp8 Master Pass): Hitting 0.735 ARC-C on the 8-bit layout proves that your Cold-Fusion flagship anchors and Migel Tissera's Synthia agent paths reached absolute geometric equilibrium. Instead of losing performance to the 0.596 Heretic collapse, the curved hypersphere calculation completely shielded the model's core intelligence. The Perplexity Sweet Spot…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers