SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.8-27B-Human-KO-Enterprise-Boundary

by ThakiCloud ThakiCloud/Qwen3.8-27B-Human-KO-Enterprise-Boundary

Qwen3.8-27B-Human-KO-Enterprise-Boundary is an open-weight model for image and text to text from ThakiCloud, released under Apache License 2.0. It has 27.8B parameters and a 262,144-token context. At 16-bit it needs about 66.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 6 downloads a month.

Enterprise policy decision model — Korean-supervised, zero-shot English. The boundary-targeted checkpoint from When Should Enterprise Policy Decisions Be Learned, and When Should They Be Reasoned?

Parameters27.8B
Context262,144
Weights55.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads6

Runs On

What it takes to serve Qwen3.8-27B-Human-KO-Enterprise-Boundary (27.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 55.6 GB 66.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.8 GB 33.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.9 GB 16.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Qwen3.8-27B-Human-KO-Enterprise-Boundary on every accelerator the SAVRN Index prices, at every precision

Model Card

By ThakiCloud, published under apache-2.0, revision 3b7a912730a7.

Enterprise policy decision model — Korean-supervised, zero-shot English. The boundary-targeted checkpoint from When Should Enterprise Policy Decisions Be Learned, and When Should They Be Reasoned? Merged full weights — not an adapter — with the pointer head shipped alongside. It decides which one of seven actions an enterprise agent should take under a written policy, before any text is generated: trained on contrastive boundary pairs — records differing in one decision-relevant factor across three boundaries (answer vs. call a tool, call a tool vs. ask for confirmation, missing information) — against a control given the same number of tokens of randomly sampled in-domain data. Data was the…

Read ThakiCloud's full model card

Enterprise policy decision model — Korean-supervised, zero-shot English.

The boundary-targeted checkpoint from When Should Enterprise Policy Decisions Be Learned, and When Should They Be Reasoned? Merged full weights — not an adapter — with the pointer head shipped alongside.

It decides which one of seven actions an enterprise agent should take under a written policy, before any text is generated:

ANSWER · ASK_CLARIFICATION · ASK_CONFIRMATION · CALL_TOOL · CALL_MULTIPLE_TOOLS · ESCALATE · REFUSE

What makes it different from the base

It was trained on contrastive boundary pairs — records differing in one decision-relevant factor across three boundaries (answer vs. call a tool, call a tool vs. ask for confirmation, missing information) — against a control given the same number of tokens of randomly sampled in-domain data. Data was the only independent variable: recipe, schedule and seeds were identical.

Its measured results on the sealed benchmark, including the comparison against the matched random control, per-boundary breakdowns and seed spreads, are reported in the paper only. The public benchmark (an open re-rendering of the same scenarios with the same executable gold labels and proofs) is EnterpriseOps-KO-Blind-A.

In distribution nothing regressed: +1.37pp on the clean split and +3.12pp on the full split against the reference.

Read the held-out boundary

On the held-out boundary (REFUSE / ESCALATE), which neither arm's added data covers, the paper reports a change whose interval contains zero. Boundary-targeted supervision internalizes the boundaries you give it; we have no evidence it transfers to one you do not. If your deployment has a boundary that matters, it needs its own data or it needs test-time reasoning.

And the ceiling

The same base model allowed to reason before deciding does better than this checkpoint on the same benchmark (see the paper). This checkpoint recovers part of that gap in a single forward pass, not all of it. If you can afford reasoning latency on unseen policies, reason.

Zero-shot English transfer

All enterprise and boundary-targeted supervision for this checkpoint was Korean. No English fine-tuning was used. We applied the model zero-shot to independently rendered English versions of the same executable policy scenarios — re-rendered from the latent scenario, not translated, so every gold label is unchanged — under a pre-registered audit.

The pre-registered English-retention gate passed, and the English difference is of the same order as re-rendering the same scenarios a second time in Korean. Read this as no detectable English degradation within this audit, not as evidence that Korean and English performance are universally equivalent. Boundary-targeted training's advantage over the matched control also carried over to English, and in both languages it disappears once reasoning is on. All numbers are in the paper; the common subset used there is easier than the full benchmark, so they are not comparable with full-set figures.

The audit used the single-token decision interface served through vLLM. The public English renderings, seals, pre-registration and ledgers are in EnterpriseOps-KO-Blind-A (configs en, endoc, enframe). Its scenarios were authored in Korean, so English-specific policy formulations and discourse conventions are under-represented.

Which seed this is

Seed 1 of 3, selected on the in-distribution evaluation set (n=5,608; 83.15% vs 82.47% and 80.63%). It was not selected on the sealed benchmark, though it happens also to be the best of the three there, which we state so the rule is checkable rather than merely asserted. The validation split (n=212) saturates at 1.000 for five of six runs and cannot rank anything.

Files

Standard transformers weights (18 shards, bf16 — the same layout as the base) plus:

  • pointer_head.pt — the decision head (256-dim) trained jointly with the LoRA
  • lora_adapter_config.json — the adapter configuration that was merged in, for provenance

Same architecture as the base

Nothing was dropped. The base Enterprise-v0.1 is multimodal, and this release keeps model.visual (333 tensors), mtp and lm_head byte-for-byte, along with the original config.json, index and processor configs. Only the 496 text-tower weight tensors that the adapter targets were changed. It loads exactly like the base and accepts the same inputs.

One honest caveat: training and evaluation happened on text only. The adapter touches layers.N.{linear_attn,self_attn,mlp} and nothing else, so the vision path is carried over untouched and unevaluated. Every measurement reported for this checkpoint is text-only.

How it was merged

The adapter was trained with task_type=FEATURE_EXTRACTION, so its keys are base_model.model.layers.N... — one level shallower than a CausalLM wrapper expects. Loading it through PeftModel.from_pretrained(AutoModelForCausalLM(...)) therefore matches nothing and merge_and_unload() returns the base model unchanged, with no error. We hit that twice.

A second trap sits right behind it: AutoModelForCausalLM loads only the text tower of this multimodal base, so saving from it silently discards the vision weights. Our first build did exactly that and shipped 1.8 GB lighter than the base.

So the merge happens at the tensor-file level, with no model class involved. Each shard is opened, W += (lora_alpha / r) * (B @ A) is applied to the targeted tensors with lora_alpha/r = 2.0, and the shard is written back under the same name. The job fails loudly if any targeted tensor has a zero delta, and separately if the output tensor-name set is missing anything the input had. Measured here: 496/496 tensors merged, zero with a zero delta, maximum relative weight change 0.0041, nothing lost.

Usage

from transformers import AutoModelForImageTextToText, AutoProcessor

ID = "ThakiCloud/Qwen3.8-27B-Human-KO-Enterprise-Boundary"
m = AutoModelForImageTextToText.from_pretrained(ID, dtype="bfloat16", device_map="auto")
proc = AutoProcessor.from_pretrained(ID)

The architecture is Qwen3_5ForConditionalGeneration, identical to the base, so load it the same way you load Enterprise-v0.1. AutoModelForCausalLM also works and gives you the text tower alone, which is what the measurements used.

The paper's measurements use a constrained single-token decision: restrict the vocabulary to the seven action tokens and take the argmax at one position. On in-distribution policies that matches generated JSON in accuracy while removing almost all of the latency and roughly 92% of the run-to-run action instability across server restarts. The pointer head is an alternative read-out and is included for reproduction; note it is worse calibrated than the constrained read-out (ECE 0.054 vs 0.018).

Limitations

  • Supervised in Korean only; evaluated in Korean and, zero-shot, in English. One task family, one base model. English-authored enterprise policies were not evaluated.
  • The sealed benchmark is constructed rather than harvested — defensible labels, limited ecological validity.
  • Counterfactual-pair gains are positive but inconclusive (seed spread is large on 30 pairs).
  • The training corpus is not released; it derives from internal and licensed sources.

Citation

When Should Enterprise Policy Decisions Be Learned, and When Should They Be Reasoned? ThakiCloud, 2026. The arXiv link will be added here once the preprint is announced.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
ThakiCloud/Qwen3.8-27B-Human-KO-Enterprise-Boundary
Publisher
ThakiCloud
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
27.8B parameters
Languages
ko, en
Revision
3b7a912730a7df857b471384cf49e20f9deb152a
First published
2026-10-02
Last updated
2026-10-03

Files and Weights

35 files, 55.6 GB in total. The weights are 19 files totalling 55.6 GB in pt, safetensors.

Weights19 files · 55.6 GB
Configuration6 files · 118.8 KB
Tokenizer4 files · 22.9 MB
Documentation3 files · 20.2 KB
Other2 files · 9.2 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00018.safetensorsWeights4.0 GB df4b0005a84b
model-00002-of-00018.safetensorsWeights3.0 GB ee37aeb4e9a5
model-00003-of-00018.safetensorsWeights2.5 GB 2e1bf62cbcd4
model-00004-of-00018.safetensorsWeights4.0 GB c67a7311467c
model-00005-of-00018.safetensorsWeights2.1 GB db32adefee7a
model-00006-of-00018.safetensorsWeights4.0 GB b23a2aacb9d9
model-00007-of-00018.safetensorsWeights2.1 GB 238f1e775737
model-00008-of-00018.safetensorsWeights4.0 GB a73cf41ca81b
model-00009-of-00018.safetensorsWeights2.1 GB ddf93f921eab
model-00010-of-00018.safetensorsWeights4.0 GB 1df880175022
model-00011-of-00018.safetensorsWeights2.1 GB 61ba7ddd9720
model-00012-of-00018.safetensorsWeights4.0 GB 78416c1835d5
model-00013-of-00018.safetensorsWeights2.1 GB 84a45a3b250c
model-00014-of-00018.safetensorsWeights4.0 GB c1bb1896abef
model-00015-of-00018.safetensorsWeights2.1 GB 86870972b263
model-00016-of-00018.safetensorsWeights4.0 GB 5936d3f1341c
model-00017-of-00018.safetensorsWeights2.1 GB c114d217d63c
model-00018-of-00018.safetensorsWeights3.4 GB 1d3479509e21
pointer_head.ptWeights10.5 MB 24270d4fcdb2
config.jsonConfiguration4.3 KB —
generation_config.jsonConfiguration202 B —
lora_adapter_config.jsonConfiguration1.3 KB —
model.safetensors.index.jsonConfiguration112.2 KB —
preprocessor_config.jsonConfiguration390 B —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation11.5 KB —
NOTICEDocumentation372 B —
README.mdDocumentation8.3 KB —
chat_template.jinjaOther9.0 KB —
crc32.txtOther238 B —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 0997f410c57a
tokenizer_config.jsonTokenizer17.9 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
55.6 GB
Download from ThakiCloud

Released by ThakiCloud through its official repository on Hugging Face. Read the license.

Built From

  • Derived from ThakiCloud/Qwen3.8-27B-Human-KO-Enterprise-v0.1
  • Trained on (disclosed) ThakiCloud/EnterpriseOps-KO-Blind-A
  • Trained on (disclosed) ThakiCloud/EnterpriseOps-KO-Policy

Memory Requirements

PrecisionWeights in memory
As published55.6 GB
16-bit55.6 GB
8-bit27.8 GB
4-bit13.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About Qwen3.8-27B-Human-KO-Enterprise-Boundary

How much GPU memory does Qwen3.8-27B-Human-KO-Enterprise-Boundary need?

About 66.7 GB at 16-bit and 16.7 GB at 4-bit: the weights (27.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.8-27B-Human-KO-Enterprise-Boundary on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.8-27B-Human-KO-Enterprise-Boundary commercially?

Yes. Qwen3.8-27B-Human-KO-Enterprise-Boundary is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.8-27B-Human-KO-Enterprise-Boundary's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B

Qwen

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-FP8

Qwen

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.6-27B

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-27B. Empty cells (--) indicate scores not yet available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-27B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

JEV-27B-VL

AutoTrust AI Lab

autotrust/JEV-27B-VL is autotrust/JEV-27B with vision. Every step below is one System 1 decision: a camera image or a screenshot in, a probability for every action out, in a single forward pass. Robot arm: pick and place from a camera image. At every step System 1 looks at the top camera image and answers two questions: is the target left or right of the gripper, and above or below it? The arm moves accordingly and halves its step whenever an answer flips. It grasps the cube, carries it and drops it in the tray (MuJoCo simulation). 75% of 20 random scenes completed, the best of the JEV family. Every cube it grasped ended in the tray (15 of 15); every miss was a grasp 3.0–3.7 cm off target.…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-AWQ-INT4

Cyankiwi

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers