Enterprise policy decision model — Korean-supervised, zero-shot English.
The boundary-targeted checkpoint from When Should Enterprise Policy Decisions Be Learned, and
When Should They Be Reasoned? Merged full weights — not an adapter — with the pointer head
shipped alongside.
It decides which one of seven actions an enterprise agent should take under a written
policy, before any text is generated:
ANSWER · ASK_CLARIFICATION · ASK_CONFIRMATION · CALL_TOOL ·
CALL_MULTIPLE_TOOLS · ESCALATE · REFUSE
What makes it different from the base
It was trained on contrastive boundary pairs — records differing in one decision-relevant
factor across three boundaries (answer vs. call a tool, call a tool vs. ask for confirmation,
missing information) — against a control given the same number of tokens of randomly
sampled in-domain data. Data was the only independent variable: recipe, schedule and seeds
were identical.
Its measured results on the sealed benchmark, including the comparison against the matched
random control, per-boundary breakdowns and seed spreads, are reported in the paper only. The
public benchmark (an open re-rendering of the same scenarios with the same executable gold labels
and proofs) is
EnterpriseOps-KO-Blind-A.
In distribution nothing regressed: +1.37pp on the clean split and +3.12pp on the full split
against the reference.
Read the held-out boundary
On the held-out boundary (REFUSE / ESCALATE), which neither arm's added data covers, the paper
reports a change whose interval contains zero. Boundary-targeted supervision internalizes the boundaries you give
it; we have no evidence it transfers to one you do not. If your deployment has a boundary
that matters, it needs its own data or it needs test-time reasoning.
And the ceiling
The same base model allowed to reason before deciding does better than this checkpoint on the
same benchmark (see the paper). This checkpoint recovers part of that gap in a single forward
pass, not all of it. If you can afford reasoning latency on unseen policies, reason.
Zero-shot English transfer
All enterprise and boundary-targeted supervision for this checkpoint was Korean. No English
fine-tuning was used. We applied the model zero-shot to independently rendered English versions
of the same executable policy scenarios — re-rendered from the latent scenario, not translated, so
every gold label is unchanged — under a pre-registered audit.
The pre-registered English-retention gate passed, and the English difference is of the same
order as re-rendering the same scenarios a second time in Korean. Read this as no detectable
English degradation within this audit, not as evidence that Korean and English performance are
universally equivalent. Boundary-targeted training's advantage over the matched control also
carried over to English, and in both languages it disappears once reasoning is on. All numbers
are in the paper; the common subset used there is easier than the full benchmark, so they are
not comparable with full-set figures.
The audit used the single-token decision interface served through vLLM. The public English
renderings, seals, pre-registration and ledgers are in
EnterpriseOps-KO-Blind-A
(configs en, endoc, enframe). Its scenarios were authored in Korean, so English-specific
policy formulations and discourse conventions are under-represented.
Which seed this is
Seed 1 of 3, selected on the in-distribution evaluation set (n=5,608; 83.15% vs 82.47% and
80.63%). It was not selected on the sealed benchmark, though it happens also to be the best
of the three there, which we state so the rule is checkable rather than merely asserted. The
validation split (n=212) saturates at 1.000 for five of six runs and cannot rank anything.
Files
Standard transformers weights (18 shards, bf16 — the same layout as the base) plus:
pointer_head.pt — the decision head (256-dim) trained jointly with the LoRA
lora_adapter_config.json — the adapter configuration that was merged in, for provenance
Same architecture as the base
Nothing was dropped. The base Enterprise-v0.1 is multimodal, and this release keeps
model.visual (333 tensors), mtp and lm_head byte-for-byte, along with the original
config.json, index and processor configs. Only the 496 text-tower weight tensors that the
adapter targets were changed. It loads exactly like the base and accepts the same inputs.
One honest caveat: training and evaluation happened on text only. The adapter touches
layers.N.{linear_attn,self_attn,mlp} and nothing else, so the vision path is carried over
untouched and unevaluated. Every measurement reported for this checkpoint is text-only.
How it was merged
The adapter was trained with task_type=FEATURE_EXTRACTION, so its keys are
base_model.model.layers.N... — one level shallower than a CausalLM wrapper expects. Loading
it through PeftModel.from_pretrained(AutoModelForCausalLM(...)) therefore matches nothing
and merge_and_unload() returns the base model unchanged, with no error. We hit that twice.
A second trap sits right behind it: AutoModelForCausalLM loads only the text tower of this
multimodal base, so saving from it silently discards the vision weights. Our first build did
exactly that and shipped 1.8 GB lighter than the base.
So the merge happens at the tensor-file level, with no model class involved. Each shard is
opened, W += (lora_alpha / r) * (B @ A) is applied to the targeted tensors with
lora_alpha/r = 2.0, and the shard is written back under the same name. The job fails loudly
if any targeted tensor has a zero delta, and separately if the output tensor-name set is
missing anything the input had. Measured here: 496/496 tensors merged, zero with a zero delta,
maximum relative weight change 0.0041, nothing lost.
Usage
from transformers import AutoModelForImageTextToText, AutoProcessor
ID = "ThakiCloud/Qwen3.8-27B-Human-KO-Enterprise-Boundary"
m = AutoModelForImageTextToText.from_pretrained(ID, dtype="bfloat16", device_map="auto")
proc = AutoProcessor.from_pretrained(ID)
The architecture is Qwen3_5ForConditionalGeneration, identical to the base, so load it the
same way you load Enterprise-v0.1. AutoModelForCausalLM also works and gives you the text
tower alone, which is what the measurements used.
The paper's measurements use a constrained single-token decision: restrict the vocabulary
to the seven action tokens and take the argmax at one position. On in-distribution policies
that matches generated JSON in accuracy while removing almost all of the latency and roughly
92% of the run-to-run action instability across server restarts. The pointer head is an
alternative read-out and is included for reproduction; note it is worse calibrated than the
constrained read-out (ECE 0.054 vs 0.018).
Limitations
- Supervised in Korean only; evaluated in Korean and, zero-shot, in English. One task family,
one base model. English-authored enterprise policies were not evaluated.
- The sealed benchmark is constructed rather than harvested — defensible labels, limited
ecological validity.
- Counterfactual-pair gains are positive but inconclusive (seed spread is large on 30 pairs).
- The training corpus is not released; it derives from internal and licensed sources.
Citation
When Should Enterprise Policy Decisions Be Learned, and When Should They Be Reasoned?
ThakiCloud, 2026. The arXiv link will be added here once the preprint is announced.