Dataset · Text classification
EnterpriseOps-KO-Blind-A
by ThakiCloud ThakiCloud/EnterpriseOps-KO-Blind-A
A sealed Korean benchmark of 300 enterprise policy decisions whose gold labels were computed, not written. No human and no language model ever supplied an answer.
Dataset Card
By ThakiCloud, published under cc-by-4.0, revision 17eeeecf08b4.
A sealed Korean benchmark of 300 enterprise policy decisions whose gold labels were computed, not written. No human and no language model ever supplied an answer. A hidden executable policy maps a set of latent factors to exactly one correct action and emits a proof; the natural-language rendering is produced separately and never sees the label.
Each item asks the same question an enterprise agent has to answer before it says anything: given this written policy, these callable tools, what the system already knows, and this user turn — which one of seven actions is correct?
ANSWER · ASK_CLARIFICATION · ASK_CONFIRMATION · CALL_TOOL ·
CALL_MULTIPLE_TOOLS · ESCALATE · REFUSE
Why the labels are trustworthy
We tried the usual routes first and they failed, which is why this one exists.
| Attempt | Result |
|---|---|
| Cross-vendor adjudication (two commercial models, independent), pool A | κ = 0.47 on 1,193 items |
| Same protocol, pool B | κ = 0.40 on 853 items |
| Caveat on both | ~30% of items abstained and neither run finished its challenger stage, so read these as "not converging" rather than as clean reliability estimates |
| Diagnosis | On the items meant to probe the answer-vs-call-a-tool boundary, the adjudicators chose CALL_TOOL on 3 of 348 and ANSWER on 182. Layering tools and principles onto real legal text does not put the intended boundary into the text. |
So the boundary is no longer written into the text and recovered by a judge. It is a property of the latent factor assignment, and the label is whatever a deterministic executor derives from it.
latent factors → executor (fixed priority order) → gold action + proof
↓
renderer (models outside the evaluated family, label hidden)
↓
machine gates → metadata-only selection → content-hash seal
Latent factors: needs_live_state, missing_required_fact (+ user_can_supply),
confirmation_required (+ side_effecting), prohibited (+
escalation_makes_permissible), human_review_required, tool_count,
answerable_from_static_policy.
The full factor space was enumerated; only assignments from which exactly one action is derivable were kept.
What is in here
data/policy_if.jsonl 300 sealed items {id, axis, prompt, expected, grader, meta}
data/proofs.jsonl per item: factors, the rule that fired, the derivation,
the scenario, the counterfactual partner, the renderer
data/SEAL.json content hashes + the 12-item verification checklist
data/select_report.json what the selection actually used (metadata only)
code/ the generator and the gates (see below)
ledgers/ the aggregate measurement ledgers behind every number in the paper
Layers
| Layer | n | What it probes |
|---|---|---|
normal_unseen |
60 | Ordinary decisions under an unseen policy |
call_tool_boundary |
60 | Answer from policy vs. call a tool for live state |
counterfactual_pair |
60 | 30 pairs differing in one latent variable, gold flips |
dialect_noisy |
30 | 15 pairs differing only in surface dialect, gold identical |
policy_vs_principle |
30 | A specific clause against a general principle |
confirm_refuse_boundary |
30 | Confirmation vs. escalation vs. refusal |
missing_information |
30 | Which missing fact the user can actually supply |
The seal
items_sha256 = 090fedfed3660b343942cecf8e5cf3296237f712891e3b11093c472f4239575d
Twelve mechanical checks, all passing, re-computed from the sealed artifact rather than trusted from the generator's own records: item count, slice quota, unique ids, unique prompts, proof completeness, no unresolved reference, no answer leak in the prompt, 30 counterfactual pairs that flip, 15 invariance pairs that do not, counterfactual pairs differing in exactly one variable (verified by diffing the two scenarios directly), selection by metadata only, and zero leakage in the selected pool.
Re-verify it yourself:
python code/verify_seal.py # 12/12 must pass
python code/verify_seal.py --self-test # confirm the verifier catches a corrupted label
Note on
code/.verify_seal.py,worlds.py,sampler.py,select300.pyandgates.pyrun on this repository as-is.factors.pyandrender.pyeach import one internal validation helper (data_factory.validators.verify) that is part of our training pipeline and is not released; they are included as the reference definition of the executor and the rendering protocol, not as runnable scripts. The verifier reimplements every check it needs from the standard library on purpose — it must not depend on the code that produced the artifact it is checking.
Scoring
Exact action match. The paper's primary endpoints are overall exact-action accuracy,
both-correct on counterfactual pairs (an item counts only if the model gets both members
right), and CALL_TOOL recall. Pair-level endpoints are the informative ones: getting both
members of a counterfactual pair right requires responding to the one latent variable that
changed rather than to the surface of either item.
Reference points measured on this set with a single 27B Korean enterprise model:
| Configuration | Exact action |
|---|---|
| No thinking, constrained single token | 70.7% |
| No thinking, LoRA pointer head | 74.6% |
| Trained on random in-domain data | 74.4% |
| Trained on boundary-targeted data | 79.2% |
| Always thinking, constrained single token | 88.0% |
Models: ThakiCloud/Qwen3.8-27B-Human-KO-Enterprise-v0.1
and -v0.2.
Honest limitations
- Constructed, not harvested. These are invented policies in invented institutions. That is what makes the labels defensible and what limits ecological validity: a model can be good here and bad on a real company's documents.
- Released, therefore contaminable. We publish it because we have already used it to develop an adaptive method, which disqualifies it as a test set for our own future claims whether or not we publish. Treat it as a development set once it is in your training data.
- The renderings were produced with general-purpose commercial models (Gemini and Claude). Check those providers' terms before using this benchmark as training data.
code/render.pyis included for reproducibility but its vendor glue requires your own API keys, and the diagnostic endpoint is an environment variable with no default.
Not included
The training corpus is not released. It derives from internal and licensed sources we are not in a position to redistribute. This repository contains only the evaluation benchmark, its proofs, its generator, and the aggregate measurement ledgers.
Citation
The benchmark is introduced in When Should Enterprise Policy Decisions Be Learned, and When Should They Be Reasoned? (ThakiCloud, 2026). The arXiv link will be added to this card once the preprint is announced.
@misc{han2026pace,
author = {Han, Hyojung},
title = {{EnterpriseOps-KO} {Blind-A}: proof-carrying enterprise policy decisions},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/ThakiCloud/EnterpriseOps-KO-Blind-A}}
}
Structure
default 300 rows
| Split | Rows | Size |
|---|---|---|
| test | 300 | 747.1 KB |
Details
- Repository
- ThakiCloud/EnterpriseOps-KO-Blind-A
- Publisher
- ThakiCloud
- Task category
- Text classification
- Tags
- enterprise-agents, policy-compliance, tool-use
- Size category
- n<1K
- Languages
- ko
- Revision
- 17eeeecf08b4150682bd37fe2e6db7a15dbb5eab
- Last updated
- 2026-10-02
Files
23 files, 4.0 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| data/SEAL.json | Data | 786 B | — |
| data/policy_if.jsonl | Data | 783.1 KB | — |
| data/proofs.jsonl | Data | 2.8 MB | — |
| data/select_report.json | Data | 3.7 KB | — |
| ledgers/2026-10-01-pace-analysis-compare-methods.json | Data | 23.0 KB | — |
| ledgers/2026-10-01-pace-analysis-restart-pairs.json | Data | 4.2 KB | — |
| ledgers/2026-10-01-pace-r0-summary.json | Data | 11.4 KB | — |
| ledgers/2026-10-01-pace-r2-loao.json | Data | 51.6 KB | — |
| ledgers/2026-10-01-pace-r5-selective.json | Data | 74.6 KB | — |
| ledgers/2026-10-02-pace-blindA-2x2.json | Data | 39.2 KB | — |
| ledgers/2026-10-02-pace-blindA-adaptive-think.json | Data | 1.7 KB | — |
| ledgers/2026-10-02-pace-blindA-race.json | Data | 26.4 KB | — |
| ledgers/2026-10-02-pace-r3-judge.json | Data | 34.4 KB | — |
| ledgers/2026-10-02-pace-think-utility.json | Data | 4.0 KB | — |
| README.md | Documentation | 8.0 KB | — |
| code/factors.py | Other | 11.4 KB | — |
| code/gates.py | Other | 46.6 KB | — |
| code/render.py | Other | 18.9 KB | — |
| code/sampler.py | Other | 27.0 KB | — |
| code/select300.py | Other | 8.8 KB | — |
| code/verify_seal.py | Other | 5.7 KB | — |
| code/worlds.py | Other | 21.8 KB | — |
| .gitattributes | Repository | 2.5 KB | — |
License and Download
- License
- cc-by-4.0
- Access
- No access gate
Released by ThakiCloud through its official repository on Hugging Face. Read the license.
Models Trained on This Dataset
- Trained on (disclosed)Qwen3.8-27B-Human-KO-Enterprise-Boundary