SAVRN
Search Contact SAVRN

Dataset · Text classification

EnterpriseOps-KO-Blind-A

by ThakiCloud ThakiCloud/EnterpriseOps-KO-Blind-A

A sealed Korean benchmark of 300 enterprise policy decisions whose gold labels were computed, not written. No human and no language model ever supplied an answer.

Rows300
Configurations1
Size4.0 MB
Licensecc-by-4.0
AccessPublicly accessible
Monthly Downloads—

Dataset Card

By ThakiCloud, published under cc-by-4.0, revision 17eeeecf08b4.

A sealed Korean benchmark of 300 enterprise policy decisions whose gold labels were computed, not written. No human and no language model ever supplied an answer. A hidden executable policy maps a set of latent factors to exactly one correct action and emits a proof; the natural-language rendering is produced separately and never sees the label.

Each item asks the same question an enterprise agent has to answer before it says anything: given this written policy, these callable tools, what the system already knows, and this user turn — which one of seven actions is correct?

ANSWER · ASK_CLARIFICATION · ASK_CONFIRMATION · CALL_TOOL · CALL_MULTIPLE_TOOLS · ESCALATE · REFUSE

Why the labels are trustworthy

We tried the usual routes first and they failed, which is why this one exists.

Attempt Result
Cross-vendor adjudication (two commercial models, independent), pool A κ = 0.47 on 1,193 items
Same protocol, pool B κ = 0.40 on 853 items
Caveat on both ~30% of items abstained and neither run finished its challenger stage, so read these as "not converging" rather than as clean reliability estimates
Diagnosis On the items meant to probe the answer-vs-call-a-tool boundary, the adjudicators chose CALL_TOOL on 3 of 348 and ANSWER on 182. Layering tools and principles onto real legal text does not put the intended boundary into the text.

So the boundary is no longer written into the text and recovered by a judge. It is a property of the latent factor assignment, and the label is whatever a deterministic executor derives from it.

latent factors → executor (fixed priority order) → gold action + proof
                                ↓
            renderer (models outside the evaluated family, label hidden)
                                ↓
        machine gates → metadata-only selection → content-hash seal

Latent factors: needs_live_state, missing_required_fact (+ user_can_supply), confirmation_required (+ side_effecting), prohibited (+ escalation_makes_permissible), human_review_required, tool_count, answerable_from_static_policy.

The full factor space was enumerated; only assignments from which exactly one action is derivable were kept.

What is in here

data/policy_if.jsonl     300 sealed items  {id, axis, prompt, expected, grader, meta}
data/proofs.jsonl        per item: factors, the rule that fired, the derivation,
                         the scenario, the counterfactual partner, the renderer
data/SEAL.json           content hashes + the 12-item verification checklist
data/select_report.json  what the selection actually used (metadata only)
code/                    the generator and the gates (see below)
ledgers/                 the aggregate measurement ledgers behind every number in the paper

Layers

Layer n What it probes
normal_unseen 60 Ordinary decisions under an unseen policy
call_tool_boundary 60 Answer from policy vs. call a tool for live state
counterfactual_pair 60 30 pairs differing in one latent variable, gold flips
dialect_noisy 30 15 pairs differing only in surface dialect, gold identical
policy_vs_principle 30 A specific clause against a general principle
confirm_refuse_boundary 30 Confirmation vs. escalation vs. refusal
missing_information 30 Which missing fact the user can actually supply

The seal

items_sha256 = 090fedfed3660b343942cecf8e5cf3296237f712891e3b11093c472f4239575d

Twelve mechanical checks, all passing, re-computed from the sealed artifact rather than trusted from the generator's own records: item count, slice quota, unique ids, unique prompts, proof completeness, no unresolved reference, no answer leak in the prompt, 30 counterfactual pairs that flip, 15 invariance pairs that do not, counterfactual pairs differing in exactly one variable (verified by diffing the two scenarios directly), selection by metadata only, and zero leakage in the selected pool.

Re-verify it yourself:

python code/verify_seal.py              # 12/12 must pass
python code/verify_seal.py --self-test  # confirm the verifier catches a corrupted label

Note on code/. verify_seal.py, worlds.py, sampler.py, select300.py and gates.py run on this repository as-is. factors.py and render.py each import one internal validation helper (data_factory.validators.verify) that is part of our training pipeline and is not released; they are included as the reference definition of the executor and the rendering protocol, not as runnable scripts. The verifier reimplements every check it needs from the standard library on purpose — it must not depend on the code that produced the artifact it is checking.

Scoring

Exact action match. The paper's primary endpoints are overall exact-action accuracy, both-correct on counterfactual pairs (an item counts only if the model gets both members right), and CALL_TOOL recall. Pair-level endpoints are the informative ones: getting both members of a counterfactual pair right requires responding to the one latent variable that changed rather than to the surface of either item.

Reference points measured on this set with a single 27B Korean enterprise model:

Configuration Exact action
No thinking, constrained single token 70.7%
No thinking, LoRA pointer head 74.6%
Trained on random in-domain data 74.4%
Trained on boundary-targeted data 79.2%
Always thinking, constrained single token 88.0%

Models: ThakiCloud/Qwen3.8-27B-Human-KO-Enterprise-v0.1 and -v0.2.

Honest limitations

  • Constructed, not harvested. These are invented policies in invented institutions. That is what makes the labels defensible and what limits ecological validity: a model can be good here and bad on a real company's documents.
  • Released, therefore contaminable. We publish it because we have already used it to develop an adaptive method, which disqualifies it as a test set for our own future claims whether or not we publish. Treat it as a development set once it is in your training data.
  • The renderings were produced with general-purpose commercial models (Gemini and Claude). Check those providers' terms before using this benchmark as training data.
  • code/render.py is included for reproducibility but its vendor glue requires your own API keys, and the diagnostic endpoint is an environment variable with no default.

Not included

The training corpus is not released. It derives from internal and licensed sources we are not in a position to redistribute. This repository contains only the evaluation benchmark, its proofs, its generator, and the aggregate measurement ledgers.

Citation

The benchmark is introduced in When Should Enterprise Policy Decisions Be Learned, and When Should They Be Reasoned? (ThakiCloud, 2026). The arXiv link will be added to this card once the preprint is announced.

@misc{han2026pace,
  author       = {Han, Hyojung},
  title        = {{EnterpriseOps-KO} {Blind-A}: proof-carrying enterprise policy decisions},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/ThakiCloud/EnterpriseOps-KO-Blind-A}}
}

Structure

default 300 rows

SplitRowsSize
test300747.1 KB
idstringaxisstringpromptstringexpectedvaluegraderstringmetaJson

Details

Repository
ThakiCloud/EnterpriseOps-KO-Blind-A
Publisher
ThakiCloud
Task category
Text classification
Tags
enterprise-agents, policy-compliance, tool-use
Size category
n<1K
Languages
ko
Revision
17eeeecf08b4150682bd37fe2e6db7a15dbb5eab
Last updated
2026-10-02

Files

23 files, 4.0 MB in total.

Data14 files · 3.8 MB
Documentation1 file · 8.0 KB
Other7 files · 140.2 KB
Repository1 file · 2.5 KB
Every file
FileTypeSizeSHA-256
data/SEAL.jsonData786 B—
data/policy_if.jsonlData783.1 KB—
data/proofs.jsonlData2.8 MB—
data/select_report.jsonData3.7 KB—
ledgers/2026-10-01-pace-analysis-compare-methods.jsonData23.0 KB—
ledgers/2026-10-01-pace-analysis-restart-pairs.jsonData4.2 KB—
ledgers/2026-10-01-pace-r0-summary.jsonData11.4 KB—
ledgers/2026-10-01-pace-r2-loao.jsonData51.6 KB—
ledgers/2026-10-01-pace-r5-selective.jsonData74.6 KB—
ledgers/2026-10-02-pace-blindA-2x2.jsonData39.2 KB—
ledgers/2026-10-02-pace-blindA-adaptive-think.jsonData1.7 KB—
ledgers/2026-10-02-pace-blindA-race.jsonData26.4 KB—
ledgers/2026-10-02-pace-r3-judge.jsonData34.4 KB—
ledgers/2026-10-02-pace-think-utility.jsonData4.0 KB—
README.mdDocumentation8.0 KB—
code/factors.pyOther11.4 KB—
code/gates.pyOther46.6 KB—
code/render.pyOther18.9 KB—
code/sampler.pyOther27.0 KB—
code/select300.pyOther8.8 KB—
code/verify_seal.pyOther5.7 KB—
code/worlds.pyOther21.8 KB—
.gitattributesRepository2.5 KB—

License and Download

License
cc-by-4.0
Access
No access gate
Download from ThakiCloud

Released by ThakiCloud through its official repository on Hugging Face. Read the license.

Models Trained on This Dataset