Laya KVP-10K — Key/Value Match Detector (noul), v2 (remediated release)
A fine-tuned Laya model
(Convai Innovations, 421M, ModernBERT-large backbone) that answers one typed
noul question: does a value correctly match its key label?
(e.g. first_name = John → true, first_name = 1992 → false).
Fine-tuned on the IBM KVP-10K
dataset via the pre-parsed community mirror
alessandrorusso21/KVP10k
(OCR pre-extracted; no OCR step needed).
Scope note. This is an experimental fine-tune on KVP-10K data for one
very specific purpose: judging whether a key label matches a paired value
inside a form-like document excerpt (English, 512-token budget). It is not
a general-purpose key-value extraction or document-understanding model, and
the reported numbers are claims about that narrow task only.
This is the v2 (remediated) release: document-level 80/10/10 data
partitioning, mixed-class frozen test set, calib-split temperature fitting,
and a frozen category-stratified release gate. It supersedes the v1 release
(sothiem/laya-kvp10k-noul,
pre-remediation, historical artifact).
This is not a generative LLM
Laya is a non-autoregressive decision model: you give it a state
(text/dict) plus typed questions (choice, noul, score) and it returns
calibrated probabilities in a single forward pass. It never generates
text. Do not use chat templates, generate(), or causal-LM losses with this
checkpoint. It is loaded with laya.Agent(<dir>), not transformers.
Model details
|
|
| Base model |
convaiinnovations/laya @ aa8c91c (421M English checkpoint) |
| Encoder |
answerdotai/ModernBERT-large |
| Fine-tuning recipe |
Official RLCD (REINFORCE with strictly proper scoring rules + GRPO-style group-mean baseline, global_std advantage normalization) — not LoRA/peft |
laya package version |
0.3.11 (torch 2.14.0+cu130, Python 3.13.7) |
| Question type |
noul (key_value_match) — options render in fixed [false, true] order; p[1] is P(true) |
| Context budget |
max_len=512, head_max_len=192 (≈320 tokens for the state) |
| Temperature |
fitted 0.3899 on the held-out calib split (accepted: calib NLL 0.1897 → 0.0812). Note: the stock laya.Agent only accepts temperatures in [0.5, 5] and clamps this to 0.5 at load — see Limitations |
| Parameters |
421.3M |
| Precision |
float32 master weights, bf16 AMP inference |
Question definition
type: noul
id: key_value_match
instructions: "The value correctly matches the key label."
criteria:
"false": "the value does not match the key"
"true": "the value is consistent with the key"
State format
state = {
"key": "first_name",
"value": "1992",
"document_excerpt": "<token-budgeted window around the value span>",
}
Usage
from laya import Agent
agent = Agent("<this-repo-dir-or-hf-id>")
state = {
"key": "first_name",
"value": "1992",
"document_excerpt": "Name: John Smith DOB: 1992-04-01 ...",
}
question = {
"type": "noul",
"id": "key_value_match",
"instructions": "The value correctly matches the key label.",
"criteria": {
"false": "the value does not match the key",
"true": "the value is consistent with the key",
},
}
p = agent.ask(state, [question])[0] # [P(false), P(true)]
print(p[1]) # P(match)
Noul options render in fixed [false, true] order and must never be
shuffled — p[1] is P(true).
Performance
Headline result: the frozen mixed-class test set (test_eval,
n = 10,538, built only from the HF test documents, evaluated exactly
once at release; the fitted temperature applied without refitting):
| model |
accuracy |
recall_1 (match) |
recall_0 (mismatch) |
bias gap |
NLL |
Brier |
ECE |
lat p50/p95 ms |
| laya-base |
0.7472 |
0.7478 |
0.7466 |
0.0011 |
0.5320 |
0.1772 |
0.0665 |
22.4 / 52.1 |
this model, raw RLCD (best) |
0.9852 |
0.9761 |
0.9943 |
0.0182 |
0.3465 |
0.0880 |
0.2698 |
23.5 / 47.9 |
| this model, temperature-refit (this checkpoint) |
0.9852 |
0.9761 |
0.9943 |
0.0182 |
0.0743 |
0.0141 |
0.0233 |
23.7 / 32.3 |
Per source (this checkpoint):
| Source |
n |
Accuracy |
Brier |
| gold (positives) |
5,269 |
97.61% |
0.0217 |
| cross-document negatives |
3,057 |
99.51% |
0.0067 |
| same-format negatives |
1,165 |
99.06% |
0.0091 |
| cross-type negatives |
1,047 |
99.62% |
0.0027 |
Notes:
- +23.8 pp accuracy over the base checkpoint (0.747 → 0.985).
- The post-training
noul temperature refit (T = 0.3899, fit on the
document-disjoint calib split) is what buys the calibration: identical
accuracy to the raw checkpoint, NLL −78%, Brier −84%, ECE −91%.
- Mismatch recall (0.9943) slightly exceeds match recall (0.9761) — gap
0.0182, below the 0.10 warn threshold. This is the expected cost of the
strict 50/50 class balance via seeded synthetic negatives; match recall is
the residual weakness.
- Training curve (4 epochs, val accuracy 97.64% → 98.27%; epoch 3 selected by
composite score):
Category release gate (frozen probe, 7 categories × 40)
| category |
n |
accuracy |
match recall |
mismatch recall |
| exact_match |
40 |
1.000 |
1.000 |
— |
| same_fmt_wrong |
40 |
0.975 |
— |
0.975 |
| reformat |
40 |
0.725 |
0.725 |
— |
| value_absent |
40 |
1.000 |
— |
1.000 |
| cross_field |
40 |
1.000 |
— |
1.000 |
| alias_key |
40 |
1.000 |
1.000 |
— |
| uppercase |
40 |
1.000 |
1.000 |
— |
Predefined thresholds (set before the run): accuracy ≥ 0.85, match recall ≥
0.85, mismatch recall ≥ 0.75, min n 20. The gate FAILED on reformat
(reformulated values such as case/spacing/$/comma variants of the gold
value): 0.725 vs the 0.85 threshold. The other six categories pass. This is a
known regression inherited from the base checkpoint (which scores
reformulated values below verbatim ones); it is disclosed here rather than
hidden by the strong aggregate.
Training data
- Partitioning: documents of the HF
train split are partitioned
before any example generator runs (seed 42): 80% train (5,486 docs,
177,282 examples) / 10% val (686 docs, 16,992) / 10% calib (686 docs,
18,090). Document sets are pairwise disjoint; every negative generator runs
per partition with its donor pool restricted to that partition.
- Positives: gold (key, value) pairs from KVP-10K
gts/<hash>.json
(kvps_list).
- Negatives (synthetic, seeded, 50/50 class balance, never a gold pair):
per positive the negative budget is split over same-format digit
perturbations (50%), cross-type typed values (25%), legacy same-doc /
cross-doc swaps (~25%); plus reformulation-equivalent positives (0.36 per
gold), a no-excerpt prior slice (~10%, train-only, base-model-agreement
filtered), and seeded Faker invoice documents (~10%, train-only).
- Document excerpts: 240-token window around the value span, built from
the pre-parsed OCR words in
ocrs/<hash>.json.
- Held out: the HF
test/ split never appears in train/val/calib and was
used only for the one-shot frozen test_eval (10,538 examples, ~50/50
classes, gold-guarded) and never for hyperparameter selection. The
temperature was fit on calib only; val was used only for checkpoint
selection and gates.
Training procedure
Single NVIDIA RTX 5090 (32 GB), bf16 AMP, 4 epochs:
| Parameter |
Value |
| Effective batch size |
32 (micro 16 × grad-accum 2) |
| Group size |
4 (REINFORCE noisy-logit samples) |
| LR (encoder / head) |
2.5e-5 / 1.0e-4, cosine → 1e-6 |
| Exploration noise σ |
0.4 → 0.1 (linear anneal) |
| Baseline |
group-mean, advantage normalized by one global std |
| Reward |
log-score 1.0 + spherical 0.75 (strictly proper) |
| CE co-training weight |
1.0 (soft-teacher mix 0.4 with base-model probabilities, clipped to [0.05, 0.95]) |
| Grad clip |
1.0 |
| Checkpoint selection |
composite score on val (accuracy penalized for miscalibration and two-point mass); epoch 3 of 4 selected |
| Temperature refit |
noul bucket only, fit on calib, accepted iff calib NLL not worsened beyond 1e-6 (fitted 0.3899 accepted) |
Limitations
reformat category fails the release gate (0.725 vs 0.85): expect
reduced match recall when the value appears in a reformulated form
(case/spacing/$/comma/spelled-out variants) rather than verbatim.
- Temperature clamp: the fitted noul temperature (0.3899) is stored in
rl_agent_config.json, but the stock laya.Agent only accepts
temperatures in [0.5, 5] and clamps it to 0.5 at load (with a
RuntimeWarning). The reported test metrics are measured at T = 0.5 and are
already well calibrated (ECE 0.023); applying the fitted T would require
rescaling outside the Agent.
- Match recall (0.9761) is the weaker class; the model is slightly
"mismatch-leaning" (bias gap 0.0182).
- English documents only (512-token budget). For non-English use
convaiinnovations/laya-multilingual (1024-token budget) and re-fine-tune.
- The model judges a key/value pair in the context of a document excerpt;
accuracy degrades if the excerpt does not cover the value span.
- Probabilities are calibrated on the KVP-10K-style distribution (form-like
documents); out-of-domain calibration is not guaranteed.
Repository contents
| File |
Description |
model.safetensors |
Fine-tuned weights (encoder + decision head), 804 MB |
rl_agent_config.json |
Laya agent config (budgets, fitted noul temperature 0.3899, question metadata) |
encoder/config.json |
ModernBERT-large encoder config |
tokenizer/ |
Tokenizer files (tokenizer.json, tokenizer_config.json) |
checkpoint_meta.json |
Checkpoint provenance (selected epoch, calib diagnostics, temperature-refit decision, manifest digests) |
progress.png |
Training curves |
laya_final_test_eval.json |
One-shot frozen test report (n = 10,538, with provenance block) |
laya_final_calib.json |
Calibration-fit diagnostics on the calib split (not final performance) |
category_gate.json |
Frozen category release-gate report and pass/fail decision |
benchmark_comparison_test_eval.md |
base / raw / refit comparison table |
run_log.jsonl |
Run manifest (config record: seed, versions, manifest digest) + per-step/epoch log |
Citations & credits
- Laya: Convai Innovations — https://huggingface.co/convaiinnovations/laya
(Apache-2.0)
- KVP-10K: IBM Research — https://research.ibm.com/blog/kvp10k-dataset
- KVP-10K pre-parsed community mirror:
https://huggingface.co/datasets/alessandrorusso21/KVP10k
- Encoder: answerdotai/ModernBERT-large
- Training/evaluation code: this repository's
src/ (RLCD loop, evaluation,
release gate)
License
Apache-2.0 (inherits the base model's license).