SAVRN
Search Contact SAVRN

Open-weight model · Text classification

laya-kvp10k-noul

by Sothiara Em sothiem/laya-kvp10k-noul

laya-kvp10k-noul is an open-weight model for text classification from Sothiara Em, released under Apache License 2.0. It has 421M parameters. At 16-bit it needs about 1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

A fine-tuned Laya model (Convai Innovations, 421M, ModernBERT-large backbone) that answers one typed noul question: does a value correctly match its key label? (e.g. firstname = John → true, firstname = 1992 → false).

Parameters421M
Context—
Weights842.6 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve laya-kvp10k-noul (421M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.8 GB 1.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.4 GB 0.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

laya-kvp10k-noul on every accelerator the SAVRN Index prices, at every precision

Model Card

By Sothiara Em, published under apache-2.0, revision 2ea229cfb3ec.

A fine-tuned Laya model (Convai Innovations, 421M, ModernBERT-large backbone) that answers one typed noul question: does a value correctly match its key label? (e.g. firstname = John → true, firstname = 1992 → false). Fine-tuned on the IBM KVP-10K dataset via the pre-parsed community mirror (OCR pre-extracted; no OCR step needed). This is the v2 (remediated) release: document-level 80/10/10 data partitioning, mixed-class frozen test set, calib-split temperature fitting, and a frozen category-stratified release gate. It supersedes the v1 release pre-remediation, historical artifact). Laya is a non-autoregressive decision model: you give it a state (text/dict) plus typed questions (choice…

Read Sothiara Em's full model card

Laya KVP-10K — Key/Value Match Detector (noul), v2 (remediated release)

A fine-tuned Laya model (Convai Innovations, 421M, ModernBERT-large backbone) that answers one typed noul question: does a value correctly match its key label? (e.g. first_name = John → true, first_name = 1992 → false).

Fine-tuned on the IBM KVP-10K dataset via the pre-parsed community mirror alessandrorusso21/KVP10k (OCR pre-extracted; no OCR step needed).

Scope note. This is an experimental fine-tune on KVP-10K data for one very specific purpose: judging whether a key label matches a paired value inside a form-like document excerpt (English, 512-token budget). It is not a general-purpose key-value extraction or document-understanding model, and the reported numbers are claims about that narrow task only.

This is the v2 (remediated) release: document-level 80/10/10 data partitioning, mixed-class frozen test set, calib-split temperature fitting, and a frozen category-stratified release gate. It supersedes the v1 release (sothiem/laya-kvp10k-noul, pre-remediation, historical artifact).

This is not a generative LLM

Laya is a non-autoregressive decision model: you give it a state (text/dict) plus typed questions (choice, noul, score) and it returns calibrated probabilities in a single forward pass. It never generates text. Do not use chat templates, generate(), or causal-LM losses with this checkpoint. It is loaded with laya.Agent(<dir>), not transformers.

Model details

Base model convaiinnovations/laya @ aa8c91c (421M English checkpoint)
Encoder answerdotai/ModernBERT-large
Fine-tuning recipe Official RLCD (REINFORCE with strictly proper scoring rules + GRPO-style group-mean baseline, global_std advantage normalization) — not LoRA/peft
laya package version 0.3.11 (torch 2.14.0+cu130, Python 3.13.7)
Question type noul (key_value_match) — options render in fixed [false, true] order; p[1] is P(true)
Context budget max_len=512, head_max_len=192 (≈320 tokens for the state)
Temperature fitted 0.3899 on the held-out calib split (accepted: calib NLL 0.1897 → 0.0812). Note: the stock laya.Agent only accepts temperatures in [0.5, 5] and clamps this to 0.5 at load — see Limitations
Parameters 421.3M
Precision float32 master weights, bf16 AMP inference

Question definition

type: noul
id: key_value_match
instructions: "The value correctly matches the key label."
criteria:
  "false": "the value does not match the key"
  "true": "the value is consistent with the key"

State format

state = {
    "key": "first_name",
    "value": "1992",
    "document_excerpt": "<token-budgeted window around the value span>",
}

Usage

from laya import Agent

agent = Agent("<this-repo-dir-or-hf-id>")

state = {
    "key": "first_name",
    "value": "1992",
    "document_excerpt": "Name: John Smith  DOB: 1992-04-01  ...",
}
question = {
    "type": "noul",
    "id": "key_value_match",
    "instructions": "The value correctly matches the key label.",
    "criteria": {
        "false": "the value does not match the key",
        "true": "the value is consistent with the key",
    },
}
p = agent.ask(state, [question])[0]   # [P(false), P(true)]
print(p[1])                           # P(match)

Noul options render in fixed [false, true] order and must never be shuffled — p[1] is P(true).

Performance

Headline result: the frozen mixed-class test set (test_eval, n = 10,538, built only from the HF test documents, evaluated exactly once at release; the fitted temperature applied without refitting):

model accuracy recall_1 (match) recall_0 (mismatch) bias gap NLL Brier ECE lat p50/p95 ms
laya-base 0.7472 0.7478 0.7466 0.0011 0.5320 0.1772 0.0665 22.4 / 52.1
this model, raw RLCD (best) 0.9852 0.9761 0.9943 0.0182 0.3465 0.0880 0.2698 23.5 / 47.9
this model, temperature-refit (this checkpoint) 0.9852 0.9761 0.9943 0.0182 0.0743 0.0141 0.0233 23.7 / 32.3

Per source (this checkpoint):

Source n Accuracy Brier
gold (positives) 5,269 97.61% 0.0217
cross-document negatives 3,057 99.51% 0.0067
same-format negatives 1,165 99.06% 0.0091
cross-type negatives 1,047 99.62% 0.0027

Notes:

  • +23.8 pp accuracy over the base checkpoint (0.747 → 0.985).
  • The post-training noul temperature refit (T = 0.3899, fit on the document-disjoint calib split) is what buys the calibration: identical accuracy to the raw checkpoint, NLL −78%, Brier −84%, ECE −91%.
  • Mismatch recall (0.9943) slightly exceeds match recall (0.9761) — gap 0.0182, below the 0.10 warn threshold. This is the expected cost of the strict 50/50 class balance via seeded synthetic negatives; match recall is the residual weakness.
  • Training curve (4 epochs, val accuracy 97.64% → 98.27%; epoch 3 selected by composite score):

Category release gate (frozen probe, 7 categories × 40)

category n accuracy match recall mismatch recall
exact_match 40 1.000 1.000 —
same_fmt_wrong 40 0.975 — 0.975
reformat 40 0.725 0.725 —
value_absent 40 1.000 — 1.000
cross_field 40 1.000 — 1.000
alias_key 40 1.000 1.000 —
uppercase 40 1.000 1.000 —

Predefined thresholds (set before the run): accuracy ≥ 0.85, match recall ≥ 0.85, mismatch recall ≥ 0.75, min n 20. The gate FAILED on reformat (reformulated values such as case/spacing/$/comma variants of the gold value): 0.725 vs the 0.85 threshold. The other six categories pass. This is a known regression inherited from the base checkpoint (which scores reformulated values below verbatim ones); it is disclosed here rather than hidden by the strong aggregate.

Training data

  • Partitioning: documents of the HF train split are partitioned before any example generator runs (seed 42): 80% train (5,486 docs, 177,282 examples) / 10% val (686 docs, 16,992) / 10% calib (686 docs, 18,090). Document sets are pairwise disjoint; every negative generator runs per partition with its donor pool restricted to that partition.
  • Positives: gold (key, value) pairs from KVP-10K gts/<hash>.json (kvps_list).
  • Negatives (synthetic, seeded, 50/50 class balance, never a gold pair): per positive the negative budget is split over same-format digit perturbations (50%), cross-type typed values (25%), legacy same-doc / cross-doc swaps (~25%); plus reformulation-equivalent positives (0.36 per gold), a no-excerpt prior slice (~10%, train-only, base-model-agreement filtered), and seeded Faker invoice documents (~10%, train-only).
  • Document excerpts: 240-token window around the value span, built from the pre-parsed OCR words in ocrs/<hash>.json.
  • Held out: the HF test/ split never appears in train/val/calib and was used only for the one-shot frozen test_eval (10,538 examples, ~50/50 classes, gold-guarded) and never for hyperparameter selection. The temperature was fit on calib only; val was used only for checkpoint selection and gates.

Training procedure

Single NVIDIA RTX 5090 (32 GB), bf16 AMP, 4 epochs:

Parameter Value
Effective batch size 32 (micro 16 × grad-accum 2)
Group size 4 (REINFORCE noisy-logit samples)
LR (encoder / head) 2.5e-5 / 1.0e-4, cosine → 1e-6
Exploration noise σ 0.4 → 0.1 (linear anneal)
Baseline group-mean, advantage normalized by one global std
Reward log-score 1.0 + spherical 0.75 (strictly proper)
CE co-training weight 1.0 (soft-teacher mix 0.4 with base-model probabilities, clipped to [0.05, 0.95])
Grad clip 1.0
Checkpoint selection composite score on val (accuracy penalized for miscalibration and two-point mass); epoch 3 of 4 selected
Temperature refit noul bucket only, fit on calib, accepted iff calib NLL not worsened beyond 1e-6 (fitted 0.3899 accepted)

Limitations

  • reformat category fails the release gate (0.725 vs 0.85): expect reduced match recall when the value appears in a reformulated form (case/spacing/$/comma/spelled-out variants) rather than verbatim.
  • Temperature clamp: the fitted noul temperature (0.3899) is stored in rl_agent_config.json, but the stock laya.Agent only accepts temperatures in [0.5, 5] and clamps it to 0.5 at load (with a RuntimeWarning). The reported test metrics are measured at T = 0.5 and are already well calibrated (ECE 0.023); applying the fitted T would require rescaling outside the Agent.
  • Match recall (0.9761) is the weaker class; the model is slightly "mismatch-leaning" (bias gap 0.0182).
  • English documents only (512-token budget). For non-English use convaiinnovations/laya-multilingual (1024-token budget) and re-fine-tune.
  • The model judges a key/value pair in the context of a document excerpt; accuracy degrades if the excerpt does not cover the value span.
  • Probabilities are calibrated on the KVP-10K-style distribution (form-like documents); out-of-domain calibration is not guaranteed.

Repository contents

File Description
model.safetensors Fine-tuned weights (encoder + decision head), 804 MB
rl_agent_config.json Laya agent config (budgets, fitted noul temperature 0.3899, question metadata)
encoder/config.json ModernBERT-large encoder config
tokenizer/ Tokenizer files (tokenizer.json, tokenizer_config.json)
checkpoint_meta.json Checkpoint provenance (selected epoch, calib diagnostics, temperature-refit decision, manifest digests)
progress.png Training curves
laya_final_test_eval.json One-shot frozen test report (n = 10,538, with provenance block)
laya_final_calib.json Calibration-fit diagnostics on the calib split (not final performance)
category_gate.json Frozen category release-gate report and pass/fail decision
benchmark_comparison_test_eval.md base / raw / refit comparison table
run_log.jsonl Run manifest (config record: seed, versions, manifest digest) + per-step/epoch log

Citations & credits

  • Laya: Convai Innovations — https://huggingface.co/convaiinnovations/laya (Apache-2.0)
  • KVP-10K: IBM Research — https://research.ibm.com/blog/kvp10k-dataset
  • KVP-10K pre-parsed community mirror: https://huggingface.co/datasets/alessandrorusso21/KVP10k
  • Encoder: answerdotai/ModernBERT-large
  • Training/evaluation code: this repository's src/ (RLCD loop, evaluation, release gate)

License

Apache-2.0 (inherits the base model's license).

Identity and Version

Repository
sothiem/laya-kvp10k-noul
Publisher
Sothiara Em
Task
Text classification
Modality
Text
Library
laya
Parameters
421M parameters
Languages
en
Revision
2ea229cfb3ecf056371e2a7c8d656db43393d920
First published
2026-09-24
Last updated
2026-09-27

Files and Weights

14 files, 846.7 MB in total. The weights are 1 file totalling 842.6 MB in safetensors.

Weights1 file · 842.6 MB
Configuration6 files · 10.4 KB
Tokenizer2 files · 3.6 MB
Documentation2 files · 12.7 KB
Other2 files · 479.4 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights842.6 MB 6a64862d42e9
category_gate.jsonConfiguration1.8 KB —
checkpoint_meta.jsonConfiguration2.8 KB —
encoder/config.jsonConfiguration2.2 KB —
laya_final_calib.jsonConfiguration1.5 KB —
laya_final_test_eval.jsonConfiguration1.3 KB —
rl_agent_config.jsonConfiguration892 B —
README.mdDocumentation11.8 KB —
benchmark_comparison_test_eval.mdDocumentation916 B —
progress.pngOther132.1 KB 6b5798266318
run_log.jsonlOther347.3 KB —
.gitattributesRepository1.6 KB —
tokenizer/tokenizer.jsonTokenizer3.6 MB —
tokenizer/tokenizer_config.jsonTokenizer351 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
842.6 MB
Download from Sothiara Em

Released by Sothiara Em through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published842.6 MB
16-bit0.8 GB
8-bit0.4 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About laya-kvp10k-noul

How much GPU memory does laya-kvp10k-noul need?

About 1 GB at 16-bit and 0.3 GB at 4-bit: the weights (421M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run laya-kvp10k-noul on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use laya-kvp10k-noul commercially?

Yes. laya-kvp10k-noul is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

laya-typed-decisions

Convai Innovations

Non-autoregressive System 1 decision model, fine-tuned on the typed-decisions workflows: agent-trace observability, customer service, invoice processing and security incidents. Part of the Laya family. 400 test cases, 2,000 decisions, measured on the official test split. +3.9 points over Jev's published 0.727, above the 0.735 teacher ceiling, with 2.4x better Brier and 1.6x better score MAE. Jev figures are third-party published, not measured here — there is no TypeSafe API access in this project, and sample sizes and prompts differ. Treat the comparison as indicative. Router will not select this checkpoint automatically unless you construct it with autotaskdetection=True — it is…

Open weights apache-2.0 421M parameters transformers

Model · Text classification

laya

Convai Innovations

Multilingual, non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with mathematically calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. Trained with reinforcement learning against strictly proper scoring rules (RLCD), so reporting honest probabilities is the only way to maximise reward. It never generates text, so there is nothing to parse and nothing to hallucinate. This repo holds all three checkpoints and is the hub for the family. The English checkpoint is at the repo root; the other two are bundled subfolders, and only the one you request is downloaded: pip install -U…

Open weights apache-2.0 421M parameters transformers

Model · Text classification

afm-de

Ariacompute

Laya-layout ModernBERT + MASK option head, RLCD fine-tune from convaiinnovations/laya. Checkpoint files: model.safetensors, rlagentconfig.json, tokenizer/, encoder/. Load with AFM-D: python -m afmd.de.eval --checkpoint --device cuda. See AFM-D / product docs. License follows the base Laya / ModernBERT stack.

Open weights 421M parameters transformers

Model · Text classification

openjevx

Muthukumaran Navaneethakrishnan

OpenJevX is an open-weight, non-autoregressive System One decision model specialized from Laya, which uses answerdotai/ModernBERT-large plus a dynamic typed-decision head. It accepts runtime-defined choice, score, and noul questions and returns calibrated probabilities in one forward pass. It is compatible with the TypeSafe Jev /v1/systemone request shape through the OpenJevX server. - CUDA p50 latency per five-question case: 22.8 ms The benchmark uses the untouched 400-case, 2,000-decision test split from LocalLLaMA/typed-decisions. Training uses only its 1,200-case train split. OpenJevX is a derivative of Laya by ConvAI Innovations and ModernBERT by Answer.AI and LightOn. Laya and…

Open weights apache-2.0 421M parameters laya

Model · Text classification

DEBATE-kor-large

Jong Rock Jeong

DEBATE-kor-large is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text. The model is initialized from mlburnham/PoliticalDEBATEDeBERTalargev1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI. Political DEBATE DeBERTa-large → PolNLI-kor → DEBATE-kor-large Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI. DEBATE-kor formulates NLI as a binary classification problem. notentailment combines…

Open weights 435M parameters 512 tokens transformers

Model · Text classification

deberta_MP_dynamic

Oriane Peter

This model is a fine-tuned version of microsoft/deberta-v3-large on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 3e-05 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1000 - Transformers 5.12.1 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.22.2

Open weights mit 435M parameters 512 tokens transformers