SAVRN
Search Contact SAVRN

Open-weight model · Text classification

laya

by Convai Innovations convaiinnovations/laya

laya is an open-weight model for text classification from Convai Innovations, released under Apache License 2.0. It has 421M parameters. At 16-bit it needs about 1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Multilingual, non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with mathematically calibrated probabilities in a single forward pass (~33 ms) across 100+ languages.

Parameters421M
Context—
Weights2.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve laya (421M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.8 GB 1.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.4 GB 0.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

laya on every accelerator the SAVRN Index prices, at every precision

Model Card

By Convai Innovations, published under apache-2.0, revision fe2b7719c095.

Multilingual, non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with mathematically calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. Trained with reinforcement learning against strictly proper scoring rules (RLCD), so reporting honest probabilities is the only way to maximise reward. It never generates text, so there is nothing to parse and nothing to hallucinate. This repo holds all three checkpoints and is the hub for the family. The English checkpoint is at the repo root; the other two are bundled subfolders, and only the one you request is downloaded: pip install -U…

Read Convai Innovations's full model card

Multilingual, non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with mathematically calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. Trained with reinforcement learning against strictly proper scoring rules (RLCD), so reporting honest probabilities is the only way to maximise reward. It never generates text, so there is nothing to parse and nothing to hallucinate.

This repo holds all three checkpoints and is the hub for the family. The English checkpoint is at the repo root; the other two are bundled subfolders, and only the one you request is downloaded:

Checkpoint Backbone Encoder Params Context Best at
convaiinnovations/laya (this repo root) ModernBERT-large 421M 512 English text, guardrails, email triage
convaiinnovations/laya-multilingual mmBERT-base 322M 1024 (up to 8k) 100+ languages, ~2.2x faster
convaiinnovations/laya-typed-decisions ModernBERT-large 421M 1024 the four typed-decisions workflows (0.766 acc)

What's new in laya 0.3.13

pip install -U laya for all of this; everything below is new since 0.3.6. The checkpoints themselves are unchanged.

  • About 10x faster loading. Checkpoints are built without the throwaway random weight initialisation, so laya.load() drops from about 22 s to about 2 s on CPU, with bit-identical answers. This also skips the pass that crashed on Windows with Python 3.14.
  • import laya no longer loads torch. Routing, language detection and e-mail cleaning work in lightweight processes.
  • Batch scoring. agent.predict_batch(states, questions) scores many states in shared forward passes, with answers identical to calling predict one state at a time.
  • Routed batches. Router.predict_batch(requests) routes each request, groups them by checkpoint and question set, and scores each group in shared forward passes, with answers identical to one predict call per request, including requests whose options arrive in a different order.
  • Prediction hooks. Opt-in hooks run around every decision, to audit, trace, redact, cache or gate results. With no hooks set, answers are unchanged.
  • Schema-driven decisions. agent.decide(state, schema=...) takes a JSON schema or pydantic model and returns typed values in one forward pass.
  • Per-language calibration. lang_temperatures= applies your own per-language temperatures, and every answer also reports answer_confidence (max(p)) next to the unchanged confidence.
  • Sturdier fast path. On a CUDA out-of-memory error the fast path switches off before falling back to CPU, and single-option choice questions work.
  • transformers 4.x and Apple GPUs. Checkpoints load with the right RoPE settings on transformers 4.x, and predict() no longer crashes on MPS builds without an autocast backend.
  • Faster paths, all opt-in. laya.load(..., fast=True) uses a TileLang GPU fast path that matches the stock bf16 forward within rounding. Agent(compile=True) enables torch.compile, and laya.onnx_agent.ONNXAgent runs an exported model on ONNX Runtime.
  • Run it your way, locally. A self-hosted Jev-compatible HTTP server (pip install "laya[serve]", then laya-serve, see Self-hosting), a laya command for quick local tests, an optional MCP server (pip install "laya[mcp]"), LangChain and LangGraph integrations (pip install "laya[langchain]"), and laya-ts, a TypeScript package for Node and the browser that gives the same answers as the Python package.
  • Better routing. Plain-ASCII Spanish, Italian, Portuguese, French and German, Brazilian Portuguese support text, CJK text with Latin brand names, romanized Bangla and Azerbaijani now reach the multilingual checkpoint. Checked on 20,000 English texts: at most 5 English sentences move, all quoting long native-script names.
  • Router defaults and hooks. Router() keeps two checkpoints resident, you can pass your own language guess with lang_guess=, and Router/Agent work as context managers.
  • Correctness and clearer errors. Long conversation lists keep the newest turn when truncated, non-ASCII instructions reach the model as text, and a malformed question is rejected with a message naming the question and what to fix. A noul criteria dict keyed anything other than true/false is rejected rather than silently replaced; use labels to change the wording.

Quickstart: Route Mode (Recommended)

Laya's built-in Router is the recommended way to use Laya in production. It evaluates any state in any language, automatically detects scripts and languages in sub-milliseconds, and dispatches to the optimal checkpoint in a single forward pass.

pip install laya
import laya
from laya import Router

# Preload checkpoints into memory for instant sub-35ms routing
router = Router(preload=True)

state = {
    "from": "[email protected]",
    "subject": "Duplicate charge on invoice #4411",
    "body": "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
}

questions = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this request?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages, system errors",
            "sales": "pricing, new contracts",
            "other": "everything else"
        }
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this request?",
        "criteria": ["not urgent", "soon", "critical deadline or blocking issue"]
    },
    "churn_risk": {
        "type": "noul",
        "instructions": "Does the user threaten to cancel or leave?"
    },
    "refund_requested": {
        "type": "noul",
        "instructions": "Does the user explicitly request a refund?"
    }
}

# 1. English state -> automatically routed to ModernBERT-large (39.5 ms)
res_en = router.predict(state, questions)
print("Department :", res_en["answers"]["department"]["choice"])  # -> billing (confidence: 0.94)
print("Routing    :", res_en["routing"]["model"])                 # -> english

# 2. Hindi state -> automatically routed to mmBERT-base (100+ languages, 32.8 ms)
res_hi = router.predict({"body": "मुझसे दो बार शुल्क लिया गया, कृपया पैसे वापस करें।"}, questions)
print("Department :", res_hi["answers"]["department"]["choice"])  # -> billing (confidence: 0.86)
print("Routing    :", res_hi["routing"]["model"])                 # -> multilingual

# 3. Explicit override when you already know the checkpoint
res_td = router.predict(state, questions, model="typed-decisions")

Every result carries full routing metadata explaining why the choice was made:

res_hi["routing"]
# {
#   'model': 'multilingual',
#   'repo': 'convaiinnovations/laya/multilingual',
#   'reason': 'non-Latin script (devanagari, 100% of letters); the English checkpoint cannot read it'
# }

Why Route: The Evidence

On a shared benchmark (17,416 questions, one T4 GPU, identical questions per model):

Benchmark / Task English (laya) Multilingual (laya-multilingual) Router (Routed)
MASSIVE intent, English 0.783 0.657 0.783
MASSIVE intent, 13 other languages 0.306 0.451 0.451
XNLI, English 0.860 0.843 0.860
XNLI, 14 other languages 0.521 0.731 0.731
Languages usable (>3x random) 23 / 51 45 / 51 45 / 51
Latency, 1 question (T4 GPU) 39.5 ms 32.8 ms 32.8 ms
Latency, 10 questions batched 158.6 ms 72.3 ms 72.3 ms

The English checkpoint collapses on non-Latin scripts (Khmer scores 0.000 accuracy at 0.952 confidence). Because the model stays confident while being wrong, confidence gating cannot save you. Router detects the script in <0.5 ms pure Python before the forward pass.

Supplying your own language detection

If you already run a language-identification model, pass its answer instead of relying on the built-in heuristic. lang_guess takes a language code or a callable, is checked after an explicit lang= and before detection, and a callable that returns None falls through to detection:

router.predict(state, questions, lang_guess="ro")              # a code you already know
router = Router(preload=True, lang_guess=my_lid)               # or install one for every request

Production Preload & Memory

A cold checkpoint build costs seconds; language detection costs microseconds. Since laya 0.3.13 the lazy default keeps two checkpoints resident (english and multilingual, the only two automatic routing chooses between), so after each language's first load a switch costs detection only. A single-language deployment never builds the second. max_loaded=1 rebuilds on every switch (measured at a 7.4 s median reload on CPU and 10.3 s on T4).

For a server or a demo, preload:

# Every checkpoint resident in memory; language flips cost detection only (<1 ms)
router = Router(preload=True)
router = Router(preload=True, device="cuda")

# Or preload only the specific checkpoints you serve:
router.preload(["english", "multilingual"])

# If your app already built an agent, attach it to avoid duplicate VRAM:
router.attach("english", existing_agent)

# Manage resident memory (default keeps two hot: english + multilingual, LRU eviction)
router = Router(max_loaded=3)       # all three hot, e.g. with auto_task_detection
router = Router(max_loaded=1)       # memory-constrained host, reloads on every switch
router.unload()                     # free memory

with Router() as r:                 # releases the models when the block ends
    r.predict(state, questions)
Deployment Mode Per-Request Latency Model Reloads
Router() (lazy, max_loaded=2) detection only (<1 ms) after each language's first load 1 the first time a language appears
Router(max_loaded=1) 7 to 10 s on every language switch 1 per switch
Router(preload=True) 32.8 ms (GPU) / 193–464 ms (CPU) none

Single-Model Mode (Direct SDK)

If you only need a single checkpoint for a dedicated pipeline:

import laya

# 1. Load from the repo root or subfolders (downloads only the requested weights)
agent = laya.load("convaiinnovations/laya")                           # English root (~808 MB)
agent_ml = laya.load("convaiinnovations/laya", subfolder="multilingual") # 100+ languages (~647 MB)
agent_td = laya.load("convaiinnovations/laya", subfolder="typed-decisions")

# 2. Run all questions in ONE single forward pass (~35 ms on GPU)
result = agent.predict(state, questions)
answers = result["answers"]

print("Department :", answers["department"]["choice"])   # -> billing (confidence: 0.94)
print("Urgency    :", answers["urgency"]["score"])        # -> 1.84 / 2.0
print("Churn Risk :", answers["churn_risk"]["noul"])       # -> 0.892 (89.2% probability)

If laya.load() hangs: transformers probes for TensorFlow at import, and when TF is installed its abseil runtime can deadlock model construction. Run with USE_TF=0.


Self-hosting: Jev-compatible HTTP server

laya-serve exposes the Router on the same POST /v1/systemone request and response shape as TypeSafe Jev, so existing TypeSafe clients work by changing their base URL:

pip install "laya[serve]"
LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-serve        # 0.0.0.0:8000, preloads the checkpoints
curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"document": "I was charged twice. Please fix this ASAP."},
  "questions": {"billing": {"type": "noul", "instructions": "Is this ticket about billing?"}}
}'

It accepts every question shape the Jev API does (for example criteria as a list), ignores unknown fields, and returns a 422 naming the problem for a malformed question. It binds 0.0.0.0 with no authentication unless LAYA_API_KEY is set, in which case it requires Authorization: Bearer <key>.


Architecture

  • Backbone: ModernBERT-large (395M, bidirectional, fully fine-tuned) + a decision head trained from scratch: 2 transformer layers, an option-marker scorer, and an act/escalate head. 421M total. (Multilingual uses mmBERT-base, 22 layers, 256k vocab, 322M total).
  • Option markers: Every option is scored at its own [MASK] token, then softmaxed over that question's options. The answer space is defined at request time, so new schemas need no retraining.
  • Budget: 512 tokens per question for English (head_max_len = 192); 1024 tokens for multilingual (head_max_len = 256).
  • Batching: Every question in a call is answered in one single forward pass.

Training

RLCD (Reinforcement Learning for Calibrated Decisions). The policy reports a distribution; exploration adds zero-mean Gaussian noise to the logits; the reward is a strictly proper scoring rule (log + spherical, plus ranked probability score for ordinal questions). Expected reward is maximised only by reporting honest probabilities. Updates are REINFORCE with a group-mean baseline (GRPO-style). Multi-turn conversations use TD(λ=1.0) over prefix slices.


Benchmarks

Measured on a Tesla T4; every checkpoint answered byte-identical questions in the same run.

Speed

questions per call laya laya-multilingual
1 39.5 ms 32.8 ms
5 84.5 ms 40.1 ms
10 158.6 ms (15.9 ms/q) 72.3 ms (7.2 ms/q)
50 771 ms 337 ms (6.8 ms/q)

103–332 questions/sec batched on a single T4. For reference, TypeSafe Jev has been independently measured at 236–276 ms p50 (AbdelStark, nibzard), so Laya answers a single question roughly 6–8× faster.

Laya (with routing) vs TypeSafe Jev

Every Laya figure is what Router().predict(...) returns — the checkpoint the router selects for that input. Jev figures are third-party published, never measured here (no TypeSafe API access); sample sizes and prompts differ.

Benchmark / Metric TypeSafe Jev 1.13.0 Laya (routed) Comparison
typed-decisions, 2,000 decisions 0.727 0.766 +0.039 (beats 0.735 teacher ceiling)
AG News, 4 labels 0.910 0.950 +0.040
DAIR Emotion, 6 labels 0.480 0.595 +0.115
Banking77 (72 vs 77 labels) 0.870 0.425 Jev leads on >20 options
ECE (lower better) 0.246 0.081 3× better (post-temperature)
p50 latency, 1 question 236–276 ms 32.8 ms 7.8× faster
Languages usable (>3x random) no published benchmark 45 of 51 Global language coverage
Weights closed API Apache 2.0 Open weights, on-premise capable
Cost $0.042 / 1M tokens $0 self-hosted 100% free

On DAIR Emotion, Jev assigned zero probability to the true label on 16% of examples.

Where Jev leads
  • High-cardinality label spaces (>20 options at default settings): On Banking77, Jev scores 0.870 (on 72 labels) while Laya scores 0.425 (on 77 labels at default 256-token head budget). Options share a fixed head_max_len budget (192 tokens on English, 256 on multilingual), so 77 options receive only ~3 to 4 tokens per label, causing text to become indistinguishable. Jev supports up to 255 options out-of-the-box. While laya-multilingual supports 1,024 context (and up to 8,192 in the encoder) and you can raise agent.cfg["head_max_len"] = 512 at runtime, Jev is currently better suited for 50+ options in a single prompt without tuning.
  • Soft distribution matching: On typed-decisions, while Laya achieves higher argmax accuracy (0.766 vs 0.727), Jev achieves higher soft accuracy (0.580 vs 0.471) against the teacher's full probability distributions.
  • Out-of-the-box raw calibration: Before temperature scaling, the base checkpoint has higher raw ECE (0.213 vs 0.144). Laya achieves its 0.081 ECE after domain temperature fitting.

Full report: BENCHMARKS.md.

typed-decisions, measured on all three checkpoints

400 cases, 2,000 decisions, four workflows — measured here.

model accuracy soft acc Brier ECE score MAE
laya-typed-decisions 0.766 0.471 0.062 0.213 0.242
laya 0.362 0.332 0.316 0.175 0.694
laya-multilingual 0.342 0.326 0.439 0.285 0.687
Jev 1.13.0 (published) 0.727 0.580 0.148 0.144 0.391
teacher self-agreement ceiling 0.735
per-question majority class 0.461

The fine-tuned checkpoint clears the teacher ceiling and wins all four workflows: invoice processing 0.804, security incidents 0.766, customer service 0.764, agent-trace observability 0.730. By primitive: noul 0.857, choice 0.733, score 0.723.

The base checkpoints sit below the majority-class baseline here — the capability on this benchmark comes from fine-tuning, which is what the fine-tuning notebook is for.


Honest Limits

  • Base checkpoints are near chance on typed-decisions zero-shot — 0.362 here and 0.352 for multilingual, against a 0.318 random and 0.461 majority-class baseline. The 0.766 belongs to the checkpoint fine-tuned on that benchmark's own training split. Laya is a fast base to specialise, not a zero-shot decision engine.
  • High-cardinality choice questions and token budgets: Sequences split into an option prompt budget (head_max_len) and the remaining document/state budget (max_len - head_max_len):
  • laya (English) defaults to 512 context (head_max_len = 192, ~320 tokens for state).
  • laya-multilingual and laya-typed-decisions default to 1,024 context (head_max_len = 256, ~768 tokens for state; mmBERT-base encoder supports up to 8,192 with RoPE). At default settings, a 77-option question like Banking77 allocates only (256 - 16) // 77 ≈ 3–4 tokens per label, causing accuracy to fall off sharply (0.425 vs Jev's 0.870). If evaluating 50+ options in a single question: 1. Raise agent.cfg["head_max_len"] = 512 and agent.cfg["max_len"] = 1024 (or up to 2048 / 4096 / 8192) so every option has enough tokens to remain distinct. 2. Or split large option sets into a two-step coarse-to-fine hierarchical choice.
  • Ordinal score questions are the weakest primitive (SST-5 0.372).
  • noul can follow its option labels instead of the state, most strongly on this English checkpoint. noul renders its two options as false: / true:, and here that label pair can dominate the answer, returning a confident "no" for clearly positive input (#156). Check noul answers on your own data. If they look stuck, ask the same question as a two-option choice with neutral keys and your yes/no wording as the descriptions:

python {"type": "choice", "instructions": "Is this review positive?", "criteria": {"A": "yes, the review is positive", "B": "no, the review is negative"}} - action.act_probability carries no usable signal yet (#185). It reads 1.0 for almost every input, and its raw logits run against correctness (AUROC 0.30 on 396 labelled decisions). Gate on confidence instead, which reaches an AUROC of 0.77 on the same items. - Ships over-confident: Refitting one temperature per (question type, option count) moves mean ECE 0.466 → 0.081 (laya) and 0.314 → 0.106 (laya-multilingual). Do this on your own data before trusting the probabilities. - English only on root: Use laya-multilingual for anything outside English.


Links

  • GitHub: https://github.com/NandhaKishorM/laya
  • PyPI: https://pypi.org/project/laya/
  • Live Demo: https://huggingface.co/spaces/convaiinnovations/laya-demo
  • Write-up: Read on Dev.to

Apache 2.0 · Convai Innovations

Identity and Version

Repository
convaiinnovations/laya
Publisher
Convai Innovations
Task
Text classification
Modality
Text
Library
transformers
Parameters
421M parameters
Languages
Not stated by the source
Revision
fe2b7719c095b82cdb2bfeedefc1fee30e506c21
First published
2026-09-18
Last updated
2026-09-24

Files and Weights

38 files, 2.4 GB in total. The weights are 3 files totalling 2.3 GB in safetensors.

Weights3 files · 2.3 GB
Configuration10 files · 39.0 KB
Tokenizer6 files · 41.5 MB
Documentation2 files · 23.0 KB
Other16 files · 1.9 MB
Repository1 file · 1.9 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights842.6 MB 891102d37268
multilingual/model.safetensorsWeights643.8 MB 9d628fd971b7
typed-decisions/model.safetensorsWeights842.6 MB 4fa56de72383
email_utils.pyConfiguration3.9 KB —
encoder/config.jsonConfiguration2.1 KB —
eval/results.jsonConfiguration3.0 KB —
multilingual/encoder/config.jsonConfiguration1.9 KB —
multilingual/rl_agent_config.jsonConfiguration472 B —
rl_agent_api.pyConfiguration4.7 KB —
rl_agent_config.jsonConfiguration745 B —
rl_common.pyConfiguration19.1 KB —
typed-decisions/encoder/config.jsonConfiguration2.1 KB —
typed-decisions/rl_agent_config.jsonConfiguration847 B —
README.mdDocumentation21.4 KB —
eval/results.mdDocumentation1.6 KB —
assets/laya_benchmark.pngOther260.9 KB e01e49f0d842
assets/laya_benchmark_common.pngOther278.2 KB 183b0b17e8d9
assets/laya_vs_jev.pngOther216.4 KB 5c06517ea7f3
assets/laya_vs_jev_full.pngOther487.0 KB ee47b751d524
assets/logo-lockup-dark.pngOther27.7 KB —
assets/logo-lockup-dark.svgOther1.6 KB —
assets/logo-lockup.pngOther26.8 KB —
assets/logo-lockup.svgOther1.6 KB —
assets/logo-mark-ink.pngOther17.4 KB —
assets/logo-mark-ink.svgOther1.3 KB —
assets/logo-mark-mono.svgOther1.4 KB —
assets/logo-mark.pngOther18.0 KB —
assets/logo-mark.svgOther1.3 KB —
eval/benchmark_comparison.pngOther467.5 KB 78e4b926bfca
eval/reliability_eval_in.pngOther65.4 KB —
eval/reliability_eval_zs.pngOther72.0 KB —
.gitattributesRepository1.9 KB —
multilingual/tokenizer/tokenizer.jsonTokenizer34.4 MB 609d8f4c067c
multilingual/tokenizer/tokenizer_config.jsonTokenizer524 B —
tokenizer/tokenizer.jsonTokenizer3.6 MB —
tokenizer/tokenizer_config.jsonTokenizer308 B —
typed-decisions/tokenizer/tokenizer.jsonTokenizer3.6 MB —
typed-decisions/tokenizer/tokenizer_config.jsonTokenizer337 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
2.3 GB
Download from Convai Innovations

Released by Convai Innovations through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published2.3 GB
16-bit0.8 GB
8-bit0.4 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About laya

How much GPU memory does laya need?

About 1 GB at 16-bit and 0.3 GB at 4-bit: the weights (421M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run laya on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use laya commercially?

Yes. laya is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

laya-typed-decisions

Convai Innovations

Non-autoregressive System 1 decision model, fine-tuned on the typed-decisions workflows: agent-trace observability, customer service, invoice processing and security incidents. Part of the Laya family. 400 test cases, 2,000 decisions, measured on the official test split. +3.9 points over Jev's published 0.727, above the 0.735 teacher ceiling, with 2.4x better Brier and 1.6x better score MAE. Jev figures are third-party published, not measured here — there is no TypeSafe API access in this project, and sample sizes and prompts differ. Treat the comparison as indicative. Router will not select this checkpoint automatically unless you construct it with autotaskdetection=True — it is…

Open weights apache-2.0 421M parameters transformers

Model · Text classification

laya-kvp10k-noul

Sothiara Em

A fine-tuned Laya model (Convai Innovations, 421M, ModernBERT-large backbone) that answers one typed noul question: does a value correctly match its key label? (e.g. firstname = John → true, firstname = 1992 → false). Fine-tuned on the IBM KVP-10K dataset via the pre-parsed community mirror (OCR pre-extracted; no OCR step needed). This is the v2 (remediated) release: document-level 80/10/10 data partitioning, mixed-class frozen test set, calib-split temperature fitting, and a frozen category-stratified release gate. It supersedes the v1 release pre-remediation, historical artifact). Laya is a non-autoregressive decision model: you give it a state (text/dict) plus typed questions (choice…

Open weights apache-2.0 421M parameters laya

Model · Text classification

afm-de

Ariacompute

Laya-layout ModernBERT + MASK option head, RLCD fine-tune from convaiinnovations/laya. Checkpoint files: model.safetensors, rlagentconfig.json, tokenizer/, encoder/. Load with AFM-D: python -m afmd.de.eval --checkpoint --device cuda. See AFM-D / product docs. License follows the base Laya / ModernBERT stack.

Open weights 421M parameters transformers

Model · Text classification

openjevx

Muthukumaran Navaneethakrishnan

OpenJevX is an open-weight, non-autoregressive System One decision model specialized from Laya, which uses answerdotai/ModernBERT-large plus a dynamic typed-decision head. It accepts runtime-defined choice, score, and noul questions and returns calibrated probabilities in one forward pass. It is compatible with the TypeSafe Jev /v1/systemone request shape through the OpenJevX server. - CUDA p50 latency per five-question case: 22.8 ms The benchmark uses the untouched 400-case, 2,000-decision test split from LocalLLaMA/typed-decisions. Training uses only its 1,200-case train split. OpenJevX is a derivative of Laya by ConvAI Innovations and ModernBERT by Answer.AI and LightOn. Laya and…

Open weights apache-2.0 421M parameters laya

Model · Text classification

DEBATE-kor-large

Jong Rock Jeong

DEBATE-kor-large is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text. The model is initialized from mlburnham/PoliticalDEBATEDeBERTalargev1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI. Political DEBATE DeBERTa-large → PolNLI-kor → DEBATE-kor-large Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI. DEBATE-kor formulates NLI as a binary classification problem. notentailment combines…

Open weights 435M parameters 512 tokens transformers

Model · Text classification

deberta_MP_dynamic

Oriane Peter

This model is a fine-tuned version of microsoft/deberta-v3-large on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 3e-05 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1000 - Transformers 5.12.1 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.22.2

Open weights mit 435M parameters 512 tokens transformers