SAVRN
Search Contact SAVRN

Open-weight model · Text classification

deberta_MP_dynamic

by Oriane Peter orpe42/deberta_MP_dynamic

deberta_MP_dynamic is an open-weight model for text classification from Oriane Peter, released under MIT License. It has 435M parameters and a 512-token context. At 16-bit it needs about 1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 146 downloads a month.

This model is a fine-tuned version of microsoft/deberta-v3-large on an unknown dataset.

Parameters435M
Context512
Weights8.7 GB
Licensemit
AccessOpen weights
Monthly Downloads146

Runs On

What it takes to serve deberta_MP_dynamic (435M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.9 GB 1.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.4 GB 0.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

deberta_MP_dynamic on every accelerator the SAVRN Index prices, at every precision

Model Card

By Oriane Peter, published under mit, revision 31d3e69f9dc9.

This model is a fine-tuned version of microsoft/deberta-v3-large on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 3e-05 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1000 - Transformers 5.12.1 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.22.2

Read Oriane Peter's full model card

This model is a fine-tuned version of microsoft/deberta-v3-large on an unknown dataset. It achieves the following results on the evaluation set: - Loss: 0.1224 - Macro F1: 0.5993 - Micro F1: 0.6698 - Macro Precision: 0.5833 - Macro Recall: 0.6198 - Micro Precision: 0.6508 - Micro Recall: 0.6898 - Exact Match Ratio: 0.0756 - Macro Roc Auc: 0.8980

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training: - learning_rate: 3e-05 - train_batch_size: 16 - eval_batch_size: 16 - seed: 42 - gradient_accumulation_steps: 2 - total_train_batch_size: 32 - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments - lr_scheduler_type: cosine - lr_scheduler_warmup_steps: 0.1 - num_epochs: 1000

Training results

Training Loss Epoch Step Validation Loss Macro F1 Micro F1 Macro Precision Macro Recall Micro Precision Micro Recall Exact Match Ratio Macro Roc Auc
0.1465 12.8312 500 0.1330 0.0747 0.2343 0.1856 0.0569 0.5726 0.1473 0.0091 0.7400
0.1375 25.6494 1000 0.1276 0.3216 0.5160 0.4262 0.3049 0.5651 0.4747 0.0287 0.8359
0.1341 38.4675 1500 0.1264 0.3569 0.5434 0.5077 0.3403 0.6199 0.4838 0.0415 0.8623
0.1311 51.2857 2000 0.1256 0.3895 0.5649 0.5364 0.3516 0.6622 0.4926 0.0550 0.8777
0.1292 64.1039 2500 0.1239 0.4202 0.5950 0.5565 0.3972 0.6184 0.5733 0.0403 0.8882
0.1282 76.9351 3000 0.1230 0.4595 0.6156 0.5226 0.4707 0.5995 0.6326 0.0375 0.8948
0.1261 89.7532 3500 0.1226 0.4813 0.6258 0.5394 0.4899 0.6158 0.6362 0.0438 0.8991
0.1244 102.5714 4000 0.1224 0.4999 0.6321 0.5574 0.4918 0.6251 0.6392 0.0513 0.8987
0.1236 115.3896 4500 0.1228 0.5070 0.6272 0.5033 0.5565 0.5710 0.6956 0.0293 0.9031
0.1226 128.2078 5000 0.1218 0.5264 0.6449 0.5336 0.5453 0.6114 0.6822 0.0427 0.9045
0.1212 141.0260 5500 0.1216 0.5281 0.6490 0.5699 0.5257 0.6485 0.6495 0.0504 0.9062
0.1199 153.8571 6000 0.1213 0.5410 0.6510 0.5528 0.5639 0.6220 0.6827 0.0507 0.9078
0.1189 166.6753 6500 0.1216 0.5464 0.6513 0.5374 0.5807 0.6130 0.6946 0.0463 0.9085
0.1183 179.4935 7000 0.1218 0.5480 0.6529 0.5846 0.5361 0.6646 0.6415 0.0559 0.9012
0.1180 192.3117 7500 0.1214 0.5578 0.6552 0.5453 0.5893 0.6191 0.6957 0.0453 0.9069
0.1170 205.1299 8000 0.1214 0.5594 0.6579 0.5841 0.5593 0.6608 0.6550 0.0594 0.9067
0.1163 217.9610 8500 0.1214 0.5698 0.6573 0.5460 0.6099 0.6075 0.7158 0.0444 0.9045
0.1151 230.7792 9000 0.1214 0.5733 0.6621 0.5715 0.5900 0.6491 0.6757 0.0559 0.9061
0.1150 243.5974 9500 0.1216 0.5749 0.6611 0.5524 0.6126 0.6266 0.6996 0.0494 0.9059
0.1147 256.4156 10000 0.1214 0.5796 0.6620 0.5694 0.6046 0.6354 0.6909 0.0572 0.9064
0.1141 269.2338 10500 0.1216 0.5753 0.6614 0.5817 0.5872 0.6524 0.6706 0.0590 0.9000
0.1137 282.0519 11000 0.1212 0.5830 0.6647 0.5830 0.5950 0.6505 0.6796 0.0628 0.9039
0.1133 294.8831 11500 0.1215 0.5792 0.6635 0.5742 0.5980 0.6497 0.6779 0.0600 0.9028
0.1130 307.7013 12000 0.1217 0.5839 0.6641 0.5602 0.6250 0.6283 0.7042 0.0545 0.9048
0.1131 320.5195 12500 0.1216 0.5812 0.6624 0.5705 0.6051 0.6436 0.6824 0.0548 0.9040
0.1121 333.3377 13000 0.1217 0.5875 0.6628 0.5576 0.6324 0.6254 0.7049 0.0516 0.9044
0.1119 346.1558 13500 0.1217 0.5869 0.6642 0.5654 0.6192 0.6437 0.6862 0.0653 0.9053
0.1116 358.9870 14000 0.1219 0.5896 0.6648 0.5710 0.6203 0.6361 0.6963 0.0568 0.9043
0.1113 371.8052 14500 0.1217 0.5900 0.6669 0.5709 0.6227 0.6419 0.6938 0.0608 0.9024
0.1109 384.6234 15000 0.1216 0.5922 0.6677 0.5722 0.6226 0.6498 0.6865 0.0607 0.9053
0.1105 397.4416 15500 0.1219 0.5900 0.6656 0.5700 0.6243 0.6400 0.6934 0.0593 0.9052
0.1103 410.2597 16000 0.1219 0.5899 0.6670 0.5743 0.6186 0.6447 0.6909 0.0616 0.9021
0.1099 423.0779 16500 0.1215 0.5927 0.6700 0.5942 0.5990 0.6696 0.6704 0.0694 0.8993
0.1098 435.9091 17000 0.1220 0.5873 0.6662 0.5714 0.6160 0.6501 0.6831 0.0662 0.9013
0.1093 448.7273 17500 0.1220 0.5902 0.6681 0.5687 0.6213 0.6435 0.6947 0.0634 0.8999
0.1094 461.5455 18000 0.1219 0.5931 0.6681 0.5787 0.6160 0.6477 0.6899 0.0650 0.8996
0.1092 474.3636 18500 0.1217 0.5928 0.6703 0.5956 0.5996 0.6690 0.6717 0.0700 0.8976
0.1088 487.1818 19000 0.1222 0.5919 0.6675 0.5822 0.6128 0.6547 0.6808 0.0648 0.8988
0.1086 500.0 19500 0.1225 0.5897 0.6652 0.5699 0.6200 0.6505 0.6806 0.0695 0.9002
0.1082 512.8312 20000 0.1219 0.5950 0.6692 0.5819 0.6169 0.6535 0.6857 0.0691 0.9007
0.1081 525.6494 20500 0.1217 0.5937 0.6695 0.5875 0.6110 0.6609 0.6783 0.0692 0.8997
0.1082 538.4675 21000 0.1224 0.5973 0.6683 0.5683 0.6364 0.6373 0.7026 0.0666 0.9016
0.1082 551.2857 21500 0.1220 0.5962 0.6691 0.5803 0.6212 0.6526 0.6865 0.0663 0.9011
0.1075 564.1039 22000 0.1222 0.5995 0.6695 0.5884 0.6186 0.6631 0.6760 0.0701 0.8978
0.1075 576.9351 22500 0.1221 0.5989 0.6705 0.5825 0.6231 0.6495 0.6928 0.0664 0.9000
0.1074 589.7532 23000 0.1219 0.5994 0.6734 0.5798 0.6260 0.6519 0.6965 0.0683 0.8998
0.1075 602.5714 23500 0.1220 0.6010 0.6721 0.5843 0.6236 0.6521 0.6934 0.0668 0.8953
0.1074 615.3896 24000 0.1222 0.5996 0.6721 0.5816 0.6256 0.6503 0.6954 0.0686 0.8985
0.1069 628.2078 24500 0.1224 0.5963 0.6696 0.5810 0.6200 0.6582 0.6814 0.0696 0.8981
0.1068 641.0260 25000 0.1224 0.5983 0.6710 0.5869 0.6172 0.6603 0.6820 0.0712 0.8958
0.1068 653.8571 25500 0.1221 0.6029 0.6723 0.5975 0.6141 0.6668 0.6778 0.0753 0.8958
0.1065 666.6753 26000 0.1221 0.5980 0.6700 0.6006 0.6033 0.6730 0.6670 0.0795 0.8947
0.1065 679.4935 26500 0.1224 0.6024 0.6728 0.5847 0.6273 0.6523 0.6947 0.0707 0.8967
0.1064 692.3117 27000 0.1223 0.6027 0.6739 0.5805 0.6318 0.6530 0.6962 0.0707 0.8977
0.1063 705.1299 27500 0.1225 0.6006 0.6704 0.5936 0.6144 0.6681 0.6726 0.0765 0.8947
0.1060 717.9610 28000 0.1226 0.6036 0.6723 0.5818 0.6310 0.6512 0.6948 0.0721 0.8969
0.1062 730.7792 28500 0.1223 0.6031 0.6724 0.5838 0.6286 0.6533 0.6926 0.0723 0.8956
0.1058 743.5974 29000 0.1223 0.6033 0.6728 0.5984 0.6122 0.6699 0.6758 0.0776 0.8939
0.1059 756.4156 29500 0.1224 0.6055 0.6738 0.5971 0.6217 0.6621 0.6859 0.0768 0.8935
0.1056 769.2338 30000 0.1224 0.6038 0.6736 0.5961 0.6165 0.6659 0.6814 0.0768 0.8936
0.1056 782.0519 30500 0.1224 0.6036 0.6731 0.5867 0.6252 0.6576 0.6892 0.0731 0.8957
0.1056 794.8831 31000 0.1225 0.6060 0.6745 0.5837 0.6345 0.6527 0.6979 0.0717 0.8978
0.1055 807.7013 31500 0.1225 0.6026 0.6737 0.5822 0.6289 0.6552 0.6934 0.0737 0.8969
0.1054 820.5195 32000 0.1225 0.6037 0.6735 0.5856 0.6278 0.6554 0.6927 0.0736 0.8964
0.1054 833.3377 32500 0.1225 0.6046 0.6747 0.5853 0.6295 0.6550 0.6957 0.0762 0.8959
0.1054 846.1558 33000 0.1226 0.6034 0.6737 0.5841 0.6295 0.6560 0.6924 0.0751 0.8962
0.1052 858.9870 33500 0.1224 0.6057 0.6740 0.5944 0.6228 0.6662 0.6819 0.0773 0.8948
0.1050 871.8052 34000 0.1226 0.6042 0.6744 0.5891 0.6253 0.6607 0.6886 0.0756 0.8954
0.1053 884.6234 34500 0.1226 0.6052 0.6752 0.5878 0.6287 0.6573 0.6942 0.0758 0.8955
0.1052 897.4416 35000 0.1226 0.6050 0.6754 0.5887 0.6264 0.6600 0.6916 0.0756 0.8953
0.1052 910.2597 35500 0.1226 0.6049 0.6745 0.5862 0.6292 0.6586 0.6912 0.0760 0.8963
0.1051 923.0779 36000 0.1225 0.6061 0.6754 0.5904 0.6268 0.6616 0.6898 0.0772 0.8954
0.1052 935.9091 36500 0.1226 0.6050 0.6751 0.5882 0.6272 0.6596 0.6914 0.0755 0.8960
0.1050 948.7273 37000 0.1226 0.6054 0.6752 0.5867 0.6301 0.6585 0.6927 0.0750 0.8960
0.1052 961.5455 37500 0.1225 0.6060 0.6753 0.5897 0.6275 0.6602 0.6912 0.0753 0.8959
0.1050 974.3636 38000 0.1225 0.6065 0.6754 0.5887 0.6296 0.6595 0.6920 0.0759 0.8961
0.1053 987.1818 38500 0.1225 0.6057 0.6752 0.5884 0.6284 0.6591 0.6921 0.0760 0.8961
0.1051 1000.0 39000 0.1225 0.6060 0.6754 0.5888 0.6286 0.6595 0.6921 0.0759 0.8961

Framework versions

  • Transformers 5.12.1
  • Pytorch 2.11.0+cu128
  • Datasets 5.0.1
  • Tokenizers 0.22.2

Configuration

Architecture
DebertaV2ForSequenceClassification
Context length (tokens)
512
Layers
24
Hidden size
1,024
Feed-forward size
4,096
Attention heads
16
Vocabulary size
128,100
Model type
deberta-v2

Identity and Version

Repository
orpe42/deberta_MP_dynamic
Publisher
Oriane Peter
Task
Text classification
Modality
Text
Library
transformers
Parameters
435M parameters
Languages
Not stated by the source
Revision
31d3e69f9dc99210bae75efdfb941a199cf8a554
First published
2026-08-06
Last updated
2026-09-30

Files and Weights

44 files, 8.7 GB in total. The weights are 9 files totalling 8.7 GB in bin, pt, pth, safetensors.

Weights9 files · 8.7 GB
Configuration4 files · 132.2 KB
Tokenizer6 files · 25.0 MB
Documentation1 file · 15.9 KB
Other23 files · 1.2 MB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
best_model/model.safetensorsWeights1.7 GB 190423dea1d6
best_model/training_args.binWeights5.3 KB a4b7760eded9
current_best/model.safetensorsWeights1.7 GB bc6248cc5ff2
current_best/optimizer.ptWeights3.5 GB 098933cebcc7
current_best/rng_state.pthWeights14.2 KB 96a3667d6833
current_best/scheduler.ptWeights1.1 KB 7eb7ccbcbe37
current_best/training_args.binWeights4.9 KB 4c5264262098
model.safetensorsWeights1.7 GB 5a80c69110e6
training_args.binWeights4.9 KB ebeedcee40d8
best_model/config.jsonConfiguration5.1 KB —
config.jsonConfiguration3.6 KB —
current_best/config.jsonConfiguration5.1 KB —
current_best/trainer_state.jsonConfiguration118.2 KB —
README.mdDocumentation15.9 KB —
best_thresholds.npyOther576 B 7b874082e7c2
runs/Aug06_12-10-47_erc-hpc-vm042/events.out.tfevents.1786014647.erc-hpc-vm042.307913.0Other63.3 KB 3c1474c7be85
runs/Aug06_12-10-47_erc-hpc-vm042/events.out.tfevents.1786031039.erc-hpc-vm042.307913.1Other811 B 65be6cb94ebc
runs/Aug06_12-13-56_erc-hpc-vm044/events.out.tfevents.1786014836.erc-hpc-vm044.79796.0Other21.9 KB d3640ea97f1d
runs/Aug06_14-27-22_erc-hpc-comp031/events.out.tfevents.1786022842.erc-hpc-comp031.2297350.0Other8.8 KB bbc23ae94833
runs/Aug06_15-42-04_erc-hpc-comp031/events.out.tfevents.1786027324.erc-hpc-comp031.2302155.0Other49.5 KB edca708fd60a
runs/Aug06_15-42-04_erc-hpc-comp031/events.out.tfevents.1786050127.erc-hpc-comp031.2302155.1Other811 B d9f33ae0fe41
runs/Aug06_16-58-04_erc-hpc-comp248/events.out.tfevents.1786031884.erc-hpc-comp248.1390536.0Other90.9 KB 020d926d2d43
runs/Aug06_16-58-04_erc-hpc-comp248/events.out.tfevents.1786073374.erc-hpc-comp248.1390536.1Other824 B ea33405baff6
runs/Aug06_23-42-40_erc-hpc-vm041/events.out.tfevents.1786056160.erc-hpc-vm041.4006523.0Other67.5 KB 39a2ec7ae027
runs/Aug07_08-40-50_erc-hpc-vm041/events.out.tfevents.1786088450.erc-hpc-vm041.4016584.0Other14.1 KB 1b7ba8284dd5
runs/Aug07_08-44-07_erc-hpc-comp248/events.out.tfevents.1786088647.erc-hpc-comp248.1584800.0Other12.3 KB b04d9f9bbc76
runs/Aug07_09-31-05_erc-hpc-vm041/events.out.tfevents.1786091465.erc-hpc-vm041.4017016.0Other31.9 KB f262cdd89e0e
runs/Aug07_09-32-07_erc-hpc-comp248/events.out.tfevents.1786091527.erc-hpc-comp248.1586555.0Other30.1 KB c6cbe1e8c47d
runs/Aug18_15-19-24_erc-hpc-comp248/events.out.tfevents.1787062764.erc-hpc-comp248.1575213.0Other145.4 KB dc621e4ea00b
runs/Aug19_14-33-28_erc-hpc-comp033/events.out.tfevents.1787146408.erc-hpc-comp033.100299.0Other149.3 KB 6209a821ec74
runs/Aug19_14-33-28_erc-hpc-comp033/events.out.tfevents.1787222161.erc-hpc-comp033.100299.1Other824 B e5929b9e8b37
runs/Aug23_16-46-45_erc-hpc-vm042/events.out.tfevents.1787500005.erc-hpc-vm042.2697604.0Other124.0 KB 523c970927f1
runs/Aug24_08-29-23_erc-hpc-comp031/events.out.tfevents.1787556563.erc-hpc-comp031.186562.0Other149.3 KB ff18ab3068d4
runs/Aug24_08-29-23_erc-hpc-comp031/events.out.tfevents.1787633956.erc-hpc-comp031.186562.1Other824 B 097ba003dba0
runs/Sep26_14-51-43_erc-hpc-comp246/events.out.tfevents.1790430703.erc-hpc-comp246.28815.0Other149.0 KB 1341975c1a80
runs/Sep26_14-51-43_erc-hpc-comp246/events.out.tfevents.1790453924.erc-hpc-comp246.28815.1Other824 B 7ab39a74dbf7
runs/Sep29_13-20-28_erc-hpc-comp039/events.out.tfevents.1790684428.erc-hpc-comp039.1459476.0Other93.1 KB 74cccc7916f7
.gitattributesRepository1.5 KB —
best_model/tokenizer.jsonTokenizer8.3 MB —
best_model/tokenizer_config.jsonTokenizer538 B —
current_best/tokenizer.jsonTokenizer8.3 MB —
current_best/tokenizer_config.jsonTokenizer538 B —
tokenizer.jsonTokenizer8.3 MB —
tokenizer_config.jsonTokenizer538 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
8.7 GB
Download from Oriane Peter

Released by Oriane Peter through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published8.7 GB
16-bit0.9 GB
8-bit0.4 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About deberta_MP_dynamic

How much GPU memory does deberta_MP_dynamic need?

About 1 GB at 16-bit and 0.3 GB at 4-bit: the weights (435M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run deberta_MP_dynamic on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use deberta_MP_dynamic commercially?

Yes. deberta_MP_dynamic is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is deberta_MP_dynamic's context length?

512 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

DEBATE-kor-large

Jong Rock Jeong

DEBATE-kor-large is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text. The model is initialized from mlburnham/PoliticalDEBATEDeBERTalargev1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI. Political DEBATE DeBERTa-large → PolNLI-kor → DEBATE-kor-large Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI. DEBATE-kor formulates NLI as a binary classification problem. notentailment combines…

Open weights 435M parameters 512 tokens transformers

Opir-multitask-large is the English, highest-accuracy multi-task checkpoint in the Opir family: an encoder-based GLiClass guardrail model for real-time LLM safety filtering. It supports binary safe/unsafe classification, toxicity detection, jailbreak and prompt-injection detection, and zero-shot harmful-content categorization over a hierarchical safety taxonomy. This card is for knowledgator/opir-multitask-large. The model is used through GLiClass zero-shot classification: pass text plus the candidate labels you want scored. Use single-label mode for binary safe/unsafe decisions and multi-label mode for taxonomy, toxicity, jailbreak, or custom policy labels. Use multi-label mode when you…

Open weights apache-2.0 439M parameters gliclass

Model · Text classification

erabi-practical-v1-experimental

Sugarknight

This is an experimental, uncalibrated choice-ranking model. It is not an official Jev model, a validated general-purpose reasoner, or an automatic decision-maker. The model ranks 2–16 user-supplied candidate texts for a natural-language context and question and returns all candidate probabilities through the ERABI code. Decisions should be reviewed by a person. - Practical V1 data consists of original synthetic Japanese, English, and Simplified Chinese examples in six task families, generated and answer-blind rejudged with DeepSeek V4.1 Flash. - The Exam-QA source was filtered and transformed with the same DeepSeek model. Symbolic answer labels were mapped to source choice text. Ambiguous…

Open weights apache-2.0 439M parameters transformers

Model · Text classification

laya-typed-decisions

Convai Innovations

Non-autoregressive System 1 decision model, fine-tuned on the typed-decisions workflows: agent-trace observability, customer service, invoice processing and security incidents. Part of the Laya family. 400 test cases, 2,000 decisions, measured on the official test split. +3.9 points over Jev's published 0.727, above the 0.735 teacher ceiling, with 2.4x better Brier and 1.6x better score MAE. Jev figures are third-party published, not measured here — there is no TypeSafe API access in this project, and sample sizes and prompts differ. Treat the comparison as indicative. Router will not select this checkpoint automatically unless you construct it with autotaskdetection=True — it is…

Open weights apache-2.0 421M parameters transformers

Model · Text classification

laya

Convai Innovations

Multilingual, non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with mathematically calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. Trained with reinforcement learning against strictly proper scoring rules (RLCD), so reporting honest probabilities is the only way to maximise reward. It never generates text, so there is nothing to parse and nothing to hallucinate. This repo holds all three checkpoints and is the hub for the family. The English checkpoint is at the repo root; the other two are bundled subfolders, and only the one you request is downloaded: pip install -U…

Open weights apache-2.0 421M parameters transformers

Model · Text classification

laya-kvp10k-noul

Sothiara Em

A fine-tuned Laya model (Convai Innovations, 421M, ModernBERT-large backbone) that answers one typed noul question: does a value correctly match its key label? (e.g. firstname = John → true, firstname = 1992 → false). Fine-tuned on the IBM KVP-10K dataset via the pre-parsed community mirror (OCR pre-extracted; no OCR step needed). This is the v2 (remediated) release: document-level 80/10/10 data partitioning, mixed-class frozen test set, calib-split temperature fitting, and a frozen category-stratified release gate. It supersedes the v1 release pre-remediation, historical artifact). Laya is a non-autoregressive decision model: you give it a state (text/dict) plus typed questions (choice…

Open weights apache-2.0 421M parameters laya