SAVRN
Search Contact SAVRN

Open-weight model · Text classification

afm-de

by Ariacompute ariacompute/afm-de

afm-de is an open-weight model for text classification from Ariacompute. It has 421M parameters. At 16-bit it needs about 1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Laya-layout ModernBERT + MASK option head, RLCD fine-tune from convaiinnovations/laya. Checkpoint files: model.safetensors, rlagentconfig.json, tokenizer/, encoder/. Load with AFM-D: python -m afmd.de.eval --checkpoint --device cuda. See AFM-D / product docs.

Parameters421M
Context—
Weights842.6 MB
License—
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve afm-de (421M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.8 GB 1.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.4 GB 0.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

afm-de on every accelerator the SAVRN Index prices, at every precision

Model Card

Laya-layout ModernBERT + MASK option head, RLCD fine-tune from convaiinnovations/laya. Checkpoint files: model.safetensors, rlagentconfig.json, tokenizer/, encoder/. Load with AFM-D: python -m afmd.de.eval --checkpoint --device cuda. See AFM-D / product docs. License follows the base Laya / ModernBERT stack.

Excerpt from the card by Ariacompute.

Identity and Version

Repository
ariacompute/afm-de
Publisher
Ariacompute
Task
Text classification
Modality
Text
Library
transformers
Parameters
421M parameters
Languages
Not stated by the source
Revision
6e24310c82bbb2da2e860db1da7ec470b4c5ebc0
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

7 files, 846.2 MB in total. The weights are 1 file totalling 842.6 MB in safetensors.

Weights1 file · 842.6 MB
Configuration2 files · 4.6 KB
Tokenizer2 files · 3.6 MB
Documentation1 file · 563 B
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights842.6 MB 24c547bb5497
encoder/config.jsonConfiguration2.1 KB —
rl_agent_config.jsonConfiguration2.5 KB —
README.mdDocumentation563 B —
.gitattributesRepository1.5 KB —
tokenizer/tokenizer.jsonTokenizer3.6 MB —
tokenizer/tokenizer_config.jsonTokenizer338 B —

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
842.6 MB
Download from Ariacompute

Released by Ariacompute through its official repository on Hugging Face.

Memory Requirements

PrecisionWeights in memory
As published842.6 MB
16-bit0.8 GB
8-bit0.4 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About afm-de

How much GPU memory does afm-de need?

About 1 GB at 16-bit and 0.3 GB at 4-bit: the weights (421M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run afm-de on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Similar Models

Model · Text classification

laya-typed-decisions

Convai Innovations

Non-autoregressive System 1 decision model, fine-tuned on the typed-decisions workflows: agent-trace observability, customer service, invoice processing and security incidents. Part of the Laya family. 400 test cases, 2,000 decisions, measured on the official test split. +3.9 points over Jev's published 0.727, above the 0.735 teacher ceiling, with 2.4x better Brier and 1.6x better score MAE. Jev figures are third-party published, not measured here — there is no TypeSafe API access in this project, and sample sizes and prompts differ. Treat the comparison as indicative. Router will not select this checkpoint automatically unless you construct it with autotaskdetection=True — it is…

Open weights apache-2.0 421M parameters transformers

Model · Text classification

laya

Convai Innovations

Multilingual, non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with mathematically calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. Trained with reinforcement learning against strictly proper scoring rules (RLCD), so reporting honest probabilities is the only way to maximise reward. It never generates text, so there is nothing to parse and nothing to hallucinate. This repo holds all three checkpoints and is the hub for the family. The English checkpoint is at the repo root; the other two are bundled subfolders, and only the one you request is downloaded: pip install -U…

Open weights apache-2.0 421M parameters transformers

Model · Text classification

laya-kvp10k-noul

Sothiara Em

A fine-tuned Laya model (Convai Innovations, 421M, ModernBERT-large backbone) that answers one typed noul question: does a value correctly match its key label? (e.g. firstname = John → true, firstname = 1992 → false). Fine-tuned on the IBM KVP-10K dataset via the pre-parsed community mirror (OCR pre-extracted; no OCR step needed). This is the v2 (remediated) release: document-level 80/10/10 data partitioning, mixed-class frozen test set, calib-split temperature fitting, and a frozen category-stratified release gate. It supersedes the v1 release pre-remediation, historical artifact). Laya is a non-autoregressive decision model: you give it a state (text/dict) plus typed questions (choice…

Open weights apache-2.0 421M parameters laya

Model · Text classification

openjevx

Muthukumaran Navaneethakrishnan

OpenJevX is an open-weight, non-autoregressive System One decision model specialized from Laya, which uses answerdotai/ModernBERT-large plus a dynamic typed-decision head. It accepts runtime-defined choice, score, and noul questions and returns calibrated probabilities in one forward pass. It is compatible with the TypeSafe Jev /v1/systemone request shape through the OpenJevX server. - CUDA p50 latency per five-question case: 22.8 ms The benchmark uses the untouched 400-case, 2,000-decision test split from LocalLLaMA/typed-decisions. Training uses only its 1,200-case train split. OpenJevX is a derivative of Laya by ConvAI Innovations and ModernBERT by Answer.AI and LightOn. Laya and…

Open weights apache-2.0 421M parameters laya

Model · Text classification

DEBATE-kor-large

Jong Rock Jeong

DEBATE-kor-large is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text. The model is initialized from mlburnham/PoliticalDEBATEDeBERTalargev1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI. Political DEBATE DeBERTa-large → PolNLI-kor → DEBATE-kor-large Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI. DEBATE-kor formulates NLI as a binary classification problem. notentailment combines…

Open weights 435M parameters 512 tokens transformers

Model · Text classification

deberta_MP_dynamic

Oriane Peter

This model is a fine-tuned version of microsoft/deberta-v3-large on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 3e-05 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1000 - Transformers 5.12.1 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.22.2

Open weights mit 435M parameters 512 tokens transformers