SAVRN
Search Contact SAVRN

Open-weight model · Text classification

laya-coreai

by Andrey Babikov AndyInQtr/laya-coreai

laya-coreai is an open-weight model for text classification from Andrey Babikov, released under Apache License 2.0. It has 322M parameters. At 16-bit it needs about 0.8 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Laya typed decisions on Apple Silicon, running on the Core AI runtime — the successor to Core ML. This is a.aimodel asset exported from via Apple's coreai-torch bridge.

Parameters322M
Context—
Weights643.8 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve laya-coreai (322M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.6 GB 0.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

laya-coreai on every accelerator the SAVRN Index prices, at every precision

Model Card

By Andrey Babikov, published under apache-2.0, revision 3b7ab85b23f7.

Laya typed decisions on Apple Silicon, running on the Core AI runtime — the successor to Core ML. This is a.aimodel asset exported from via Apple's coreai-torch bridge. It outputs choice / score / noul probabilities (and RL action logits) with zero generated tokens and no PyTorch, Core ML, Transformers, or cloud API at inference time. macOS 27+ (Core AI runtime), Python 3.10+. Validated on M3 Max / macOS 27.2. Validate the download end-to-end (all three specializations, timing, contract checks): Snake demo with the model (terminal game, reuses the laya-coreml UI + safety shield; automatically uses the B3 asset when present for ~2x game throughput): ~3× faster per pass than the fastest Core…

Read Andrey Babikov's full model card

laya-multilingual-coreai-f16

Laya typed decisions on Apple Silicon, running on the Core AI runtime — the successor to Core ML. This is a .aimodel asset exported from convaiinnovations/laya-multilingual via Apple's coreai-torch bridge. It outputs choice / score / noul probabilities (and RL action logits) with zero generated tokens and no PyTorch, Core ML, Transformers, or cloud API at inference time.

Run

macOS 27+ (Core AI runtime), Python 3.10+. Validated on M3 Max / macOS 27.2.

hf download AndyInQtr/laya-coreai --local-dir laya-coreai
pip install coreai-core==1.0.0b2 laya-coreml numpy   # tokenizer/prompt builders come from laya-coreml
import laya_coreai  # from the downloaded repo's laya_coreai/ package

agent = laya_coreai.load("laya-coreai", unit="gpu")  # or "ne" / "cpu"
out = agent.predict(
    "The customer asks for a refund of a duplicate payment.",
    {"refund": {"type": "noul", "instructions": "Does the customer request a refund?"}},
)
print(out["answers"])

Validate the download end-to-end (all three specializations, timing, contract checks):

python laya-coreai/scripts/validate.py laya-coreai
# PASS predict(gpu/ne/cpu) ... gpu p50≈5.0ms · ne p50≈5.0ms · cpu p50≈11.3ms
# RESULT: all checks passed

Snake demo with the model (terminal game, reuses the laya-coreml UI + safety shield; automatically uses the B3 asset when present for ~2x game throughput):

pip install rich
python laya-coreai/scripts/snake_play.py --fps 12

Measured on M3 Max / macOS 27.2

Backend Single-pass p50 Snake decision parity vs Core ML W8
this model, GPU (recommended) 4.6–5.0 ms 40/40 argmax, max KL 1.5e-4
this model, Neural Engine 4.6–5.0 ms 40/40 argmax, max KL 1.5e-4
this model, CPU-only 10.8–11.3 ms 40/40 argmax, max KL 4.1e-4
Core ML W8 (laya-coreml ane-w8, CPU+ANE) 13.5 ms baseline
Core ML FP16 flexible (CPU+GPU) 17.1 ms exact weights
MLX FP16 (laya-mlx) ~10.8 ms exact weights

~3× faster per pass than the fastest Core ML bundle at fixed shape, with better fidelity than the W8 quantization it is compared against (W8 drifts up to 0.014 calibration on the upstream fixture; this asset drifts 2.1e-4).

Always pin the compute unit

laya_coreai.load(unit="gpu") (default) or unit="ne". The unpinned default specialization routes part of this graph to the ANE compiler, which fails type inference (anec.scaled_elementwise memref mismatch) on macOS 27.2; the failed program load SIGABRTs the process. GPU is the robust choice; NE works but inherits that fragility.

Batched asset: one pass per snake decision

laya-f16-b3.aimodel (same repo) exports batch 3, letting the three snake questions (move / risk / food) share one forward pass:

Backend (decision = 3 questions) p50 per decision
this model, B3 asset, GPU 5.6–6.0 ms (175 decisions/s)
this model, B1 asset, GPU 14.3–14.6 ms (3 passes)
Core ML W8 bundle (CPU+ANE) 11.6–13.5 ms
MLX FP16 ~10.8 ms + prompt build

B3 parity: 100/100 argmax agreement vs both the B1 asset (max KL 0.00000) and the Core ML W8 bundle (max KL 0.00014). Load it with laya_coreai.load("laya-coreai", unit="gpu", asset="laya-f16-b3.aimodel").

Format and limits

  • Fixed shape: 96 tokens total, 32 option slots, per question (same capacity class as the Core ML ane-w8 bundle). B1 asset runs one question per pass; B3 asset batches three. Over-capacity prompts raise; nothing is silently truncated.
  • FP16 weights (main.mlirb, 615 MB), logits/action outputs FP32.
  • Rebuild from the pinned source with scripts/build_aimodel.py (torch export → coreai-torch → .aimodel, ~90 s, sha256-gated provenance).
  • coreai_config.json records shapes, provenance, and the validation record.

Provenance

  • Original checkpoint: convaiinnovations/laya-multilingual at 052592a15d198d9ad47da779604259b10b47b7aa.
  • Original weights SHA256: 9d628fd971b700382ac6f65920a86f149777b2e748e0c955fb3b19695aa8f204.
  • Graph: laya_coreml.torch_model.DecisionModel (Apache-2.0); upstream implementation NandhaKishorM/laya.
  • Converter: apple/coreai-torch 0.4.2 / coreai-core 1.0.0b2 (BSD-3).
  • Original model by Convai Innovations and contributors, Apache-2.0.
  • Independent conversion; not an official Convai Innovations or Apple release.

See LICENSE and NOTICE. This is an inference port, not a newly trained decision model; task/language limitations originate with Laya.


Own MLX engine (engine="mlx") — immune to the Apple runtime leak

Known issue in coreai-core 1.0.0b2 (Apple beta): its GPU/NE paths allocate every inference output from a process-global Metal ioSurface pool that never recycles blocks. After ~6,000–8,000 rapid calls the pool dries and the runtime hard-traps the process: CoreAIRuntime/NDArray+Pool.swift:77: Fatal error: Failed to allocate storage for NDArray with byteCount: 24, sk: ioSurface. At game speeds (--max-speed) that is reached in ~90 s; it is an Apple bug (fixed only upstream), not a model or wrapper bug. Workaround for paced play: --unit ne at ≤60 fps.

This repo therefore ships a from-scratch MLX engine for the same graph (laya_coreai/mlx_engine.py): ModernBERT encoder + decision head + scorer + action head, implemented directly against the original FP32 checkpoint and fused with mx.compile. No CoreAIRuntime involvement — leak-free by construction.

Verified against the laya-f16-b3.aimodel asset on the same inputs (B=3 × L=96, macOS 27.2, M3 Max):

own MLX engine Core AI asset (GPU)
3-question decision p50 6.0 ms 5.6 ms
decision parity (40 prompts) 40/40 argmax, max KL 2.0e-4 baseline
noul parity max |Δ| 0.003, 80/80 sign agreement baseline
leak soak 20,000 calls, fds flat, RSS flat hard-trap at ~6k

This repo ships the engine's checkpoint (model.safetensors, 633 MB, fp32 originals of the same weights the .aimodel carries, Apache-2.0 like the upstream convaiinnovations/laya-multilingual), so a full bundle download is all you need:

hf download AndyInQtr/laya-coreai --local-dir laya-coreai
cd laya-coreai && pip install mlx laya-coreml numpy
python scripts/snake_play.py --engine mlx --max-speed     # leak-free, ~6 ms/decision
python scripts/validate.py                                 # certifies every backend

Python API:

import laya_coreai
agent = laya_coreai.load(engine="mlx")     # finds model.safetensors in the bundle
out = agent.predict("Safe route: yes.", questions)       # one fused pass

Identity and Version

Repository
AndyInQtr/laya-coreai
Publisher
Andrey Babikov
Task
Text classification
Modality
Text
Library
coreai
Parameters
322M parameters
Languages
Not stated by the source
Revision
3b7ab85b23f758153f131922f09793badfbc5592
First published
2026-09-20
Last updated
2026-09-21

Files and Weights

24 files, 2.0 GB in total. The weights are 1 file totalling 643.8 MB in safetensors.

Weights1 file · 643.8 MB
Configuration13 files · 49.5 KB
Tokenizer2 files · 34.4 MB
Documentation3 files · 18.5 KB
Other4 files · 1.3 GB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights643.8 MB 9d628fd971b7
coreai_config.jsonConfiguration1.3 KB —
encoder/config.jsonConfiguration1.9 KB —
laya-f16-b3.aimodel/metadata.jsonConfiguration105 B —
laya-f16.aimodel/metadata.jsonConfiguration105 B —
laya_coreai/__init__.pyConfiguration10.6 KB —
laya_coreai/mlx_engine.pyConfiguration12.7 KB —
rl_agent_config.jsonConfiguration472 B —
scripts/acceptance_parity.pyConfiguration2.4 KB —
scripts/build_aimodel.pyConfiguration4.4 KB —
scripts/race_b3.pyConfiguration4.0 KB —
scripts/snake_play.pyConfiguration4.6 KB —
scripts/test_fd_leak.pyConfiguration1.4 KB —
scripts/validate.pyConfiguration5.6 KB —
LICENSEDocumentation10.2 KB —
NOTICEDocumentation1.2 KB —
README.mdDocumentation7.2 KB —
laya-f16-b3.aimodel/main.hashOther32 B —
laya-f16-b3.aimodel/main.mlirbOther644.4 MB 8486db39f429
laya-f16.aimodel/main.hashOther32 B —
laya-f16.aimodel/main.mlirbOther644.4 MB af81dbc3dbdd
.gitattributesRepository1.7 KB —
tokenizer/tokenizer.jsonTokenizer34.4 MB 609d8f4c067c
tokenizer/tokenizer_config.jsonTokenizer502 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
643.8 MB
Download from Andrey Babikov

Released by Andrey Babikov through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published643.8 MB
16-bit0.6 GB
8-bit0.3 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About laya-coreai

How much GPU memory does laya-coreai need?

About 0.8 GB at 16-bit and 0.2 GB at 4-bit: the weights (322M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run laya-coreai on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use laya-coreai commercially?

Yes. laya-coreai is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

laya-multilingual

Convai Innovations

Non-autoregressive System 1 decision model covering 100+ languages. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with probabilities in a single forward pass. No text generation, so nothing to parse and nothing to hallucinate. Part of the Laya family — use this checkpoint for anything that is not English. Since laya 0.3.13 the default Router() keeps both english and this checkpoint resident, so a mixed workload no longer swaps checkpoints on every language change. For a server, load them up front so even the first request of each language is just a forward pass: router.attach("multilingual", agent) registers an Agent you already built, so a…

Open weights apache-2.0 322M parameters transformers

Model · Text classification

laya-pt-es-typed

Telepatia

This checkpoint fine-tunes convaiinnovations/laya-multilingual for native choice, score, and noul decisions in Portuguese and Spanish. It keeps the original 322M-parameter mmBERT architecture. It adds no inference component and does not generate text. It returns typed answers and probabilities in one forward pass. This is a text model. Inference takes a textual state plus typed questions. The second training stage used text decisions derived from public speech corpora, but this checkpoint does not accept audio by itself. The separate audio projector is not included. The official Laya SDK defines these primitives as follows: - choice: selects one key from a runtime-defined criteria object.…

Open weights apache-2.0 322M parameters laya

Model · Text classification

laya-multilingual

Scott Lamkin

Non-autoregressive System 1 decision model covering 100+ languages. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with probabilities in a single forward pass. No text generation, so nothing to parse and nothing to hallucinate. Part of the Laya family — use this checkpoint for anything that is not English. The default Router() keeps both english and this checkpoint resident, so a mixed workload no longer swaps checkpoints on every language change. For a server, load them up front so even the first request of each language is just a forward pass: router.attach("multilingual", agent) registers an Agent you already built, so a process that loaded…

Open weights apache-2.0 322M parameters transformers

Classifies GitHub issues written in any language as bug, feature, question or docs. A fine-tune of Laya multilingual (mmBERT-base) used by the laya-triage GitHub Action for non-English issues, next to the English model laya-triage-en. The same 500 NLBSE'23 validation issues, machine-translated with NLLB-200 into 13 languages. Accuracy (±3 points per language): laya-triage and Jev are within noise of each other across languages; both are far ahead of the untuned base. Translations can flatter a model trained on translations, so we also checked real issues: on 367 non-English issues opened in 2026 (never seen, written by people, not translated) accuracy went from 47.1% to 65.7%. Use it…

Open weights apache-2.0 322M parameters

Model · Text classification

rex

NguyenThanhDat

System 1 calibrated decision model for zero-latency threat triage

Open weights apache-2.0 322M parameters

Model · Text classification

sentinel-laya-multilingual

Sep Lol

Fine-tune of convaiinnovations/laya-multilingual (Apache-2.0) for prompt-injection and jailbreak detection. The base model is mmBERT, vocab 256k. Code: 3p3r/sentinel-laya. The English checkpoint is a different model: 3p3r/sentinel-laya. This revision starts from the inject-label checkpoint and adds a round on 3p3r/short-role-attacks, a synthetic English and German set of short role swaps and short orders, each with a benign twin. The five evaluation sets below were not part of training. Noul temperature is 2.2534, kept from the first revision. Threshold is 0.5. Positive class: jailbreak or prompt injection. Metric: Binary F1. Same prompts as the English card, one RTX 3090, batch size 64.…

Open weights apache-2.0 322M parameters