laya-multilingual-coreai-f16
Laya typed decisions on Apple Silicon, running on the Core AI runtime — the
successor to Core ML. This is a .aimodel asset exported from
convaiinnovations/laya-multilingual
via Apple's coreai-torch bridge. It
outputs choice / score / noul probabilities (and RL action logits) with
zero generated tokens and no PyTorch, Core ML, Transformers, or cloud API
at inference time.
Run
macOS 27+ (Core AI runtime), Python 3.10+. Validated on M3 Max / macOS 27.2.
hf download AndyInQtr/laya-coreai --local-dir laya-coreai
pip install coreai-core==1.0.0b2 laya-coreml numpy # tokenizer/prompt builders come from laya-coreml
import laya_coreai # from the downloaded repo's laya_coreai/ package
agent = laya_coreai.load("laya-coreai", unit="gpu") # or "ne" / "cpu"
out = agent.predict(
"The customer asks for a refund of a duplicate payment.",
{"refund": {"type": "noul", "instructions": "Does the customer request a refund?"}},
)
print(out["answers"])
Validate the download end-to-end (all three specializations, timing, contract checks):
python laya-coreai/scripts/validate.py laya-coreai
# PASS predict(gpu/ne/cpu) ... gpu p50≈5.0ms · ne p50≈5.0ms · cpu p50≈11.3ms
# RESULT: all checks passed
Snake demo with the model (terminal game, reuses the laya-coreml UI + safety
shield; automatically uses the B3 asset when present for ~2x game throughput):
pip install rich
python laya-coreai/scripts/snake_play.py --fps 12
Measured on M3 Max / macOS 27.2
| Backend |
Single-pass p50 |
Snake decision parity vs Core ML W8 |
| this model, GPU (recommended) |
4.6–5.0 ms |
40/40 argmax, max KL 1.5e-4 |
| this model, Neural Engine |
4.6–5.0 ms |
40/40 argmax, max KL 1.5e-4 |
| this model, CPU-only |
10.8–11.3 ms |
40/40 argmax, max KL 4.1e-4 |
Core ML W8 (laya-coreml ane-w8, CPU+ANE) |
13.5 ms |
baseline |
| Core ML FP16 flexible (CPU+GPU) |
17.1 ms |
exact weights |
MLX FP16 (laya-mlx) |
~10.8 ms |
exact weights |
~3× faster per pass than the fastest Core ML bundle at fixed shape, with
better fidelity than the W8 quantization it is compared against (W8 drifts up
to 0.014 calibration on the upstream fixture; this asset drifts 2.1e-4).
Always pin the compute unit
laya_coreai.load(unit="gpu") (default) or unit="ne". The unpinned
default specialization routes part of this graph to the ANE compiler, which
fails type inference (anec.scaled_elementwise memref mismatch) on macOS
27.2; the failed program load SIGABRTs the process. GPU is the robust choice;
NE works but inherits that fragility.
Batched asset: one pass per snake decision
laya-f16-b3.aimodel (same repo) exports batch 3, letting the three snake
questions (move / risk / food) share one forward pass:
| Backend (decision = 3 questions) |
p50 per decision |
| this model, B3 asset, GPU |
5.6–6.0 ms (175 decisions/s) |
| this model, B1 asset, GPU |
14.3–14.6 ms (3 passes) |
| Core ML W8 bundle (CPU+ANE) |
11.6–13.5 ms |
| MLX FP16 |
~10.8 ms + prompt build |
B3 parity: 100/100 argmax agreement vs both the B1 asset (max KL 0.00000) and
the Core ML W8 bundle (max KL 0.00014). Load it with
laya_coreai.load("laya-coreai", unit="gpu", asset="laya-f16-b3.aimodel").
Format and limits
- Fixed shape: 96 tokens total, 32 option slots, per question (same
capacity class as the Core ML
ane-w8 bundle). B1 asset runs one question
per pass; B3 asset batches three. Over-capacity prompts raise; nothing is
silently truncated.
- FP16 weights (
main.mlirb, 615 MB), logits/action outputs FP32.
- Rebuild from the pinned source with
scripts/build_aimodel.py (torch
export → coreai-torch → .aimodel, ~90 s, sha256-gated provenance).
coreai_config.json records shapes, provenance, and the validation record.
Provenance
- Original checkpoint:
convaiinnovations/laya-multilingual at 052592a15d198d9ad47da779604259b10b47b7aa.
- Original weights SHA256:
9d628fd971b700382ac6f65920a86f149777b2e748e0c955fb3b19695aa8f204.
- Graph:
laya_coreml.torch_model.DecisionModel (Apache-2.0); upstream implementation NandhaKishorM/laya.
- Converter: apple/coreai-torch 0.4.2 / coreai-core 1.0.0b2 (BSD-3).
- Original model by Convai Innovations and contributors, Apache-2.0.
- Independent conversion; not an official Convai Innovations or Apple release.
See LICENSE and NOTICE. This is an inference port, not a newly trained
decision model; task/language limitations originate with Laya.
Own MLX engine (engine="mlx") — immune to the Apple runtime leak
Known issue in coreai-core 1.0.0b2 (Apple beta): its GPU/NE paths allocate
every inference output from a process-global Metal ioSurface pool that
never recycles blocks. After ~6,000–8,000 rapid calls the pool dries and the
runtime hard-traps the process:
CoreAIRuntime/NDArray+Pool.swift:77: Fatal error: Failed to allocate storage
for NDArray with byteCount: 24, sk: ioSurface. At game speeds (--max-speed)
that is reached in ~90 s; it is an Apple bug (fixed only upstream), not a
model or wrapper bug. Workaround for paced play: --unit ne at ≤60 fps.
This repo therefore ships a from-scratch MLX engine for the same graph
(laya_coreai/mlx_engine.py): ModernBERT encoder + decision head + scorer +
action head, implemented directly against the original FP32 checkpoint and
fused with mx.compile. No CoreAIRuntime involvement — leak-free by
construction.
Verified against the laya-f16-b3.aimodel asset on the same inputs
(B=3 × L=96, macOS 27.2, M3 Max):
|
own MLX engine |
Core AI asset (GPU) |
| 3-question decision p50 |
6.0 ms |
5.6 ms |
| decision parity (40 prompts) |
40/40 argmax, max KL 2.0e-4 |
baseline |
| noul parity |
max |Δ| 0.003, 80/80 sign agreement |
baseline |
| leak soak |
20,000 calls, fds flat, RSS flat |
hard-trap at ~6k |
This repo ships the engine's checkpoint (model.safetensors, 633 MB,
fp32 originals of the same weights the .aimodel carries, Apache-2.0 like
the upstream convaiinnovations/laya-multilingual), so a full bundle
download is all you need:
hf download AndyInQtr/laya-coreai --local-dir laya-coreai
cd laya-coreai && pip install mlx laya-coreml numpy
python scripts/snake_play.py --engine mlx --max-speed # leak-free, ~6 ms/decision
python scripts/validate.py # certifies every backend
Python API:
import laya_coreai
agent = laya_coreai.load(engine="mlx") # finds model.safetensors in the bundle
out = agent.predict("Safe route: yes.", questions) # one fused pass