SAVRN Model Hub
Open-Weight Models
An open-weight model is an AI model whose trained weights are published for anyone to download. The weights are what the model learned in training. With a copy of them you can run the model on hardware you control and train it further on your own data.
Open weights are not the same as open source. Many publishers release the weights without the training data or code, and the license sets what you may do with the model. This library puts each model's full card, architecture, files, license and published evaluations on one page.
Updated 2026-09-20 · How the library is built
3,247 models, sorted by most downloaded.
Laya typed decisions on Apple Silicon, using CPU + GPU. This is a portable Core ML bundle for laya-coreml, converted from convaiinnovations/laya. It outputs choice, score, and noul probabilities with zero generated tokens. Inference needs no PyTorch, Transformers, MLX, remote code, or cloud API. Apple Silicon, macOS 15+, Python 3.11–3.13. Tested on M3 Max / macOS 27.2. To download explicitly and then run entirely offline: Use laya.load("aac6fef/laya-coreml", localfilesonly=True) for a cached snapshot or pass a local directory. Use revision=" " to pin a remote revision. This FP16 export retains the original model architecture and decision schema. The enumerated-length GPU export is the…
Laya Multilingual is a fast, non-autoregressive System 1 decision engine supporting over 100 languages natively. It evaluates typed schemas (choice, score, noul) over unstructured state (text, email, ticket, or JSON) in a single ~35ms forward pass on a GPU, returning mathematically calibrated probabilities and confidence scores with zero text generation and zero hallucinations. Trained on over 1,000,000 multi-domain records using RLCD (Reinforcement Learning for Calibrated Decisions) with strictly proper scoring rules. Released under the Apache 2.0 License by Convai Innovations.
Laya typed decisions on Apple Silicon, using CPU + GPU. This is a portable Core ML bundle for laya-coreml, converted from convaiinnovations/laya-multilingual. It outputs choice, score, and noul probabilities with zero generated tokens. Inference needs no PyTorch, Transformers, MLX, remote code, or cloud API. Apple Silicon, macOS 15+, Python 3.11–3.13. Tested on M3 Max / macOS 27.2. To download explicitly and then run entirely offline: Use laya.load("aac6fef/laya-multilingual-coreml", localfilesonly=True) for a cached snapshot or pass a local directory. Use revision=" " to pin a remote revision. This FP16 export retains the original model architecture and decision schema. The…
Laya typed decisions on Apple Silicon, using CPU + Neural Engine. This is a portable Core ML bundle for laya-coreml, converted from convaiinnovations/laya-multilingual. It outputs choice, score, and noul probabilities with zero generated tokens. Inference needs no PyTorch, Transformers, MLX, remote code, or cloud API. Apple Silicon, macOS 15+, Python 3.11–3.13. Tested on M3 Max / macOS 27.2. To download explicitly and then run entirely offline: Use laya.load("aac6fef/laya-multilingual-coreml-ane", localfilesonly=True) for a cached snapshot or pass a local directory. Use revision=" " to pin a remote revision. This FP16 conversion retains the original trained parameters. The ANE graph's host…
Laya typed decisions on Apple Silicon, using CPU + Neural Engine. This is a portable Core ML bundle for laya-coreml, converted from convaiinnovations/laya-multilingual. It outputs choice, score, and noul probabilities with zero generated tokens. Inference needs no PyTorch, Transformers, MLX, remote code, or cloud API. Apple Silicon, macOS 15+, Python 3.11–3.13. Tested on M3 Max / macOS 27.2. To download explicitly and then run entirely offline: Use laya.load("aac6fef/laya-multilingual-coreml-ane-w8", localfilesonly=True) for a cached snapshot or pass a local directory. Use revision=" " to pin a remote revision. This is a separately identified 8-bit grouped K-means weight-palette variant.…
Laya typed decisions on Apple Silicon, using CPU + GPU. This is a portable Core ML bundle for laya-coreml, converted from convaiinnovations/laya-multilingual. It outputs choice, score, and noul probabilities with zero generated tokens. Inference needs no PyTorch, Transformers, MLX, remote code, or cloud API. Apple Silicon, macOS 15+, Python 3.11–3.13. Tested on M3 Max / macOS 27.2. To download explicitly and then run entirely offline: Use laya.load("aac6fef/laya-multilingual-coreml-snake", localfilesonly=True) for a cached snapshot or pass a local directory. Use revision=" " to pin a remote revision. This FP16 export retains the original model architecture and decision schema. The…
Laya typed decisions on Apple Silicon, using CPU + GPU. This is a portable Core ML bundle for laya-coreml, converted from convaiinnovations/laya-typed-decisions. It outputs choice, score, and noul probabilities with zero generated tokens. Inference needs no PyTorch, Transformers, MLX, remote code, or cloud API. Apple Silicon, macOS 15+, Python 3.11–3.13. Tested on M3 Max / macOS 27.2. To download explicitly and then run entirely offline: Use laya.load("aac6fef/laya-typed-decisions-coreml", localfilesonly=True) for a cached snapshot or pass a local directory. Use revision=" " to pin a remote revision. This FP16 export retains the original model architecture and decision schema. The…
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019). - PEFT 0.21.0
This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,511 tokens. The source MXFP4 checkpoint was explicitly dequantized to BF16 before scoring and pruning. The source-row selection hash is…
This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…
This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…
This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…
Lightning is a small, autoregressive transformer which utilizes FlashAttention and MHA. This model is trained on a variety of books from a dataset(300 MB). This is the expanded version of Lightning-60m. Lightning utilizes FlashAttention and AdamW for performance and capability. Lightning is designed to provide quick, coherent outputs, improved with a larger size and weight. Lightning is intended to be used for research, analysis and fine-tuning, stories and other. It is not intended to be used for professional advice, real writing or any kind of heavy work as generated outputs may be incorrect. Lightning can be used directly for text generation, experimentation, and conversational…
Public, reproducible inference for the Tartan IMU Challenge (IROS 2026). Task. Given 6-axis IMU windows (acc 3 + gyro 3, length 200) covering a motion window plus its temporal context, predict body-frame velocity (vx, vy, vz) for all four platforms (car / dog / drone / human) with a single unified model. Model. lihuaimuiros2026v1: a shared ResNet1D encoder with an FFT branch, followed by a Transformer over a triplet of consecutive windows (prev / curr / next), and a 4-expert Mixture-of-Experts head. The expert gate is a soft router driven purely by IMU-derived features — experts are learned motion regimes (general / slow / fast / agile), not platforms. A single frozen checkpoint (seed 0) is…
Launch manifests for the lil launcher. Each entry is a directory holding one lil.yaml that names the repository holding the weights, optionally pins the commit the local inference lab qualified, and states the serving policy the launcher cannot read from the checkpoint itself. Checkpoint facts come from the weight repository at the resolved commit. Entries of kind: draft describe speculative-decoding drafts and are not served directly. The entry directory name is the launch name: The schema is documented in the launcher repository.
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Open-Weight Models Explained
What is an open-weight model?
An AI model whose trained weights are published for anyone to download, so it can be run, tested and fine-tuned on hardware the user controls.
Is an open-weight model the same as open source?
Not always. Open weights means the trained model can be downloaded. Open source usually also means the training code and data are available and the license allows broad reuse. Many open-weight models release the weights only.
Can I use an open-weight model commercially?
It depends on the license. Apache 2.0 and MIT allow commercial use. Other licenses limit it, for example to non-commercial use or below a set number of users. Every model page here shows its license.
How much memory does an open-weight model need?
About two bytes per parameter at 16-bit precision, so a 7-billion-parameter model needs roughly 14 GB for its weights, plus memory for the context it processes. Each model page lists its parameter count and the size of its files.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.

