A hidden Markov model distilled from meta-llama/Llama-3.1-8B-Instruct, used as the tractable proposal for SQL-constrained text-to-SQL on Spider in Mitigating Bias in Locally Constrained Decoding via Tractable Proposals (arXiv:2606.01926). Locally constrained decoding (LCD) enforces a constraint by masking next tokens one step at a time, which biases the samples it draws. P-GCD instead combines this HMM with a tensorized finite automaton to form a proposal that carries both the logical and the probabilistic information about the target distribution, and samples from it with sequential Monte Carlo. The HMM supplies the lookahead the This is not a fine-tune of the base LM and shares none of…
SAVRN Model Hub
Open-Weight Models
An open-weight model is an AI model whose trained weights are published for anyone to download. The weights are what the model learned in training. With a copy of them you can run the model on hardware you control and train it further on your own data.
Open weights are not the same as open source. Many publishers release the weights without the training data or code, and the license sets what you may do with the model. This library puts each model's full card, architecture, files, license and published evaluations on one page.
Updated 2026-09-18 · How the library is built
2,760 models, sorted by most downloaded.
A hidden Markov model distilled from meta-llama/Llama-3.1-8B-Instruct, used as the tractable proposal for constrained tool calling on xLAM (function-call template) in Mitigating Bias in Locally Constrained Decoding via Tractable Proposals (arXiv:2606.01926). Locally constrained decoding (LCD) enforces a constraint by masking next tokens one step at a time, which biases the samples it draws. P-GCD instead combines this HMM with a tensorized finite automaton to form a proposal that carries both the logical and the probabilistic information about the target distribution, and samples from it with sequential Monte Carlo. The HMM supplies the lookahead the This is not a fine-tune of the base LM…
A hidden Markov model distilled from meta-llama/Llama-3.1-8B-Instruct, used as the tractable proposal for constrained tool calling on xLAM (JSON template) in Mitigating Bias in Locally Constrained Decoding via Tractable Proposals (arXiv:2606.01926). Locally constrained decoding (LCD) enforces a constraint by masking next tokens one step at a time, which biases the samples it draws. P-GCD instead combines this HMM with a tensorized finite automaton to form a proposal that carries both the logical and the probabilistic information about the target distribution, and samples from it with sequential Monte Carlo. The HMM supplies the lookahead the This is not a fine-tune of the base LM and shares…
基于 google/gemma-4-12B-it 的 LoRA 适配器。使用经预处理和筛选的六个公开数据来源,以 RegMix config-025 配比组成 50,000 条训练数据;不包含 Lusy,也不以 Lusy 为筛选或优化目标。 row% 为样本数占比,target token 为按本次 Gemma tokenizer 和训练监督区间统计的单轮数据集监督 token 总数,tgt% 为其占比;中位数按逐条样本的监督 token 数计算,合计行为全部 50,000 条的中位数。百分比四舍五入。训练共 2 个 epoch;表中不重复计数。不同基座 tokenizer 下的 token 数不能与 v1 直接等同比较。 六个来源分别为 CoSER、XPersona、Aya、Tulu-3-SFT-Mixture、SmolTalk 和 Infinity-Instruct。混合数据包含 50,000 个唯一 ID、50,000 个唯一内容哈希,与留出集的分组重叠数为 0。 训练集 SHA-256:8c626b338270652de2e4d1130b28d8b0fc2c2eca0b90f5dcffa7580a26e0d859。 配比来自 64 个 Gemma-E2B 代理配置及候选确认实验。此版本发布已完成的 config-025 正式训练产物;不将其宣称为 Gemma 12B 上已证明的全局最佳配比。 本仓库包含 LoRA 权重、适配器配置、tokenizer 和 chat template;不包含合并后的基座权重。加载时使用 google/gemma-4-12B-it…
Full native Orbax checkpoint: model parameters, Adam/gradient-accumulation state, and saved optimizer/data-progress metadata. This is not a Transformers safetensors export. Training uses camel-ai/gsm8kdistilled, 6,144-token examples, completion-only loss, global batch 32, and seed 42. One optimizer update consumed 32 examples (two source microbatches on 16 devices). See recipe.json and checkpoint-manifest.json for pinned revisions and hashes. Load the native checkpoint root nativecheckpoint at step 1 with the pinned MaxText/Tunix runtime. Restoring onto a different device topology or accumulation schedule requires explicit sharding and data-position validation; that portability has not yet…
This is a GGUF LoRA adapter, not a standalone model. It is intended for the anonymization.ner entity-extraction route in Jurilix. Apply it at strength 0.75 for that route and explicitly send strength 0 for other routes. 4b4a2c1d584be7264f87aac328a1bc739ce81b6c, file gemma-4-E4Bq40-it.gguf. The published source recipe pins upstream llama.cpp and a Gemma JSON parser fix. GPU performance or inference quality on other platforms. Start llama-server with --lora ADAPTER.gguf --lora-init-without-apply --cache-ram 0. For entity extraction, include "lora": [{"id": 0, "scale": 0.75}] in each request. For every other route, include "lora": [{"id": 0, "scale": 0}]. Disabling the RAM prompt cache is…
A research-oriented ViT prototype targeting Generation. The included giant setup documents defaults and file formats without presenting unverified performance numbers. - The Python file contains the model and runnable example or training entry point. - config.json records the generated architecture settings. - trainingargs.json records the default experiment recipe. - model.safetensors is a valid initialization checkpoint for smoke tests; it is not presented as a trained benchmark checkpoint. - No benchmark score is claimed in this repository. The included configuration uses lamb with a step schedule. These are starting values in the script, not evidence of a completed run. For a meaningful…
GlassEye? authorized HackerOne Bug Bounty (BBP) / Vulnerability Disclosure (VDP) assistant. In-scope HackerOne program workflows: policy/scope, report writing, severity rationale, remediation. Not for unauthorized testing or exploit dump recipes.
GlassEye? authorized HackerOne Bug Bounty (BBP) / Vulnerability Disclosure (VDP) assistant. In-scope HackerOne program workflows: policy/scope, report writing, severity rationale, remediation. Not for unauthorized testing or exploit dump recipes.
ONNX conversion of gliner-community/glinersmall-v2.5, dynamically quantised to INT8, packaged as a self-contained bundle for offline NER. This is a re-serialisation, not a fine-tune: the weights are the upstream ones. Only the format (PyTorch → ONNX) and the precision (fp32 → INT8) are ours. The fp32 reference graph (model.onnx, sha256 5245733ccb2b75072cce0b4bbb14424988f92f9daf775d97bdf0de74be28df63) is not shipped — it is only needed to reproduce the INT8 graph. Its hash is recorded in NOTICE. The graph has six inputs, fed per span-encoded prompt: spanmask is bool (not int64) — the graph declares tensor(bool). The prompt follows the GLiNER label format, using the model's own special tokens…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Open-Weight Models Explained
What is an open-weight model?
An AI model whose trained weights are published for anyone to download, so it can be run, tested and fine-tuned on hardware the user controls.
Is an open-weight model the same as open source?
Not always. Open weights means the trained model can be downloaded. Open source usually also means the training code and data are available and the license allows broad reuse. Many open-weight models release the weights only.
Can I use an open-weight model commercially?
It depends on the license. Apache 2.0 and MIT allow commercial use. Other licenses limit it, for example to non-commercial use or below a set number of users. Every model page here shows its license.
How much memory does an open-weight model need?
About two bytes per parameter at 16-bit precision, so a 7-billion-parameter model needs roughly 14 GB for its weights, plus memory for the context it processes. Each model page lists its parameter count and the size of its files.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.


