SAVRN
Search Contact SAVRN

Open-weight model · Text generation

A11OY-MINI

by SZL Holdings SZLHOLDINGS/A11OY-MINI

749/14/163 · Λ = Conjecture 1 (advisory) · a-11-oy.com GGUFs are on this repo. GGUF AVAILABLE. Bytes MEASURED (size + sha256 below, ATELIER-verified). Evals none-this-run. Not MEASURED as an eval. publicationeligible: false. Quality ROADMAP in prose only.

Parameters
Context
Weights2.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads163

Model Card

By SZL Holdings, published under apache-2.0, revision c936dc749743.

749/14/163 · Λ = Conjecture 1 (advisory) · a-11-oy.com GGUFs are on this repo. GGUF AVAILABLE. Bytes MEASURED (size + sha256 below, ATELIER-verified). Evals none-this-run. Not MEASURED as an eval. publicationeligible: false. Quality ROADMAP in prose only. Lab load forbidden. GPU UNAVAILABLE. SKU of live SZLHOLDINGS/chaski merged shard 1c55df8. Not a new train. Not SZLHOLDINGS/chaski-5050. Not a Hub basemodelrelation: quantized child until a Chaski eval gate passes. A11oy is a Space and a substrate. The mini is the voice you can actually load. Not the organ. Not the mesh. A governed command voice that fits in llama.cpp. Nobody else ships this combination. That is the point of a one-of-one.…

Read SZL Holdings's full model card

Rebuilt 2026-09-17 from the chaski-r2 winner (named-N gate champion). Legacy GGUFs below are DEPRECATED failed-parent lineage. Not flagship. Not a11oy production.

A 1 1 O Y  M I N I

Kanchay — Quechua for light. Folded small enough to carry.

KANCHAY · Doctrine v11 · Lean 749/14/163 · Λ = Conjecture 1 (advisory) · a-11-oy.com

GGUFs are on this repo. GGUF AVAILABLE. Bytes MEASURED (size + sha256 below, ATELIER-verified). Evals none-this-run. Not MEASURED as an eval. publication_eligible: false. Quality ROADMAP in prose only. Lab load forbidden. GPU UNAVAILABLE.

SKU of live SZLHOLDINGS/chaski merged shard 1c55df8. Not a new train. Not SZLHOLDINGS/chaski-5050. Not a Hub base_model_relation: quantized child until a Chaski eval gate passes.

The cut

A11oy is a Space and a substrate. The mini is the voice you can actually load. Not the organ. Not the mesh.

A governed command voice that fits in llama.cpp.

Silhouette → leave → SZL

Leader Take, then tweak
Anthropic A small Claude-shaped mouth with a much stricter constitution.
NVIDIA A NIM-less local runtime.
Unsloth GGUF path.

Nobody else ships this combination. That is the point of a one-of-one.

Intended use

Local doctrine voice. Proposal-only if wired to tools.

Limitations

  • Derived quant.
  • Not a11oy-v19-substrate.

Canonical GitHub: szl-holdings/szl-forge

SKU SZLHOLDINGS/A11OY-MINI
Parent live SZLHOLDINGS/chaski merged shard 1c55df8652e9d0f7b84356b1e2d54849165ae884
Not parent SZLHOLDINGS/chaski-5050
Silhouette base Qwen/Qwen3.5-0.8B (named in prose only; no YAML base_model)
F16 a11oy-mini-f16.gguf 1557662240 sha256 a5df00e4e3ca07f65a4b43aad4ef1505625952a3105e6dcc0dba87f2fa35fc57 MEASURED
Q4_K_M a11oy-mini-q4_k_m.gguf 541903392 sha256 06136ba385b2e052cf4cdb3dc8d333948e0b612bd15a541b314e170399c2faa6 MEASURED
R2 parent chaski-r2 local named-N winner (GitHub receipts only; no Hub page)
R2 Q4_K_M a11oy-mini-r2-Q4_K_M.gguf 541903328 sha256 6d42341c932a76e91b2c04a859a4248e7d2c77308f998a32771f43802f097b62 MEASURED - GGUF gate 5/5 + 6/6 (ollama, 2026-09-17)
R2 mmproj a11oy-mini-r2-BF16-mmproj.gguf 207346048 sha256 4855efe034435b9b3b289b2c07d09b263c088749d1f85c43c1cf4672bc7fcbf2 MEASURED (vision projector pair)
Convert llama.cpp F16 then Q4_K_M
Banned Direct safetensors→Ollama. Lab load.
Evals none-this-run. Not 5/5. Bytes MEASURED is not an eval.
Quality ROADMAP (prose only; YAML tag omitted)
Publication publication_eligible: false
Lab Forbidden. House CPU lab stays Khipu GGUF. Do not load this SKU in inference-lab.
License Apache-2.0
Doctrine v11 LOCKED. Seed 11. Λ = Conjecture 1 (advisory, never a theorem).

What this is NOT

  • Not a published model eval and not 5/5
  • Not chaski-5050
  • Not a base_model_relation: quantized Hub child (Chaski eval gate has not passed)
  • Not a lab pin and not loadable in the house lab
  • Not a tokens/s claim
  • Not a new train and not a third LLM

Hub: SZLHOLDINGS/A11OY-MINI

Identity and Version

Repository
SZLHOLDINGS/A11OY-MINI
Publisher
SZL Holdings
Task
Text generation
Modality
Text
Library
llama.cpp
Parameters
Not stated by the source
Languages
en
Revision
c936dc749743c94586706345a0142c79094a581c
First published
2026-08-28
Last updated
2026-09-18

Files and Weights

12 files, 2.9 GB in total. The weights are 5 files totalling 2.8 GB in gguf.

Weights5 files · 2.8 GB
Configuration4 files · 7.0 KB
Documentation1 file · 5.5 KB
Other1 file · 2.1 MB
Repository1 file · 1.8 KB
Every file
FileTypeSizeSHA-256
Modelfile.a11oy-r2.ggufWeights142 B
a11oy-mini-f16.ggufWeights1.6 GB a5df00e4e3ca
a11oy-mini-q4_k_m.ggufWeights541.9 MB 06136ba385b2
a11oy-mini-r2-BF16-mmproj.ggufWeights207.3 MB 4855efe03443
a11oy-mini-r2-Q4_K_M.ggufWeights541.9 MB 6d42341c932a
bom/model-bom.cdx.jsonConfiguration3.4 KB
conversion_receipt.jsonConfiguration2.2 KB
gguf_gate_receipt.jsonConfiguration636 B
hub_put_receipt.jsonConfiguration794 B
README.mdDocumentation5.5 KB
og-card.pngOther2.1 MB c96e2f2c60b8
.gitattributesRepository1.8 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
2.8 GB
Download from SZL Holdings

Released by SZL Holdings through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published2.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About A11OY-MINI

Can I use A11OY-MINI commercially?

Yes. A11OY-MINI is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers

Uncensored Qwen3.8-27B, published as GGUF quantizations with the multi token prediction (MTP) head retained and verified. Refusal behaviour has been substantially reduced, not eliminated. See Measured behaviour for the numbers. Capabilities, training data, and architecture are otherwise unchanged. - Refusal directions removed with Heretic, which co minimizes refusal count against KL divergence from the base model. No handwritten refusal removal code, no finetuning, no additional training data. - Abliteration runs at bf16 (no 4 bit quantization). the resulting LoRA is merged into the bf16 base, so the published weights are not a quantized round trip. - mtp. tensors are copied verbatim from…

Open weights apache-2.0 llama.cpp