SAVRN
Search Contact SAVRN

Open-weight model · Text generation

StandardOne-3B-GGUF

by Standard Thinking StandardThinking/StandardOne-3B-GGUF

StandardOne-3B-GGUF is an open-weight model for text generation from Standard Thinking, released under Apache License 2.0. Its published files total 24.2 GB. It draws 1.4k downloads a month.

GGUF builds of Standard One 3B (Ministral 3 3B text + Pixtral vision tower, mistral3 architecture) for use with llama.cpp.

Parameters—
Context—
Weights24.2 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.4k

Model Card

By Standard Thinking, published under apache-2.0, revision 0a2e0a975270.

GGUF builds of Standard One 3B (Ministral 3 3B text + Pixtral vision tower, mistral3 architecture) for use with llama.cpp. Most quant levels below (marked "imatrix") were built with an importance matrix calibrated on 1,512 prompts sampled from our own training rows (see "Importance-matrix calibration" below), which recovers some of the accuracy quantization would otherwise lose. Q80 and BF16 don't need one. SHA256 checksums: SHA256SUMS. Source revisions, conversion tool version and full validation Q4KM note: this is the imatrix-calibrated version, not a plain quantization. We generated both and chose whichever scored higher on the mean of 10 held-out and public-dataset decision suites (no…

Read Standard Thinking's full model card

Updated weights (v2, 2026-09-26). These files are built from Standard One 3B v2. If you downloaded them before, download them again or pin revision="v2". Earlier versions stay available under the tags v1 and v1.1.

Version: v2

GGUF builds of Standard One 3B (Ministral 3 3B text + Pixtral vision tower, mistral3 architecture) for use with llama.cpp.

Files

Most quant levels below (marked "imatrix") were built with an importance matrix calibrated on 1,512 prompts sampled from our own training rows (see "Importance-matrix calibration" below), which recovers some of the accuracy quantization would otherwise lose. Q8_0 and BF16 don't need one.

File Quant Size imatrix
StandardOne-3B-BF16.gguf BF16 (no quantization) 6.9 GB —
StandardOne-3B-Q8_0.gguf Q8_0 3.7 GB no
StandardOne-3B-Q5_K_M.gguf Q5_K_M 2.5 GB yes
StandardOne-3B-Q4_K_M.gguf Q4_K_M 2.1 GB yes
StandardOne-3B-IQ4_XS.gguf IQ4_XS 2.0 GB yes
StandardOne-3B-Q3_K_M.gguf Q3_K_M 1.8 GB yes
StandardOne-3B-IQ3_M.gguf IQ3_M 1.7 GB yes
StandardOne-3B-Q2_K.gguf Q2_K 1.5 GB yes
StandardOne-3B-IQ2_M.gguf IQ2_M 1.3 GB yes
mmproj-StandardOne-3B.gguf F16 vision projector 840 MB —

SHA256 checksums: SHA256SUMS. Source revisions, conversion tool version and full validation numbers: release-manifest.json.

Q4_K_M note: this is the imatrix-calibrated version, not a plain quantization. We generated both and chose whichever scored higher on the mean of 10 held-out and public-dataset decision suites (no JevBench items; imatrix 69.77 vs. plain 69.63).

Usage

Text-only:

llama-cli -m StandardOne-3B-Q4_K_M.gguf -ngl 99 -p "Your prompt"

With vision (image input):

llama-server -m StandardOne-3B-Q4_K_M.gguf --mmproj mmproj-StandardOne-3B.gguf -ngl 99

The GGUF's own embedded chat template (converted from the model's chat_template.jinja) is applied automatically; no extra flags needed for chat formatting.

Importance-matrix calibration

Q5_K_M down to IQ2_M were quantized with llama-imatrix calibrated on 1,512 prompts (42 per cohort across 36 training-data cohorts; training data only — no benchmark/held-out file was used), context 2048. Q8_0 and BF16 don't use an imatrix (high enough precision that it doesn't move the needle).

Validation

Accuracy was checked by comparing next-token logits over the option letters on the JevBench public suites (easy/original/hard) against the served BF16 baseline; see release-manifest.json for full methodology and gguf-validation.md (in the release kit) for the complete writeup. Measured (accuracy %, n=48/72/111 for easy/original/hard; prefill tokens/sec is the mean over the 231 scored decisions):

Quant Easy Original Hard Overall Prefill tok/s
served BF16 (reference) 100.0 88.89 46.85 — —
BF16-GGUF 100.0 87.50 51.35 72.73 11,672
Q8_0 100.0 87.50 51.35 72.73 7,842
Q5_K_M 100.0 88.89 46.85 71.00 4,966
Q4_K_M (shipped, imatrix) 100.0 88.89 45.95 70.56 6,770
IQ4_XS 100.0 90.28 52.25 74.03 6,181
Q3_K_M 100.0 87.50 47.75 71.00 6,565
IQ3_M 100.0 83.33 45.95 68.83 6,632
Q2_K 100.0 80.56 40.54 65.37 4,774
IQ2_M 97.92 70.83 49.55 66.23 5,082

Note on the mmproj conversion

llama.cpp's stock --mmproj converter (as of the commit used here) drops the [IMG_BREAK] token embedding for HF-format Mistral3ForConditionalGeneration checkpoints (a filter meant to strip text-model tensors also strips the one row of the text embedding matrix the vision projector needs), so the mmproj file it produces fails to load in llama-server/llama-cli ("unable to find tensor v.token_embd.img_break"). The mmproj file in this folder was built with a small local patch that lets that one tensor through; see release-manifest.json -> known_issues_fixed for details. It loads and runs correctly with --mmproj.

License

Apache License 2.0 — see LICENSE and NOTICE. Same terms as the source StandardOne-3B release; this GGUF conversion adds no additional restrictions.

Identity and Version

Repository
StandardThinking/StandardOne-3B-GGUF
Publisher
Standard Thinking
Task
Text generation
Modality
Text
Library
gguf
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
0a2e0a9752708897d32303b3e82b3268cf3fd971
First published
2026-09-24
Last updated
2026-09-27

Files and Weights

16 files, 24.2 GB in total. The weights are 10 files totalling 24.2 GB in gguf.

Weights10 files · 24.2 GB
Configuration1 file · 8.6 KB
Documentation3 files · 17.2 KB
Other1 file · 922 B
Repository1 file · 2.1 KB
Every file
FileTypeSizeSHA-256
StandardOne-3B-BF16.ggufWeights6.9 GB 09128a1bdda5
StandardOne-3B-IQ2_M.ggufWeights1.3 GB 6ccba092abf6
StandardOne-3B-IQ3_M.ggufWeights1.7 GB 5dfc313a4ae4
StandardOne-3B-IQ4_XS.ggufWeights2.0 GB 1e672182fe56
StandardOne-3B-Q2_K.ggufWeights1.5 GB 72c41bb2d7c4
StandardOne-3B-Q3_K_M.ggufWeights1.8 GB fd18648a83ca
StandardOne-3B-Q4_K_M.ggufWeights2.1 GB b7445f4d4cc9
StandardOne-3B-Q5_K_M.ggufWeights2.5 GB 8848fd1787fa
StandardOne-3B-Q8_0.ggufWeights3.7 GB 4d743351fc7e
mmproj-StandardOne-3B.ggufWeights840.3 MB 4b0688be9f8d
release-manifest.jsonConfiguration8.6 KB —
LICENSEDocumentation11.3 KB —
NOTICEDocumentation1.2 KB —
README.mdDocumentation4.6 KB —
SHA256SUMSOther922 B —
.gitattributesRepository2.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
24.2 GB
Download from Standard Thinking

Released by Standard Thinking through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published24.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About StandardOne-3B-GGUF

Can I use StandardOne-3B-GGUF commercially?

Yes. StandardOne-3B-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers