# StandardOne-3B by Standard Thinking: Open-Weight Model
Source: https://savrn.com/models/standardone-3b
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve StandardOne-3B (3.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 7.7 GB | 9.2 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 3.8 GB | 4.6 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 1.9 GB | 2.3 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[StandardOne-3B on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/standardone-3b/gpus)

## Model Card

By Standard Thinking, published under apache-2.0, revision 65f9140a1294.

Updated weights (v2, 2026-09-26). If you downloaded this model before, download it again or pin revision="v2". Earlier versions stay available under the tags v1 and v1.1.

Version: v2

Standard One scores a bounded set of answers for a supplied scenario and returns probabilities through POST /v1/systemone. It does not generate free-form response text. This repository contains the merged BF16 3B checkpoint; the server code is in [StandardOne-8B](https://savrn.com/models/standardone-8b).

| If you need | Repository |
| --- | --- |
| Merged 3B checkpoint | [StandardOne-3B](https://savrn.com/models/standardone-3b) (this repository) |
| 3B adapter weights and merge recipe | [StandardOne-3B-LoRA](https://savrn.com/models/standardone-3b-lora) |
| Larger merged checkpoint and server code | [StandardOne-8B](https://savrn.com/models/standardone-8b) |
| 8B adapter weights and merge recipe | [StandardOne-8B-LoRA](https://savrn.com/models/standardone-8b-lora) |

In the reported served evaluations, 3B has a lower median latency on the measured short-request profile; 8B scores higher on the public standard and hard tiers. See Benchmarks for the measurement conditions and limitations.

The figure combines results from different measurement paths. See Benchmarks for served versus offline conditions; measured 24–26 September 2026.

### At a glance

[Read the full model card (2,029 words)](https://savrn.com/models/standardone-3b/card)

## Configuration

Architecture

Mistral3ForConditionalGeneration

Context length (tokens)

262,144

Layers

26

Hidden size

3,072

Feed-forward size

9,216

Attention heads

32

Key/value heads

8

Head dimension

128

Vocabulary size

131,072

Model type

mistral3

## Identity and Version

Repository

StandardThinking/StandardOne-3B

Publisher

Standard Thinking

Task

Text generation

Modality

Text

Library

transformers

Parameters

3.8B parameters

Languages

en, ja, zh, es, fr, de, pt, ru

Revision

65f9140a1294c430495b306e98e2e618a21167b1

First published

2026-09-24

Last updated

2026-09-27

## Files and Weights

92 files, 7.7 GB in total. The weights are 2 files totalling 7.7 GB in safetensors.

Weights2 files · 7.7 GB

Configuration48 files · 17.3 MB

Tokenizer4 files · 17.3 MB

Documentation19 files · 139.7 KB

Other17 files · 2.0 MB

Repository2 files · 2.0 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model-00001-of-00002.safetensors | Weights | 5.0 GB | 7029a2415131 |
| model-00002-of-00002.safetensors | Weights | 2.7 GB | ad184b56624e |
| MERGE_REPORT.json | Configuration | 2.5 KB | — |
| config.json | Configuration | 1.7 KB | — |
| evidence/calibration-leaderboard-ablate-nosys.json | Configuration | 21.4 KB | — |
| evidence/calibration-leaderboard-merged-native.json | Configuration | 11.0 KB | — |
| evidence/calibration-per-type-served.json | Configuration | 14.2 KB | — |
| evidence/prompt-wording-choice.json | Configuration | 2.3 KB | — |
| evidence/served-nosys-easy-report.json | Configuration | 6.9 KB | — |
| evidence/served-nosys-hard-report.json | Configuration | 14.9 KB | — |
| evidence/served-nosys-latency-idle-gpu.json | Configuration | 43.6 KB | — |
| evidence/served-nosys-original-report.json | Configuration | 9.8 KB | — |
| generation_config.json | Configuration | 131 B | — |
| model.safetensors.index.json | Configuration | 45.6 KB | — |
| params.json | Configuration | 1.1 KB | — |
| processor_config.json | Configuration | 976 B | — |
| release-manifest.json | Configuration | 1.1 KB | — |
| server/benchmarks/data/jevbench-easy/public.manifest.json | Configuration | 3.3 KB | — |
| server/benchmarks/data/jevbench-hard/public.manifest.json | Configuration | 3.7 KB | — |
| server/benchmarks/data/jevbench-original/public.manifest.json | Configuration | 3.4 KB | — |
| server/benchmarks/run_matrix.py | Configuration | 9.4 KB | — |
| server/benchmarks/verify_metrics.py | Configuration | 5.1 KB | — |
| server/examples/request.json | Configuration | 542 B | — |
| server/examples/smoke.py | Configuration | 1.9 KB | — |
| server/jev_adapter/__init__.py | Configuration | 73 B | — |
| server/jev_adapter/__main__.py | Configuration | 7.5 KB | — |
| server/jev_adapter/backend.py | Configuration | 892 B | — |
| server/jev_adapter/benchmarks/__init__.py | Configuration | 81 B | — |
| server/jev_adapter/benchmarks/compare.py | Configuration | 2.8 KB | — |
| server/jev_adapter/benchmarks/data.py | Configuration | 8.5 KB | — |
| server/jev_adapter/benchmarks/jevbench.py | Configuration | 12.1 KB | — |
| server/jev_adapter/benchmarks/metrics.py | Configuration | 12.1 KB | — |
| server/jev_adapter/benchmarks/prepare.py | Configuration | 8.3 KB | — |
| server/jev_adapter/benchmarks/run.py | Configuration | 18.4 KB | — |
| server/jev_adapter/protocol.py | Configuration | 15.6 KB | — |
| server/jev_adapter/server.py | Configuration | 4.4 KB | — |
| server/jev_adapter/service.py | Configuration | 7.0 KB | — |
| server/jev_adapter/sglang.py | Configuration | 19.5 KB | — |
| server/tests/test_benchmark_data.py | Configuration | 8.7 KB | — |
| server/tests/test_benchmark_matrix.py | Configuration | 3.4 KB | — |
| server/tests/test_benchmark_run.py | Configuration | 10.0 KB | — |
| server/tests/test_disconnect_cleanup.py | Configuration | 5.1 KB | — |
| server/tests/test_http_integration.py | Configuration | 2.0 KB | — |
| server/tests/test_jevbench.py | Configuration | 11.0 KB | — |
| server/tests/test_main.py | Configuration | 11.5 KB | — |
| server/tests/test_protocol.py | Configuration | 17.8 KB | — |
| server/tests/test_service.py | Configuration | 17.1 KB | — |
| server/tests/test_sglang.py | Configuration | 16.0 KB | — |
| special_tokens_map.json | Configuration | 147.1 KB | — |
| tekken.json | Configuration | 16.8 MB | 600bb2794656 |
| LICENSE | Documentation | 11.3 KB | — |
| NOTICE | Documentation | 1.2 KB | — |
| QUICKSTART.md | Documentation | 4.5 KB | — |
| README.md | Documentation | 17.1 KB | — |
| docs/BENCHMARKS.md | Documentation | 26.0 KB | — |
| docs/public-classification-suites.md | Documentation | 3.3 KB | — |
| server/LICENSE | Documentation | 11.3 KB | — |
| server/NOTICE | Documentation | 931 B | — |
| server/README.md | Documentation | 17.5 KB | — |
| server/benchmarks/DECISION_BENCHMARK_SELECTION.md | Documentation | 5.2 KB | — |
| server/benchmarks/METRIC_VALIDATION.md | Documentation | 3.8 KB | — |
| server/benchmarks/PUBLIC_DATASETS.md | Documentation | 13.4 KB | — |
| server/benchmarks/README.md | Documentation | 12.8 KB | — |
| server/benchmarks/data/jevbench-easy/LICENSE | Documentation | 1.1 KB | — |
| server/benchmarks/data/jevbench-easy/THIRD-PARTY.md | Documentation | 2.7 KB | — |
| server/benchmarks/data/jevbench-hard/LICENSE | Documentation | 1.1 KB | — |
| server/benchmarks/data/jevbench-hard/THIRD-PARTY.md | Documentation | 2.7 KB | — |
| server/benchmarks/data/jevbench-original/LICENSE | Documentation | 1.1 KB | — |
| server/benchmarks/data/jevbench-original/THIRD-PARTY.md | Documentation | 2.7 KB | — |
| SHA256SUMS | Other | 8.8 KB | — |
| SYSTEM_PROMPT.txt | Other | 2.4 KB | — |
| chat_template.jinja | Other | 11.9 KB | — |
| docs/assets/00-benchmark-card.png | Other | 298.1 KB | 1395bca8ffbf |
| docs/assets/00-benchmark-card.svg | Other | 58.7 KB | — |
| docs/assets/02-latency-vs-qwen.png | Other | 114.5 KB | 7313f2b0645e |
| docs/assets/02-latency-vs-qwen.svg | Other | 11.3 KB | — |
| docs/assets/04-throughput.png | Other | 97.7 KB | — |
| docs/assets/04-throughput.svg | Other | 8.7 KB | — |
| docs/assets/05-gain-over-base.png | Other | 179.7 KB | 4900de357045 |
| docs/assets/05-gain-over-base.svg | Other | 16.6 KB | — |
| docs/assets/06-vs-jev.png | Other | 266.4 KB | 0274d259fe3e |
| docs/assets/06-vs-jev.svg | Other | 23.0 KB | — |
| server/benchmarks/data/jevbench-easy/public.jsonl | Other | 60.7 KB | — |
| server/benchmarks/data/jevbench-hard/public.jsonl | Other | 708.2 KB | — |
| server/benchmarks/data/jevbench-original/public.jsonl | Other | 91.9 KB | — |
| server/pyproject.toml | Other | 808 B | — |
| .gitattributes | Repository | 1.9 KB | — |
| server/.gitignore | Repository | 125 B | — |
| server/jev_adapter/native_tokenizer.py | Tokenizer | 7.5 KB | — |
| server/tests/test_native_tokenizer.py | Tokenizer | 14.1 KB | — |
| tokenizer.json | Tokenizer | 17.1 MB | d5f6046775b1 |
| tokenizer_config.json | Tokenizer | 198.1 KB | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

7.7 GB

[Download from Standard Thinking](https://huggingface.co/StandardThinking/StandardOne-3B)

Released by Standard Thinking through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from mistralai/Ministral-3-3B-Instruct-2512-BF16

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 7.7 GB |
| 16-bit | 7.7 GB |
| 8-bit | 3.8 GB |
| 4-bit | 1.9 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Built on This Model

- Quantized from[StandardOne-3B-GGUF](https://savrn.com/models/standardone-3b-gguf)
- Derived from[StandardOne-3B-GGUF](https://savrn.com/models/standardone-3b-gguf)
- Quantized from[StandardOne-3B-FP8](https://savrn.com/models/standardone-3b-fp8)
- Derived from[StandardOne-3B-FP8](https://savrn.com/models/standardone-3b-fp8)

## Questions About StandardOne-3B

### How much GPU memory does StandardOne-3B need?

About 9.2 GB at 16-bit and 2.3 GB at 4-bit: the weights (3.8B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run StandardOne-3B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use StandardOne-3B commercially?

Yes. StandardOne-3B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

### What is StandardOne-3B's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Text generation

### [StandardOne-3B-FP8](https://savrn.com/models/standardone-3b-fp8)

[Standard Thinking](https://savrn.com/model-publishers/standardthinking)

StandardOne-3B-FP8 is an FP8 (compressed-tensors, float8e4m3 weights, dynamic per-token activations) quantization of the released StandardOne-3B decision model. The language-model linear projections (q/k/v/o, gate/up/down) are quantized per-channel FP8 E4M3 with dynamic FP8 activations (llm-compressor's data-free FP8DYNAMIC recipe, no calibration data required); the vision tower, multi-modal projector, embeddings and lmhead are left unquantized in BF16. It was produced from source revision 68dafd17ead9b8cf6f85c4f08f7f2f3a1e7b9e5c of StandardOne-3B on 2026-09-25 using llm-compressor 0.14.0 (torch 2.14.0, transformers 5.17.0, compressed-tensors 0.19.0); results below. Served through SGLang…

Open weights apache-2.0 3.8B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/standardone-3b-fp8)

Model · Text generation

### [phi4-clinical-mlx](https://savrn.com/models/phi4-clinical-mlx)

[Web](https://savrn.com/model-publishers/charakaweb)

A specialized 3.8B biomedical & clinical reasoning model built on Microsoft's Phi-4-mini-instruct, optimized natively for Apple Silicon Metal acceleration via Apple MLX. The model underwent a 3-stage transfer learning curriculum: 1. Stage 1 (STEM Foundation): 116,000 instruction pairs across NCERT Classes 6–12 (Physics, Chemistry, Biology) eliminating foundational science hallucinations. 2. Stage 2 (PubMed 2026 Evidence): 12 recent 2026 clinical update archives from NCBI FTP covering survival outcomes (OS, PFS, HR), targeted therapeutics, and clinical trial endpoints. 3. Stage 3 (Comprehensive Internal Medicine): Balanced multi-specialty clinical curriculum (cardiology, nephrology…

Open weights mit 3.8B parameters 131,072 tokens mlx

[View model](https://savrn.com/models/phi4-clinical-mlx)

Model · Text generation

### [phi4-clinical](https://savrn.com/models/phi4-clinical)

[Web](https://savrn.com/model-publishers/charakaweb)

A specialized 3.8B biomedical & clinical reasoning foundation model built on Microsoft's Phi-4-mini-instruct, formatted for standard Hugging Face transformers and PyTorch. 1. Stage 1 (STEM Foundation): 116,000 instruction pairs across NCERT Classes 6–12 (Physics, Chemistry, Biology) eliminating foundational science hallucinations. 2. Stage 2 (PubMed 2026 Evidence): 12 recent 2026 clinical update archives from NCBI FTP covering survival outcomes (OS, PFS, HR), targeted therapeutics, and clinical trial endpoints. 3. Stage 3 (Comprehensive Internal Medicine): Balanced multi-specialty clinical curriculum (cardiology, nephrology, endocrinology, pulmonology) with an active oncology replay buffer.…

Open weights mit 3.8B parameters 131,072 tokens transformers

[View model](https://savrn.com/models/phi4-clinical)

Model · Text generation

### [NULLXES-SHINRA-4B-SMOKE](https://savrn.com/models/nullxes-shinra-4b-smoke)

[Maga](https://savrn.com/model-publishers/magistrtheone)

NULLXES SHINRA-4B-INSTRUCT is the Language Intelligence Layer of the NULLXES system. SHINRA is responsible for multilingual understanding, coding intelligence, instruction following, structured outputs, and agent preparation. This checkpoint is the instruction-tuned (and optionally DPO-aligned) 4B-class dense decoder. Proprietary ShinraForCausalLM (not a Llama / Mistral / Qwen / GPT-NeoX wrapper). RMSNorm → GQA+RoPE → residual → RMSNorm → SwiGLU → residual then final RMSNorm and tied LM head. Special tokens:. Generation stop is. Document stop is. Three stages. Pretrain → NULLXES SHINRA-4B-BASE SHINRAPRETRAINV1: 40% FineWeb-Edu, 20% code (python-edu + licensed Stack), 15% math/science…

Open weights other 3.9B parameters 32,768 tokens transformers

[View model](https://savrn.com/models/nullxes-shinra-4b-smoke)

Model · Text generation

### [NULLXES-SHINRA-4B-INSTRUCT](https://savrn.com/models/nullxes-shinra-4b-instruct)

[Maga](https://savrn.com/model-publishers/magistrtheone)

Language Intelligence layer of the NULLXES Intelligence Stack. SHINRA Our llm. Release line 1. NULLXES SHINRA-4B-BASE — pretrain 2. NULLXES SHINRA-4B-INSTRUCT — instruction tuning ← this model 3. NULLXES SHINRA-4B-INSTRUCT (aligned) — DPO / preference optimization Parameter budget Architecture source: configs/shinra4b.yaml at 903a639e03bc. The total also matches the published safetensors metadata. Custom SentencePiece Unigram, trained in-house on a web + wiki + code mix. Special tokens trustremotecode=True is required: SHINRA ships a custom ShinraConfig and modeling code, not a reused architecture class. The HF inference widget cannot load custom code, hence inference: false. Data pipeline…

Open weights other 3.9B parameters 32,768 tokens transformers

[View model](https://savrn.com/models/nullxes-shinra-4b-instruct)

Model · Text generation

### [Wrench-4B-Qwen3.6-8E](https://savrn.com/models/wrench-4b-qwen3-6-8e)

[Stan Chen](https://savrn.com/model-publishers/stancsz)

Wrench is a pruned, task-specific developer-tool SLM derived from Qwen3.6-35B-A3B. It contains 3,881,244,016 parameters and stays below the 4.25B parameter ceiling. This is a public experimental artifact. It is downloadable and reproducible, but it is not a claim that the final 4M retrieval-quality or MiniMax-parity gates have passed. The canonical distribution format is Hugging Face Safetensors. The package embeds the tokenizer hook, deterministic mechanical lookup runtime, long-context overlay, hash-bound package metadata, and the FreeToken launcher. It is intended to feel like one model directory, not a separately installed harness. The bundled launcher requires a compatible FreeToken…

Open weights apache-2.0 3.9B parameters 2,000,000 tokens transformers

[View model](https://savrn.com/models/wrench-4b-qwen3-6-8e)

## Standard Thinking

[All models and datasets](https://savrn.com/model-publishers/standardthinking)

## Versions

- [65f9140a1294](https://savrn.com/models/standardone-3b/versions/65f9140a1294) · current 2026-09-27
- [68dafd17ead9](https://savrn.com/models/standardone-3b/versions/68dafd17ead9) 2026-09-25

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-09-27.
- [Hugging Face record](https://huggingface.co/StandardThinking/StandardOne-3B)
- [How the hub is built](https://savrn.com/model-hub/methodology)
