Prism ML's ternary Ternary-Bonsai-2-27B build of Qwen/Qwen3.8-27B, repacked for chad, a Claude-Code-style local coding agent for Apple Silicon, with its speculative decoder bundled in. This is chad's default model. Created using Bonsai by Prism ML. with, already quantized. Nothing is built on first run. Every projection of Qwen3.8-27B (a dense qwen35 hybrid: 64 layers, 48 GatedDeltaNet + 16 full attention) is stored in a Hadamard-rotated basis: multiplied by a fixed sign vector and put through a blockwise Walsh-Hadamard transform offline, then quantized to 2-bit affine group-128 whose three levels reproduce the ternary set {−s, 0, +s}. The rotation costs no extra bits and no extra weight…
Darwin-27B-RSI is an open-weight model for text generation from FINAL_Bench, released under Apache License 2.0. It has 26.9B parameters and a 262,144-token context. At 16-bit it needs about 64.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 459 downloads a month.
Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions.
Runs On
What it takes to serve Darwin-27B-RSI (26.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 53.8 GB | 64.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 26.9 GB | 32.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 13.4 GB | 16.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.
Darwin-27B-RSI on every accelerator the SAVRN Index prices, at every precision
Model Card
By FINAL_Bench, published under apache-2.0, revision 08d44a01bbe4.
Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution). Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests. As the reasoning engine of Darwin-27B-JEV on…
Read FINAL_Bench's full model card
Darwin-27B-RSI: A Model That Improved Itself — Learning Only From Its Own Solutions
Qwen3.5-27B family · 27B dense · Thinking mode · BF16 · Apache 2.0 No human-written answers. The model generated its own learning signal — and got measurably better.
Abstract
Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution).
Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests.
As the reasoning engine of Darwin-27B-JEV on the Decision Index, it lifts the hardest reasoning decisions: GPQA Diamond skill 0.31 → 0.71, GSM8K 0.61 → 0.97, MMLU-Pro 0.60 → 0.82.
Model-level RSI vs. harness-level RSI
Darwin-27B-RSI is Model-level RSI: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. You download a new model file, and it is smarter on its own.
Harness-level RSI (for example, Google's RRSI) improves the prompts, tools and workflow around a fixed model — like rewriting an employee's manual, while Model-level RSI is the employee getting smarter. The two are complementary: a harness-level loop can run on top of a Model-level RSI model.
What Is RSI?
Most models improve only when people write more answers for them. Recursive Self-Improvement removes that bottleneck: the model works on problems, judges its own work, and learns from what it produced — then repeats. Each improved model becomes the starting point for the next improvement.
Darwin-27B-RSI demonstrates this loop on a 27B model:
- Human-written solutions or reasoning traces used: 0
- Direction of change: measurably better on held-out graduate-level science
- Contamination check: training problems share 0 items with the evaluation sets reported here
The training procedure itself is not released.
Results
Science reasoning (same protocol for both models)
| Benchmark | Darwin-27B-Opus | Darwin-27B-RSI | Δ |
|---|---|---|---|
| GPQA Diamond (1 sample) | 72.85 | 78.09 | +5.24 |
| GPQA Diamond (majority@16) | 79.80 | 83.59 | +3.79 |
| SuperGPQA (1 sample) | +4.03 |
Both models were measured under the same protocol (single sample, identical sampling settings and token budget), so numbers differ from the Darwin-27B-Opus card, which reports a different protocol. All gains are statistically significant in paired tests.
Decision Index — as the reasoning engine of Darwin-27B-JEV
The Decision Index scores typed-decision engines on 43 benchmarks and ~121K decisions (chance-corrected: 0 = random, 1 = perfect). Darwin-27B-RSI handles the decisions that need real thinking:
| Benchmark (skill) | before | with Darwin-27B-RSI |
|---|---|---|
| GPQA Diamond | 0.31 | 0.71 |
| GSM8K | 0.61 | 0.97 |
| CRUXEval | 0.61 | 0.87 |
| MMLU-Pro | 0.60 | 0.82 |
| BBH | 0.68 | 0.83 |
| CLadder | 0.49 | 0.70 |
= gold benchmark (weighted 1.2× on the board). Darwin-27B-JEV: ≈ 61.1 under the v0.2.1 board rules (our recomputation; official score pending review). Full run: FINAL-Bench/Darwin-27B-JEV-decision-index.
Usage
Darwin-27B-RSI is a thinking model. Give it room to reason and read the answer after the reasoning block.
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "FINAL-Bench/Darwin-27B-RSI"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto")
messages = [{"role": "user", "content": "A ball is thrown upward at 40 m/s. For how long is it above 40 m? (g = 10 m/s²) Think, then give the final answer."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=8192, temperature=0.6, top_p=0.95, do_sample=True)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
vLLM
vllm serve FINAL-Bench/Darwin-27B-RSI --max-model-len 32768
Recommended: temperature 0.6, top_p 0.95, generous token budget (8K–16K) for hard problems.
Model Details
| Parent | FINAL-Bench/Darwin-27B-Opus |
| Architecture | Qwen3.5 family, 27B dense |
| Precision | BF16 |
| Improvement method | Recursive Self-Improvement, no human labels |
| License | Apache 2.0 |
| Developer | VIDRAFT · FINAL-Bench |
Limitations and Disclosure
- Gains were measured on graduate-level science; other domains may change less.
- As a thinking model, it can produce long reasoning; cap
max_new_tokensfor latency-sensitive use. - 31 training problems (0.22% of the benchmark) overlap with the Decision Index MMLU set; the model learned only from its own solutions to them.
- Not affiliated with TypeSafe AI or its Jev product.
Citation
@misc{darwin27b_rsi_2026,
title = {Darwin-27B-RSI: Recursive Self-Improvement without Human Labels},
author = {VIDRAFT and FINAL-Bench},
year = {2026},
url = {https://huggingface.co/FINAL-Bench/Darwin-27B-RSI}
}
Configuration
- Architecture
- Qwen3_5ForCausalLM
- Context length (tokens)
- 262,144
- Layers
- 64
- Hidden size
- 5,120
- Feed-forward size
- 17,408
- Attention heads
- 24
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Model type
- qwen3_5_text
Identity and Version
- Repository
- FINAL-Bench/Darwin-27B-RSI
- Publisher
- FINAL_Bench
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 26.9B parameters
- Languages
- en, ko
- Revision
- 08d44a01bbe4762fbc95171c952f0536bd5d46c6
- First published
- 2026-09-27
- Last updated
- 2026-09-30
Files and Weights
12 files, 53.8 GB in total. The weights are 2 files totalling 53.8 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00002.safetensors | Weights | 49.8 GB | 7a2da2e8cf5c |
| model-00002-of-00002.safetensors | Weights | 4.0 GB | 81917aca054b |
| config.json | Configuration | 2.7 KB | — |
| generation_config.json | Configuration | 244 B | — |
| model.safetensors.index.json | Configuration | 83.9 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| README.md | Documentation | 8.4 KB | — |
| chat_template.jinja | Other | 7.8 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 06b9509352d2 |
| tokenizer_config.json | Tokenizer | 1.1 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 53.8 GB
Released by FINAL_Bench through its official repository on Hugging Face. Read the license.
Built From
- Derived from FINAL-Bench/Darwin-27B-Opus
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 53.8 GB |
| 16-bit | 53.8 GB |
| 8-bit | 26.9 GB |
| 4-bit | 13.4 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Built on This Model
- Quantized fromDarwin-27B-RSI-i1-GGUF
- Derived fromDarwin-27B-RSI-i1-GGUF
- Quantized fromFINAL-Bench_Darwin-27B-RSI-GGUF
- Derived fromFINAL-Bench_Darwin-27B-RSI-GGUF
Questions About Darwin-27B-RSI
How much GPU memory does Darwin-27B-RSI need?
About 64.6 GB at 16-bit and 16.1 GB at 4-bit: the weights (26.9B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Darwin-27B-RSI on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Darwin-27B-RSI commercially?
Yes. Darwin-27B-RSI is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is Darwin-27B-RSI's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via JiRackDeltaNetTokenizer - JiRack DeltaNet understand video and images that best for Robotics also A fast and efficient 27B model optimized for CPU inference. Built on a Qwen3.8-style DeltaNet architecture (hybrid attention + SSM), with an updated tokenizer that includes Routing, Media, Vision, Sound, Tool call, and Robotics tags. Ready-to-run GGUF quantizations, and native Ollama support with reasoning disabled by default for fast, direct responses. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert…
Full 27B-class reasoning in ternary transformer weights — on everyday laptops - \~7.2 GB deployed footprint (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 95% of FP16 intelligence retained: 80.49 average across 15 thinking-mode benchmarks — a higher score than the conventional IQ2XXS build (72.73) at less than two-thirds of its footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01 - End-to-end ternary language weights across embeddings, attention projections, MLP…
V3 applies iterative refinement on top of V2's complementary blend, with targeted corpus expansion. The result: genuine liberation — not just removal of hard refusals but elimination of safety-lecture deflections. - Genuinely answers restricted queries — provides real substance instead of safety lectures - 20/20 on code generation tasks — functional implementations, not disclaimers - Thinking ON compatible — no refusals in either thinking mode - Honest scoring — every response manually audited for real substance, not just absence of "I cannot" - -2.1pp MMLU — modest capability cost for genuine liberation If you're using this model in an agent harness (coding agent, pentest framework, etc.)…
18+ only. This model generates sexually explicit fiction. A fine-tune of Qwen/Qwen3.8-27B for Chinese adult creative writing. It was trained with LoRA on about 4.2k instruction examples built from Chinese adult fiction, and the LoRA was then merged into the base weights, so this repo loads like a normal full model. Requires a recent transformers with Qwen3.8 support. The ~50 GB of bf16 weights need multiple GPUs or CPU offload. Avoid greedy decoding, which makes Qwen models more likely to repeat themselves. - Content is fictional and intended for adult readers only. - Trained mostly on long-form prose; instruction following outside creative writing may be weaker than the base model. - May…
This repository contains an MXFP8-weight checkpoint derived from exported in compressed-tensors format. The checkpoint retains the source multimodal components, but the evaluation reported here covers text tasks only. The exported checkpoint contains 11,635 F8E4M3 weight tensors, 838 BF16 weight tensors, 11,635 U8 block-scale tensors, and 171 FP32 scale/metadata tensors. Routed-expert weights and source-present text self-attention projections are MXFP8; the router, shared MLP, vision tower, and embeddings remain BF16. Measured on 2026-09-30 with lm-eval 0.4.13 and vLLM 0.29.0. The tested settings were TRITONATTN, tensor parallelism 2, pipeline parallelism 1, batch size 64, maxnumseqs=64…