SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Darwin-27B-RSI

by FINAL_Bench FINAL-Bench/Darwin-27B-RSI

Darwin-27B-RSI is an open-weight model for text generation from FINAL_Bench, released under Apache License 2.0. It has 26.9B parameters and a 262,144-token context. At 16-bit it needs about 64.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 459 downloads a month.

Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions.

Parameters26.9B
Context262,144
Weights53.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads459

Runs On

What it takes to serve Darwin-27B-RSI (26.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 53.8 GB 64.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 26.9 GB 32.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.4 GB 16.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Darwin-27B-RSI on every accelerator the SAVRN Index prices, at every precision

Model Card

By FINAL_Bench, published under apache-2.0, revision 08d44a01bbe4.

Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution). Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests. As the reasoning engine of Darwin-27B-JEV on…

Read FINAL_Bench's full model card

Darwin-27B-RSI: A Model That Improved Itself — Learning Only From Its Own Solutions

Qwen3.5-27B family · 27B dense · Thinking mode · BF16 · Apache 2.0 No human-written answers. The model generated its own learning signal — and got measurably better.


Abstract

Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution).

Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests.

As the reasoning engine of Darwin-27B-JEV on the Decision Index, it lifts the hardest reasoning decisions: GPQA Diamond skill 0.31 → 0.71, GSM8K 0.61 → 0.97, MMLU-Pro 0.60 → 0.82.


Model-level RSI vs. harness-level RSI

Darwin-27B-RSI is Model-level RSI: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. You download a new model file, and it is smarter on its own.

Harness-level RSI (for example, Google's RRSI) improves the prompts, tools and workflow around a fixed model — like rewriting an employee's manual, while Model-level RSI is the employee getting smarter. The two are complementary: a harness-level loop can run on top of a Model-level RSI model.

What Is RSI?

Most models improve only when people write more answers for them. Recursive Self-Improvement removes that bottleneck: the model works on problems, judges its own work, and learns from what it produced — then repeats. Each improved model becomes the starting point for the next improvement.

Darwin-27B-RSI demonstrates this loop on a 27B model:

  • Human-written solutions or reasoning traces used: 0
  • Direction of change: measurably better on held-out graduate-level science
  • Contamination check: training problems share 0 items with the evaluation sets reported here

The training procedure itself is not released.


Results

Science reasoning (same protocol for both models)

Benchmark Darwin-27B-Opus Darwin-27B-RSI Δ
GPQA Diamond (1 sample) 72.85 78.09 +5.24
GPQA Diamond (majority@16) 79.80 83.59 +3.79
SuperGPQA (1 sample) +4.03

Both models were measured under the same protocol (single sample, identical sampling settings and token budget), so numbers differ from the Darwin-27B-Opus card, which reports a different protocol. All gains are statistically significant in paired tests.

Decision Index — as the reasoning engine of Darwin-27B-JEV

The Decision Index scores typed-decision engines on 43 benchmarks and ~121K decisions (chance-corrected: 0 = random, 1 = perfect). Darwin-27B-RSI handles the decisions that need real thinking:

Benchmark (skill) before with Darwin-27B-RSI
GPQA Diamond 0.31 0.71
GSM8K 0.61 0.97
CRUXEval 0.61 0.87
MMLU-Pro 0.60 0.82
BBH 0.68 0.83
CLadder 0.49 0.70

= gold benchmark (weighted 1.2× on the board). Darwin-27B-JEV: ≈ 61.1 under the v0.2.1 board rules (our recomputation; official score pending review). Full run: FINAL-Bench/Darwin-27B-JEV-decision-index.


Usage

Darwin-27B-RSI is a thinking model. Give it room to reason and read the answer after the reasoning block.

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "FINAL-Bench/Darwin-27B-RSI"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto")

messages = [{"role": "user", "content": "A ball is thrown upward at 40 m/s. For how long is it above 40 m? (g = 10 m/s²) Think, then give the final answer."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=8192, temperature=0.6, top_p=0.95, do_sample=True)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))

vLLM

vllm serve FINAL-Bench/Darwin-27B-RSI --max-model-len 32768

Recommended: temperature 0.6, top_p 0.95, generous token budget (8K–16K) for hard problems.


Model Details

Parent FINAL-Bench/Darwin-27B-Opus
Architecture Qwen3.5 family, 27B dense
Precision BF16
Improvement method Recursive Self-Improvement, no human labels
License Apache 2.0
Developer VIDRAFT · FINAL-Bench

Limitations and Disclosure

  • Gains were measured on graduate-level science; other domains may change less.
  • As a thinking model, it can produce long reasoning; cap max_new_tokens for latency-sensitive use.
  • 31 training problems (0.22% of the benchmark) overlap with the Decision Index MMLU set; the model learned only from its own solutions to them.
  • Not affiliated with TypeSafe AI or its Jev product.

Citation

@misc{darwin27b_rsi_2026,
  title  = {Darwin-27B-RSI: Recursive Self-Improvement without Human Labels},
  author = {VIDRAFT and FINAL-Bench},
  year   = {2026},
  url    = {https://huggingface.co/FINAL-Bench/Darwin-27B-RSI}
}

Configuration

Architecture
Qwen3_5ForCausalLM
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5_text

Identity and Version

Repository
FINAL-Bench/Darwin-27B-RSI
Publisher
FINAL_Bench
Task
Text generation
Modality
Text
Library
transformers
Parameters
26.9B parameters
Languages
en, ko
Revision
08d44a01bbe4762fbc95171c952f0536bd5d46c6
First published
2026-09-27
Last updated
2026-09-30

Files and Weights

12 files, 53.8 GB in total. The weights are 2 files totalling 53.8 GB in safetensors.

Weights2 files · 53.8 GB
Configuration5 files · 87.7 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 8.4 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights49.8 GB 7a2da2e8cf5c
model-00002-of-00002.safetensorsWeights4.0 GB 81917aca054b
config.jsonConfiguration2.7 KB —
generation_config.jsonConfiguration244 B —
model.safetensors.index.jsonConfiguration83.9 KB —
preprocessor_config.jsonConfiguration390 B —
video_preprocessor_config.jsonConfiguration385 B —
README.mdDocumentation8.4 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
53.8 GB
Download from FINAL_Bench

Released by FINAL_Bench through its official repository on Hugging Face. Read the license.

Built From

  • Derived from FINAL-Bench/Darwin-27B-Opus

Memory Requirements

PrecisionWeights in memory
As published53.8 GB
16-bit53.8 GB
8-bit26.9 GB
4-bit13.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About Darwin-27B-RSI

How much GPU memory does Darwin-27B-RSI need?

About 64.6 GB at 16-bit and 16.1 GB at 4-bit: the weights (26.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Darwin-27B-RSI on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Darwin-27B-RSI commercially?

Yes. Darwin-27B-RSI is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Darwin-27B-RSI's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Prism ML's ternary Ternary-Bonsai-2-27B build of Qwen/Qwen3.8-27B, repacked for chad, a Claude-Code-style local coding agent for Apple Silicon, with its speculative decoder bundled in. This is chad's default model. Created using Bonsai by Prism ML. with, already quantized. Nothing is built on first run. Every projection of Qwen3.8-27B (a dense qwen35 hybrid: 64 layers, 48 GatedDeltaNet + 16 full attention) is stored in a Hadamard-rotated basis: multiplied by a fixed sign vector and put through a blockwise Walsh-Hadamard transform offline, then quantized to 2-bit affine group-128 whose three levels reproduce the ternary set {−s, 0, +s}. The rotation costs no extra bits and no extra weight…

Open weights apache-2.0 26.9B parameters 262,144 tokens mlx

Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via JiRackDeltaNetTokenizer - JiRack DeltaNet understand video and images that best for Robotics also A fast and efficient 27B model optimized for CPU inference. Built on a Qwen3.8-style DeltaNet architecture (hybrid attention + SSM), with an updated tokenizer that includes Routing, Media, Vision, Sound, Tool call, and Robotics tags. Ready-to-run GGUF quantizations, and native Ollama support with reasoning disabled by default for fast, direct responses. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert…

Open weights mit 27.3B parameters 262,144 tokens

Model · Text generation

Ternary-Bonsai-27B-mlx-2bit

Prism ML

Full 27B-class reasoning in ternary transformer weights — on everyday laptops - \~7.2 GB deployed footprint (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 95% of FP16 intelligence retained: 80.49 average across 15 thinking-mode benchmarks — a higher score than the conventional IQ2XXS build (72.73) at less than two-thirds of its footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01 - End-to-end ternary language weights across embeddings, attention projections, MLP…

Open weights apache-2.0 27.4B parameters 262,144 tokens mlx

Model · Text generation

Qwen3.8-27B-OBLITERATED

OBLITERATUS

V3 applies iterative refinement on top of V2's complementary blend, with targeted corpus expansion. The result: genuine liberation — not just removal of hard refusals but elimination of safety-lecture deflections. - Genuinely answers restricted queries — provides real substance instead of safety lectures - 20/20 on code generation tasks — functional implementations, not disclaimers - Thinking ON compatible — no refusals in either thinking mode - Honest scoring — every response manually audited for real substance, not just absence of "I cannot" - -2.1pp MMLU — modest capability cost for genuine liberation If you're using this model in an agent harness (coding agent, pentest framework, etc.)…

Open weights apache-2.0 27.8B parameters 262,144 tokens mlx

18+ only. This model generates sexually explicit fiction. A fine-tune of Qwen/Qwen3.8-27B for Chinese adult creative writing. It was trained with LoRA on about 4.2k instruction examples built from Chinese adult fiction, and the LoRA was then merged into the base weights, so this repo loads like a normal full model. Requires a recent transformers with Qwen3.8 support. The ~50 GB of bf16 weights need multiple GPUs or CPU offload. Avoid greedy decoding, which makes Qwen models more likely to repeat themselves. - Content is fictional and intended for adult readers only. - Trained mostly on long-form prose; instruction following outside creative writing may be weaker than the base model. - May…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

This repository contains an MXFP8-weight checkpoint derived from exported in compressed-tensors format. The checkpoint retains the source multimodal components, but the evaluation reported here covers text tasks only. The exported checkpoint contains 11,635 F8E4M3 weight tensors, 838 BF16 weight tensors, 11,635 U8 block-scale tensors, and 171 FP32 scale/metadata tensors. Routed-expert weights and source-present text self-attention projections are MXFP8; the router, shared MLP, vision tower, and embeddings remain BF16. Measured on 2026-09-30 with lm-eval 0.4.13 and vLLM 0.29.0. The tested settings were TRITONATTN, tensor parallelism 2, pipeline parallelism 1, batch size 64, maxnumseqs=64…

Open weights apache-2.0 25.8B parameters 262,144 tokens transformers