SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Darwin-180B-RSI

by FINAL_Bench FINAL-Bench/Darwin-180B-RSI

Darwin-180B-RSI is an open-weight model for image and text to text from FINAL_Bench, released under other. It has 180B parameters and a 262,144-token context. At 16-bit it needs about 432 GB of GPU memory, which fits on 2x MI325X from $4.00 an hour; at 4-bit, 108 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index. It draws 226 downloads a month.

flagship of the Darwin family — #1 on AIME 2026, HMMT Feb 2026, GPQA Diamond, MMLU-Pro, MMMU-Pro, LEXam and LEXam-hard, and a model that gets better by learning from its own verified work.

Parameters180B
Context262,144
Weights360.0 GB
Licenseother
AccessOpen weights
Monthly Downloads226

Runs On

What it takes to serve Darwin-180B-RSI (180B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 360.0 GB 432.0 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55
8-bit 180.0 GB 216.0 GB 1x MI325X (256 GB)
Vultr
$2.00 1x MI355X $2.59 · 2x MI300X $3.70
4-bit 90.0 GB 108.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Darwin-180B-RSI on every accelerator the SAVRN Index prices, at every precision

Model Card

flagship of the Darwin family — #1 on AIME 2026, HMMT Feb 2026, GPQA Diamond, MMLU-Pro, MMMU-Pro, LEXam and LEXam-hard, and a model that gets better by learning from its own verified work. Scores as listed on the Hugging Face official benchmark leaderboards (self-reported by each model's publisher). = #1 on that leaderboard. "—" = not reported. Darwin is VIDRAFT's measurement-driven reasoning model family — 50+ official models, 400+ community derivatives, and now two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3). Darwin treats a strong open model as a parent. It measures where the parent is weak, and strengthens exactly those parts — instead of re-training…

Excerpt from the card by FINAL_Bench, licensed other.

Configuration

Architecture
Qwen4ExpForConditionalGeneration
Context length (tokens)
262,144
Layers
48
Hidden size
2,560
Attention heads
24
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
512
Experts active per token
10
Model type
qwen4_exp

Identity and Version

Repository
FINAL-Bench/Darwin-180B-RSI
Publisher
FINAL_Bench
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
180B parameters
Languages
en, ko, zh, ja
Revision
bc3c7b0410b40c085b78084e13f01c12df31087b
First published
2026-09-28
Last updated
2026-09-30

Files and Weights

157 files, 360.0 GB in total. The weights are 132 files totalling 360.0 GB in npz, safetensors.

Weights132 files · 360.0 GB
Configuration14 files · 184.1 KB
Tokenizer4 files · 22.9 MB
Documentation3 files · 26.9 KB
Other3 files · 217.6 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00131.safetensorsWeights1.0 GB e5fbeecd1688
model-00002-of-00131.safetensorsWeights3.4 GB 1da92de7cf72
model-00003-of-00131.safetensorsWeights1.8 GB 33ecd23a8006
model-00004-of-00131.safetensorsWeights3.4 GB 18b756227729
model-00005-of-00131.safetensorsWeights3.4 GB a404db3a71e4
model-00006-of-00131.safetensorsWeights3.2 GB 8277a48e1fd7
model-00007-of-00131.safetensorsWeights3.2 GB 0fba369354e4
model-00008-of-00131.safetensorsWeights3.2 GB 024f59b6f2cb
model-00009-of-00131.safetensorsWeights3.2 GB 23fbef154f7f
model-00010-of-00131.safetensorsWeights3.2 GB b3a303529a19
model-00011-of-00131.safetensorsWeights3.2 GB d1215943b484
model-00012-of-00131.safetensorsWeights3.2 GB f63d649ea848
model-00013-of-00131.safetensorsWeights3.2 GB 5e0f0fb335f4
model-00014-of-00131.safetensorsWeights3.2 GB 5ee73bd21c73
model-00015-of-00131.safetensorsWeights3.2 GB b35421bf2172
model-00016-of-00131.safetensorsWeights3.2 GB 73c4b1bd1f37
model-00017-of-00131.safetensorsWeights3.2 GB ee60c1269974
model-00018-of-00131.safetensorsWeights3.2 GB 72cf3696aa1a
model-00019-of-00131.safetensorsWeights3.2 GB bdad649be407
model-00020-of-00131.safetensorsWeights3.2 GB 2258dad55787
model-00021-of-00131.safetensorsWeights3.2 GB 775952ac572f
model-00022-of-00131.safetensorsWeights3.2 GB 7d205070bb89
model-00023-of-00131.safetensorsWeights3.2 GB 598d4f772853
model-00024-of-00131.safetensorsWeights3.2 GB 09e534942e7d
model-00025-of-00131.safetensorsWeights3.2 GB 482d18f0e515
model-00026-of-00131.safetensorsWeights3.2 GB ac479fe29ed3
model-00027-of-00131.safetensorsWeights3.2 GB 4981d4e63672
model-00028-of-00131.safetensorsWeights3.2 GB 1b45a957b1f3
model-00029-of-00131.safetensorsWeights3.2 GB 40d2f2291ff4
model-00030-of-00131.safetensorsWeights3.2 GB 893a9b59e412
model-00031-of-00131.safetensorsWeights3.2 GB 9b6bf01d60f2
model-00032-of-00131.safetensorsWeights3.2 GB 7453fcb58721
model-00033-of-00131.safetensorsWeights3.2 GB 5b2f577ce409
model-00034-of-00131.safetensorsWeights3.2 GB 35b552266865
model-00035-of-00131.safetensorsWeights3.2 GB 4fe611218d36
model-00036-of-00131.safetensorsWeights3.2 GB 237789c0cef5
model-00037-of-00131.safetensorsWeights1.7 GB eccd977ce363
model-00038-of-00131.safetensorsWeights3.4 GB 063965d1d57b
model-00039-of-00131.safetensorsWeights1.7 GB 389e2701fa26
model-00040-of-00131.safetensorsWeights3.4 GB ed06bdece1c7
model-00041-of-00131.safetensorsWeights1.9 GB 8b4d057b91a1
model-00042-of-00131.safetensorsWeights3.4 GB 071bb12fa66a
model-00043-of-00131.safetensorsWeights1.8 GB 1828f06fc670
model-00044-of-00131.safetensorsWeights3.4 GB bab6f9fd62f1
model-00045-of-00131.safetensorsWeights1.8 GB 69715abe3df2
model-00046-of-00131.safetensorsWeights3.4 GB f838e842fb3f
model-00047-of-00131.safetensorsWeights1.7 GB 88f5f07b3339
model-00048-of-00131.safetensorsWeights3.4 GB dbbf3a6e2b81
model-00049-of-00131.safetensorsWeights1.9 GB 98fc84d028fe
model-00050-of-00131.safetensorsWeights3.4 GB 6014a2b71b74
model-00051-of-00131.safetensorsWeights1.8 GB 695a63f753c2
model-00052-of-00131.safetensorsWeights3.4 GB 1914762c6994
model-00053-of-00131.safetensorsWeights1.8 GB 8fa311b30c88
model-00054-of-00131.safetensorsWeights3.4 GB 22daa7622152
model-00055-of-00131.safetensorsWeights1.7 GB 6656e0822efe
model-00056-of-00131.safetensorsWeights3.4 GB 86ef0b5385fe
model-00057-of-00131.safetensorsWeights1.9 GB ee1f7ac39e5c
model-00058-of-00131.safetensorsWeights3.4 GB b7d2ab70d914
model-00059-of-00131.safetensorsWeights1.7 GB c6144faa339e
model-00060-of-00131.safetensorsWeights3.5 GB 45c7990b4307
model-00061-of-00131.safetensorsWeights3.4 GB cf5bff0a9663
model-00062-of-00131.safetensorsWeights3.5 GB d552485685ad
model-00063-of-00131.safetensorsWeights3.4 GB e35a86ae511b
model-00064-of-00131.safetensorsWeights1.8 GB b041fea4f308
model-00065-of-00131.safetensorsWeights3.4 GB 040b73058fe0
model-00066-of-00131.safetensorsWeights1.7 GB 7866bd0554c0
model-00067-of-00131.safetensorsWeights3.4 GB 9223da6ff56d
model-00068-of-00131.safetensorsWeights1.9 GB eef7805635dd
model-00069-of-00131.safetensorsWeights3.4 GB f7d460aa332a
model-00070-of-00131.safetensorsWeights1.8 GB cb2f85c44bfa
model-00071-of-00131.safetensorsWeights3.4 GB 184e042c9702
model-00072-of-00131.safetensorsWeights1.8 GB b876df1e1c78
model-00073-of-00131.safetensorsWeights3.4 GB 19de39abb6e1
model-00074-of-00131.safetensorsWeights1.7 GB a5f0f4a6b33d
model-00075-of-00131.safetensorsWeights3.4 GB fc43b07ed481
model-00076-of-00131.safetensorsWeights1.9 GB 4f0730fdc904
model-00077-of-00131.safetensorsWeights3.4 GB a75bf91dc472
model-00078-of-00131.safetensorsWeights1.8 GB 612636b5bdb6
model-00079-of-00131.safetensorsWeights3.4 GB e4cd07a12651
model-00080-of-00131.safetensorsWeights1.7 GB a08999e48c9f
model-00081-of-00131.safetensorsWeights3.4 GB 4cf05b67fcb8
model-00082-of-00131.safetensorsWeights1.9 GB 67215b93e90d
model-00083-of-00131.safetensorsWeights3.4 GB f097129297fd
model-00084-of-00131.safetensorsWeights1.7 GB b456694f42a4
model-00085-of-00131.safetensorsWeights3.4 GB 69083ff9730a
model-00086-of-00131.safetensorsWeights1.9 GB 28467f04e137
model-00087-of-00131.safetensorsWeights3.4 GB 9b9b8704d856
model-00088-of-00131.safetensorsWeights1.8 GB 5cea3ed9a6d5
model-00089-of-00131.safetensorsWeights3.4 GB 640ea88172a3
model-00090-of-00131.safetensorsWeights1.8 GB 67c46787ce53
model-00091-of-00131.safetensorsWeights3.4 GB 6ff4b9718910
model-00092-of-00131.safetensorsWeights1.7 GB 221ef7c514ce
model-00093-of-00131.safetensorsWeights3.4 GB 3d4f44c44752
model-00094-of-00131.safetensorsWeights1.9 GB 435a2ff39929
model-00095-of-00131.safetensorsWeights3.4 GB 9e680b695a5b
model-00096-of-00131.safetensorsWeights1.8 GB baa9b72239e8
model-00097-of-00131.safetensorsWeights3.4 GB 19be76593bce
model-00098-of-00131.safetensorsWeights1.8 GB 462ee17913ec
model-00099-of-00131.safetensorsWeights3.4 GB 21daceed016e
model-00100-of-00131.safetensorsWeights1.7 GB f0399d687021
model-00101-of-00131.safetensorsWeights3.4 GB 2a596073cf66
model-00102-of-00131.safetensorsWeights1.9 GB ed2eee2a9957
model-00103-of-00131.safetensorsWeights3.4 GB cc197d5c0d5a
model-00104-of-00131.safetensorsWeights1.8 GB 7f86abdad5d9
model-00105-of-00131.safetensorsWeights3.4 GB 4040abee30dc
model-00106-of-00131.safetensorsWeights1.9 GB 9d9d5f11aabd
model-00107-of-00131.safetensorsWeights3.4 GB 08fb87b54a96
model-00108-of-00131.safetensorsWeights1.9 GB 874378e17559
model-00109-of-00131.safetensorsWeights3.4 GB 67252c3ea4da
model-00110-of-00131.safetensorsWeights1.7 GB f2dc42581413
model-00111-of-00131.safetensorsWeights3.4 GB fe316b2426ea
model-00112-of-00131.safetensorsWeights1.9 GB 08a9de7fa0ce
model-00113-of-00131.safetensorsWeights3.4 GB ef41245305a2
model-00114-of-00131.safetensorsWeights1.8 GB 23f0a3f8b7d2
model-00115-of-00131.safetensorsWeights3.4 GB b536b5337a35
model-00116-of-00131.safetensorsWeights1.8 GB 7c3156fd3011
model-00117-of-00131.safetensorsWeights3.4 GB 27b091d8a571
model-00118-of-00131.safetensorsWeights1.7 GB 59a128ec6743
model-00119-of-00131.safetensorsWeights3.4 GB fae555008bac
model-00120-of-00131.safetensorsWeights1.9 GB d4e0e73be80c
model-00121-of-00131.safetensorsWeights3.4 GB e00e79982774
model-00122-of-00131.safetensorsWeights1.8 GB d9903575c6c2
model-00123-of-00131.safetensorsWeights3.4 GB bb8e5b70612d
model-00124-of-00131.safetensorsWeights1.7 GB 5df237c9a257
model-00125-of-00131.safetensorsWeights3.4 GB b91a643909b5
model-00126-of-00131.safetensorsWeights1.9 GB 2df33d913639
model-00127-of-00131.safetensorsWeights3.4 GB fa0c609906b1
model-00128-of-00131.safetensorsWeights1.8 GB d01708d70ff2
model-00129-of-00131.safetensorsWeights3.4 GB d24efe734057
model-00130-of-00131.safetensorsWeights3.0 GB ce20a503e4f0
model-00131-of-00131.safetensorsWeights1.3 GB 50be0ccd11e4
ztc/ztc_probe_darwin180rsi.npzWeights32.5 KB 3b520ad7c791
.eval_results/aime_2026.yamlConfiguration401 B —
.eval_results/gpqa_diamond.yamlConfiguration337 B —
.eval_results/hmmt_feb_2026.yamlConfiguration418 B —
.eval_results/lexam.yamlConfiguration396 B —
.eval_results/lexam_hard.yamlConfiguration596 B —
.eval_results/mmlu_pro.yamlConfiguration323 B —
.eval_results/mmmu_pro.yamlConfiguration349 B —
config.jsonConfiguration4.7 KB —
generation_config.jsonConfiguration202 B —
handler.pyConfiguration2.8 KB —
model.safetensors.index.jsonConfiguration170.7 KB —
preprocessor_config.jsonConfiguration390 B —
video_preprocessor_config.jsonConfiguration385 B —
ztc/usage.pyConfiguration2.0 KB —
LICENSEDocumentation3.2 KB —
README.mdDocumentation22.0 KB —
ztc/README.mdDocumentation1.6 KB —
assets/bench_field.pngOther137.3 KB 438ef769eb7e
assets/bench_vs_china.pngOther71.4 KB —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 0997f410c57a
tokenizer_config.jsonTokenizer17.9 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
other
Access
Open weights, no gate
Download size
360.0 GB
Download from FINAL_Bench

Released by FINAL_Bench through its official repository on Hugging Face.

Built From

  • Described by arXiv:2605.14386
  • Described by arXiv:2609.20269

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
AIME 2026 Task Competition MathematicsMetric Accuracy (majority vote, 16 samples, 131K thinking)Comparison conditions not established 100 FINAL-Bench
Publisher reported
Evaluated revision not stated —
AIME 2026 Task Competition MathematicsMetric Mean accuracy over 16 samplesComparison conditions not established 98.75 FINAL-Bench
Publisher reported
Evaluated revision not stated —
GPQA Diamond Configuration gpqa_diamondTask Graduate-Level ReasoningMetric Accuracy (majority vote, up to 16 samples, 131K thinking)Comparison conditions not established 94.44 FINAL-Bench
Publisher reported
Evaluated revision not stated —
HMMT Feb 2026 Task Competition MathematicsMetric Accuracy (majority vote, 16 samples, 131K thinking)Comparison conditions not established 100 FINAL-Bench
Publisher reported
Evaluated revision not stated —
HMMT Feb 2026 Task Competition MathematicsMetric Mean accuracy over 16 samplesComparison conditions not established 96.59 FINAL-Bench
Publisher reported
Evaluated revision not stated —
Idavidrein/gpqa Task diamondMetric diamondSetup GPQA Diamond, all 198 items; majority vote over up to 16 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16Comparison conditions not established 94.44 Model Card
Reported by a third party
Evaluated revision not stated 2026-09-28
LEXam (MCQ Configuration mcq_4_choicesTask Legal ReasoningMetric Accuracy (majority vote, 4 samples, 32K thinking)Comparison conditions not established 68.94 FINAL-Bench
Publisher reported
Evaluated revision not stated —
LEXam-Benchmark/LEXam Task mcq_4_choicesMetric mcq_4_choicesSetup LEXam mcq_4_choices, all 1,655 items; majority vote over 4 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 32,768 tokens; bf16. Single-sample 60.54, mean of 4 samples 61.42Comparison conditions not established 68.94 Model Card
Reported by a third party
Evaluated revision not stated 2026-09-29
LEXam-hard Task Legal Reasoning (open-ended)Metric Judge score (DeepSeek-R1-0528, single sample)Comparison conditions not established 45.72 FINAL-Bench
Publisher reported
Evaluated revision not stated —
MMLU-Pro Task Multi-discipline Knowledge & ReasoningMetric Accuracy (single sample, 131K thinking)Comparison conditions not established 88.12 FINAL-Bench
Publisher reported
Evaluated revision not stated —
MMMU-Pro (vision) Configuration visionTask Multimodal Expert ReasoningMetric Accuracy (majority vote, 3 samples, 131K thinking)Comparison conditions not established 79.48 FINAL-Bench
Publisher reported
Evaluated revision not stated —
MMMU/MMMU_Pro Task mmmu_pro_visionMetric mmmu_pro_visionSetup MMMU-Pro vision setting, all 1,730 items; majority vote over 3 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16Comparison conditions not established 79.48 Model Card
Reported by a third party
Evaluated revision not stated 2026-09-28
MathArena/aime_2026 Task MathArena/aime_2026Metric MathArena/aime_2026Setup AIME 2026, all 30 problems; majority vote over 16 samples (maj@16) = 100.0; mean accuracy over 16 samples = 98.75; temperature=1.0, top_p=0.95, top_k=20; thinking budget 131,072 tokens; bf16Comparison conditions not established 100 Model Card
Reported by a third party
Evaluated revision not stated 2026-09-28
MathArena/hmmt_feb_2026 Task MathArena/hmmt_feb_2026Metric MathArena/hmmt_feb_2026Setup HMMT February 2026, all 33 problems; majority vote over 16 samples (maj@16) = 100.0; mean accuracy over 16 samples = 96.59; temperature=1.0, top_p=0.95, top_k=20; thinking budget 131,072 tokens; bf16Comparison conditions not established 100 Model Card
Reported by a third party
Evaluated revision not stated 2026-09-28
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proSetup MMLU-Pro test, all 12,032 items; single sample (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16Comparison conditions not established 88.12 Model Card
Reported by a third party
Evaluated revision not stated 2026-09-28
joelniklaus/LEXam-hard Task lexam_hardMetric lexam_hardSetup LEXam-hard, all 518 open questions; single sample (temperature=1.0, top_p=0.95, top_k=20); thinking budget 32,768 tokens, responses truncated at 32K regenerated with a 120K budget (60 items); judged by DeepSeek-R1-0528 with the LEXam paper judge prompt (eval.yaml); score = mean of German and…Comparison conditions not established 45.72 Model Card
Reported by a third party
Evaluated revision not stated 2026-09-30

Memory Requirements

PrecisionWeights in memory
As published360.0 GB
16-bit360.0 GB
8-bit180.0 GB
4-bit90.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Darwin-180B-RSI

How much GPU memory does Darwin-180B-RSI need?

About 432 GB at 16-bit and 108 GB at 4-bit: the weights (180B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Darwin-180B-RSI on?

At 16-bit, 2x MI325X from $4.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Darwin-180B-RSI released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Darwin-180B-RSI's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-Flash-Next

Qwen

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next-Uncensored-NVFP4

OrcaRouter

This model has had its safety alignment substantially removed via abliteration (orthogonalizing the refusal direction out of the residual stream). It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment. Use must comply with the Apache 2.0 License inherited from the base model and all applicable law. The authors accept no liability for misuse. - A Blackwell GPU…

Access requested at publisher apache-2.0 180B parameters transformers

Model · Image and text to text

Qwen3.8-Flash-Next-MLX-oQ3-MTP

Robot Haus

A sensitivity-guided, mixed-precision MLX conversion of Qwen/Qwen3.8-Flash-Next, rebuilt directly from the official BF16 checkpoint with the model's matching native MTP block preserved. oQ3 uses a 3-bit affine base and spends additional precision on sensitive modules. Layer sensitivity was measured with a validated quantized calibration proxy, while every released weight was quantized from the official BF16 checkpoint. The result is a compact model with 746 higher-precision module overrides rather than a uniform 3-bit layout. The upstream tokenizer, current chat template, vision processor, generation configuration, licence, and native MTP configuration are retained. In a compatible oMLX…

Open weights other 180B parameters 262,144 tokens mlx

Model · Image and text to text

Qwen3.8-Flash-Next

Tai Hua

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next

Ahmed

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

clover-1-150b-preview

Kitani

libraryname: transformers pipelinetag: image-text-to-text basemodel: - Qwen/Qwen3.8-27B - clover - mixture-of-experts - multimodal - reasoning - open-weights - custom-code Clover 1 150B Preview is the first public checkpoint of Clover 1, an experimental MoE model we're developing at Kitani. It's based on Qwen3.8-27B, but it isn't just a finetune with a different name. We kept much of Qwen3.8's underlying architecture, including the tokenizer, multimodal components, hybrid language backbone, embeddings, and LM head, while replacing the language model's dense FFNs with our own sparse mixture-of-experts setup. Clover currently has 8 full-sized experts per language layer with top-1 routing. The…

Open weights apache-2.0 147.6B parameters 262,144 tokens transformers