SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Kimi-K3

by Moonshot AI moonshotai/Kimi-K3

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date.

Parameters2.8T
Context1,048,576
Weights1.6 TB
Licenseother
AccessOpen weights
Monthly Downloads2.2M

Runs On

What it takes to serve Kimi-K3 (2.8T parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 5559.9 GB 6671.8 GB More than one server of any accelerator the SAVRN Index prices.
8-bit 2779.9 GB 3335.9 GB More than one server of any accelerator the SAVRN Index prices.
4-bit 1390.0 GB 1668.0 GB 7x MI325X (256 GB)
Vultr
$14.00 6x MI355X $15.54 · 7x B300 $46.20

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Kimi-K3

Nobody runs this one on a single box. At 16-bit the weights are 5,559.9 GB with 6,671.8 GB needed, more than one server of anything we price, and 8-bit still spills over. 4-bit is the first fit: 1,390.0 GB of weights, 1,668.0 GB needed, on seven MI325X cards with 256 GB each at $14.00 per hour. That buys Moonshot AI's 2.8T-parameter model, image and text to text, with a 1,048,576-token context window built for long-horizon coding, knowledge work and reasoning.

The license field reads other, with no summary in our file, so we cannot say whether commercial use is permitted; read the publisher's text before planning a deployment. Then weigh the rent. Baseten, Fireworks and Together AI serve it at $3.00 per million input tokens and $15.00 per million output, DeepInfra at $2.85 and $14.25, and that bill is what your seven-card cluster has to beat at your volume.

Model Card

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning. All Kimi K3 results are obtained with reasoning effort set to 'max' and temperature = 1.0. For single-step tasks, such as GPQA Diamond, HLE-Full, and vision benchmarks without tools, we set top-p = 0.95; for agentic tasks, we set top-p = 1.0. For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision…

Excerpt from the card by Moonshot AI, licensed other.

Configuration

Architecture
KimiK3ForConditionalGeneration
Context length (tokens)
1,048,576
Layers
93
Hidden size
7,168
Feed-forward size
33,792
Attention heads
96
Key/value heads
96
Vocabulary size
163,840
Experts
896
Model type
kimi_k3

Identity and Version

Repository
moonshotai/Kimi-K3
Publisher
Moonshot AI
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
2.8T parameters
Languages
Not stated by the source
Revision
f831ab66814297da540d832a5235f8e904f29d06
First published
2026-06-13
Last updated
2026-09-02

Files and Weights

119 files, 1.6 TB in total. The weights are 96 files totalling 1.6 TB in safetensors.

Weights96 files · 1.6 TB
Configuration17 files · 60.0 MB
Tokenizer2 files · 2.8 MB
Documentation2 files · 48.3 KB
Other1 file · 88.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-000096.safetensorsWeights2.3 GB 975584c00f85
model-00002-of-000096.safetensorsWeights17.0 GB 26a3284e1d2c
model-00003-of-000096.safetensorsWeights17.0 GB e54af9de4c55
model-00004-of-000096.safetensorsWeights16.6 GB 5955fd8feda8
model-00005-of-000096.safetensorsWeights17.0 GB d60d68ad0381
model-00006-of-000096.safetensorsWeights17.0 GB b1d480576747
model-00007-of-000096.safetensorsWeights17.0 GB fb1120fef34c
model-00008-of-000096.safetensorsWeights16.6 GB 2318dda54fc1
model-00009-of-000096.safetensorsWeights17.0 GB 8b66cdde34f5
model-00010-of-000096.safetensorsWeights17.0 GB d34f55f7b734
model-00011-of-000096.safetensorsWeights17.0 GB 1738856a4cca
model-00012-of-000096.safetensorsWeights16.6 GB c6b9bef38415
model-00013-of-000096.safetensorsWeights17.0 GB 3cbf43d56d9c
model-00014-of-000096.safetensorsWeights17.0 GB ce5f343f07d4
model-00015-of-000096.safetensorsWeights17.0 GB b55caa801334
model-00016-of-000096.safetensorsWeights16.6 GB 5a63e63ced65
model-00017-of-000096.safetensorsWeights17.0 GB 622bfa605205
model-00018-of-000096.safetensorsWeights17.0 GB 838b265ebb86
model-00019-of-000096.safetensorsWeights17.0 GB 599a8ecf88f2
model-00020-of-000096.safetensorsWeights16.6 GB 8e01b61ab766
model-00021-of-000096.safetensorsWeights17.0 GB 1944265bd024
model-00022-of-000096.safetensorsWeights17.0 GB 2d32d3e3c8da
model-00023-of-000096.safetensorsWeights17.0 GB c291f2ec1578
model-00024-of-000096.safetensorsWeights16.6 GB 278855ab81d4
model-00025-of-000096.safetensorsWeights17.0 GB 6e3ae6ad868f
model-00026-of-000096.safetensorsWeights17.0 GB cdd79fb52c7a
model-00027-of-000096.safetensorsWeights17.0 GB 1d974b40c4ae
model-00028-of-000096.safetensorsWeights16.6 GB 1bec58e89ab7
model-00029-of-000096.safetensorsWeights17.0 GB c3bf0e738aa5
model-00030-of-000096.safetensorsWeights17.0 GB 7b996410482a
model-00031-of-000096.safetensorsWeights17.0 GB c26689540a24
model-00032-of-000096.safetensorsWeights16.6 GB f7dc9e726d46
model-00033-of-000096.safetensorsWeights17.0 GB 615afa33b69c
model-00034-of-000096.safetensorsWeights17.0 GB a53b27fe92df
model-00035-of-000096.safetensorsWeights17.0 GB 9f4b44c89e49
model-00036-of-000096.safetensorsWeights16.6 GB 55aef33fab36
model-00037-of-000096.safetensorsWeights17.0 GB e95dd3599d69
model-00038-of-000096.safetensorsWeights17.0 GB 1ce472771309
model-00039-of-000096.safetensorsWeights17.0 GB 9c75b18c0d3a
model-00040-of-000096.safetensorsWeights16.6 GB 6f664598c00a
model-00041-of-000096.safetensorsWeights17.0 GB 375f05f94a59
model-00042-of-000096.safetensorsWeights17.0 GB b65947611d6c
model-00043-of-000096.safetensorsWeights17.0 GB b5a425f100bb
model-00044-of-000096.safetensorsWeights16.6 GB 113ee0120442
model-00045-of-000096.safetensorsWeights17.0 GB eb0698659da5
model-00046-of-000096.safetensorsWeights17.0 GB 0270727a399c
model-00047-of-000096.safetensorsWeights17.0 GB b38f63eb0803
model-00048-of-000096.safetensorsWeights16.6 GB 131e243c02cf
model-00049-of-000096.safetensorsWeights17.0 GB 72c91dcf2909
model-00050-of-000096.safetensorsWeights17.0 GB 0a627a082cd3
model-00051-of-000096.safetensorsWeights17.0 GB 38a37bcec20a
model-00052-of-000096.safetensorsWeights16.6 GB 9703b6321741
model-00053-of-000096.safetensorsWeights17.0 GB 413ec9ea0b69
model-00054-of-000096.safetensorsWeights17.0 GB 0e3da201b765
model-00055-of-000096.safetensorsWeights17.0 GB e7c9f4e44f8a
model-00056-of-000096.safetensorsWeights16.6 GB efd6176016b9
model-00057-of-000096.safetensorsWeights17.0 GB e6982c96ac91
model-00058-of-000096.safetensorsWeights17.0 GB c93a23cbd653
model-00059-of-000096.safetensorsWeights17.0 GB a62eb8220710
model-00060-of-000096.safetensorsWeights16.6 GB 9bccbaa71b98
model-00061-of-000096.safetensorsWeights17.0 GB 1ae3969540fc
model-00062-of-000096.safetensorsWeights17.0 GB 96babdc24f22
model-00063-of-000096.safetensorsWeights17.0 GB 8a81caa697a7
model-00064-of-000096.safetensorsWeights16.6 GB 325c72d6ca5a
model-00065-of-000096.safetensorsWeights17.0 GB 276d1cce1d8d
model-00066-of-000096.safetensorsWeights17.0 GB 2ebd83fea628
model-00067-of-000096.safetensorsWeights17.0 GB f0228892f819
model-00068-of-000096.safetensorsWeights16.6 GB fa75764056d1
model-00069-of-000096.safetensorsWeights17.0 GB 9375584663bd
model-00070-of-000096.safetensorsWeights17.0 GB ab5346414872
model-00071-of-000096.safetensorsWeights17.0 GB 28ac0d3286ff
model-00072-of-000096.safetensorsWeights16.6 GB 0a269faaf8ea
model-00073-of-000096.safetensorsWeights17.0 GB a1b2e79e1bb7
model-00074-of-000096.safetensorsWeights17.0 GB df945022b493
model-00075-of-000096.safetensorsWeights17.0 GB 8c6d7cb12f7c
model-00076-of-000096.safetensorsWeights16.6 GB d5174eb5de19
model-00077-of-000096.safetensorsWeights17.0 GB 8707eacfd69e
model-00078-of-000096.safetensorsWeights17.0 GB 2773d41de168
model-00079-of-000096.safetensorsWeights17.0 GB da02c8a46b81
model-00080-of-000096.safetensorsWeights16.6 GB 96b5accdf3bc
model-00081-of-000096.safetensorsWeights17.0 GB 01562aa616ea
model-00082-of-000096.safetensorsWeights17.0 GB 8cba090734d9
model-00083-of-000096.safetensorsWeights17.0 GB 6ceebf8ce621
model-00084-of-000096.safetensorsWeights16.6 GB c3f1318e7e1c
model-00085-of-000096.safetensorsWeights17.0 GB 633b2e3b86ec
model-00086-of-000096.safetensorsWeights17.0 GB add4056ab3ec
model-00087-of-000096.safetensorsWeights17.0 GB f6d608f2c40b
model-00088-of-000096.safetensorsWeights16.6 GB 87afe43b8a71
model-00089-of-000096.safetensorsWeights17.0 GB 24016b28cfdf
model-00090-of-000096.safetensorsWeights17.0 GB 1ecd85dbd77c
model-00091-of-000096.safetensorsWeights17.0 GB a4e666132aed
model-00092-of-000096.safetensorsWeights16.6 GB 359848294be5
model-00093-of-000096.safetensorsWeights16.6 GB d31d58d1bd3f
model-00094-of-000096.safetensorsWeights4.7 GB ad66e1cb96b8
model-00095-of-000096.safetensorsWeights92.3 MB 01d41139abb8
model-00096-of-000096.safetensorsWeights802.4 MB 9d10c74fc101
.eval_results/apex-agents.yamlConfiguration157 B
.eval_results/deep-swe.yamlConfiguration156 B
.eval_results/gpqa.yamlConfiguration152 B
.eval_results/hle.yamlConfiguration139 B
.eval_results/moonshotai__Kimi-K3.yamlConfiguration211 B
config.jsonConfiguration7.0 KB
configuration_kimi_k3.pyConfiguration11.3 KB
encoding_k3.pyConfiguration26.3 KB
generation_config.jsonConfiguration53 B
kimi_k3_processor.pyConfiguration7.7 KB
kimi_k3_vision_processing.pyConfiguration6.7 KB
media_utils.pyConfiguration13.8 KB
model.safetensors.index.jsonConfiguration59.8 MB a1c5210650ce
modeling_kimi_k3.pyConfiguration53.4 KB
modeling_kimi_linear.pyConfiguration51.5 KB
preprocessor_config.jsonConfiguration1.0 KB
tokenization_kimi.pyConfiguration16.1 KB
LICENSEDocumentation3.1 KB
README.mdDocumentation45.3 KB
assets/kimi-logo.pngOther88.0 KB
.gitattributesRepository1.6 KB
tiktoken.modelTokenizer2.8 MB b6c497a7469b
tokenizer_config.jsonTokenizer3.5 KB

License and Download

License
other
Access
Open weights, no gate
Download size
1.6 TB
Download from Moonshot AI

Released by Moonshot AI through its official repository on Hugging Face.

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Delores-Lin/MDPBench Task arMetric arComparison conditions not established 77.4 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task deMetric deComparison conditions not established 89.1 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task digitalMetric digitalComparison conditions not established 90.8 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task enMetric enComparison conditions not established 87.2 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task esMetric esComparison conditions not established 80.2 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task frMetric frComparison conditions not established 80 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task hiMetric hiComparison conditions not established 77.5 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task idMetric idComparison conditions not established 86.9 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task itMetric itComparison conditions not established 92.7 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task jpMetric jpComparison conditions not established 74.9 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task koMetric koComparison conditions not established 89.9 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task latinMetric latinComparison conditions not established 86.2 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task nlMetric nlComparison conditions not established 86 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task non_latinMetric non_latinComparison conditions not established 80.7 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task overallMetric overallComparison conditions not established 83.6 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task photographedMetric photographedComparison conditions not established 81.2 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task privateMetric privateComparison conditions not established 85.6 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task ptMetric ptComparison conditions not established 88.9 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task ruMetric ruComparison conditions not established 82.4 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task thMetric thComparison conditions not established 72.1 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task viMetric viComparison conditions not established 84.8 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task zhMetric zhComparison conditions not established 89.5 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Delores-Lin/MDPBench Task zh_tMetric zh_tComparison conditions not established 81.9 MDPBench leaderboard
Reported by a third party
Evaluated revision not stated 2026-08-15
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 93.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-07-27
IntelligenceLab/Long-Horizon-Terminal-Bench Task lhtb_solvedMetric lhtb_solvedSetup 6/46 tasks solved at reward >= 0.95 ([email protected]); mean reward x100 = 37.8; official LHTB Harbor harnessComparison conditions not established 6 LHTB leaderboard
Reported by a third party
Evaluated revision not stated 2026-07-22
cais/hle Task hleMetric hleComparison conditions not established 56 Model Card
Reported by a third party
Evaluated revision not stated 2026-07-27
crosbylegal/RedlineBench Task redline_overallMetric redline_overallSetup agent=kimi-k3; 3-LLM judge panel (majority vote); turn-weighted weighted pass rate (0-100); post-publication runComparison conditions not established 49.3 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-20
datacurve/deep-swe Task deep_sweMetric deep_sweComparison conditions not established 67.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-07-27
hkust-nlp/Toolathlon Task toolathlon_verifiedMetric toolathlon_verifiedComparison conditions not established 76.5 moonshotai/Kimi-K3 model card
Reported by a third party
Evaluated revision not stated 2026-08-20
internlm/WildClawBench Task avg_costMetric avg_costComparison conditions not established 40.08 WildClawBench
Reported by a third party
Evaluated revision not stated 2026-08-11
internlm/WildClawBench Task avg_timeMetric avg_timeComparison conditions not established 488 WildClawBench
Reported by a third party
Evaluated revision not stated 2026-08-11
internlm/WildClawBench Task overallMetric overallComparison conditions not established 54.5 WildClawBench
Reported by a third party
Evaluated revision not stated 2026-08-11
joelniklaus/LEXam-hard Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established 29.54 SwissLegalEvals per-sample details (lighteval)
Reported by a third party
Evaluated revision not stated 2026-08-03
llamaindex/ExtractBench Task longMetric longSetup Pipeline name: kimi_k3_extract_oneshot_structured_output_file; served via Fireworks hosted APIComparison conditions not established 4.99 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-25
llamaindex/ExtractBench Task meanMetric meanSetup Pipeline name: kimi_k3_extract_oneshot_structured_output_file; served via Fireworks hosted APIComparison conditions not established 83.17 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-25
llamaindex/ExtractBench Task mediumMetric mediumSetup Pipeline name: kimi_k3_extract_oneshot_structured_output_file; served via Fireworks hosted APIComparison conditions not established 69.64 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-25
llamaindex/ExtractBench Task shortMetric shortSetup Pipeline name: kimi_k3_extract_oneshot_structured_output_file; served via Fireworks hosted APIComparison conditions not established 94.64 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-25
mercor/apex-agents Task apex-agentsMetric apex-agentsComparison conditions not established 41 Model Card
Reported by a third party
Evaluated revision not stated 2026-07-27

Memory Requirements

PrecisionWeights in memory
As published1.6 TB
16-bit5559.9 GB
8-bit2779.9 GB
4-bit1390.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
Baseten$3.00 / $15.00input / output, per million tokensSep 18, 2026
DeepInfra$2.85 / $14.25input / output, per million tokensSep 18, 2026
Fireworks$3.00 / $15.00input / output, per million tokensSep 18, 2026
Fireworks$3.00 / $15.00input / output, per million tokensSep 18, 2026
Together AI$3.00 / $15.00input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Built on This Model

Questions About Kimi-K3

How much GPU memory does Kimi-K3 need?

About 6671.8 GB at 16-bit and 1668 GB at 4-bit: the weights (2.8T parameters) plus a working margin. A long context needs more.

What license is Kimi-K3 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Kimi-K3's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Kimi-K3-W4A16-RTN

Aaron Beckley

Kimi K3 on a single NVIDIA A100 80GB. A weight-only quantisation of Moonshot AI's Kimi K3 (2.8T total / 104B activated parameters) that loads and generates on one A100 80GB GPU, with the routed experts held in host RAM. No Ampere-targeted K3 build existed for vLLM, so this was made to fix that gap. Routed experts and attention re-encoded from MXFP4/BF16 into compressed-tensors pack-quantized, served by vLLM's Marlin kernels. Activations stay BF16 (W4A16 / W8A16). Round-to-nearest only — no calibration data, so no dataset is baked into these weights. Errors were measured by round-tripping each tensor through compressed-tensors' compress()/decompress(). Weight error is a proxy, not a quality…

Open weights other 2.7T parameters 1,048,576 tokens vllm

Model · Image and text to text

Qwen3.5-397B-A17B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. MathVision:our model’s score is…

Open weights apache-2.0 403.4B parameters 262,144 tokens transformers

Model · Image and text to text

GLM-5.3-Flash

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.3-Flash blog and GLM-5 Technical report. Use GLM-5.3-Flash API services on Z.ai API Platform. We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear…

Open weights mit 321.3B parameters 1,048,576 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next

Qwen

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next-Uncensored-NVFP4

OrcaRouter

This model has had its safety alignment substantially removed via abliteration (orthogonalizing the refusal direction out of the residual stream). It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment. Use must comply with the Apache 2.0 License inherited from the base model and all applicable law. The authors accept no liability for misuse. - A Blackwell GPU…

Access requested at publisher apache-2.0 180B parameters transformers

Model · Image and text to text

Qwen3.8-Flash-Next-MLX-oQ3-MTP

Robot Haus

A sensitivity-guided, mixed-precision MLX conversion of Qwen/Qwen3.8-Flash-Next, rebuilt directly from the official BF16 checkpoint with the model's matching native MTP block preserved. oQ3 uses a 3-bit affine base and spends additional precision on sensitive modules. Layer sensitivity was measured with a validated quantized calibration proxy, while every released weight was quantized from the official BF16 checkpoint. The result is a compact model with 746 higher-precision module overrides rather than a uniform 3-bit layout. The upstream tokenizer, current chat template, vision processor, generation configuration, licence, and native MTP configuration are retained. In a compatible oMLX…

Open weights other 180B parameters 262,144 tokens mlx