Kimi K3 on a single NVIDIA A100 80GB. A weight-only quantisation of Moonshot AI's Kimi K3 (2.8T total / 104B activated parameters) that loads and generates on one A100 80GB GPU, with the routed experts held in host RAM. No Ampere-targeted K3 build existed for vLLM, so this was made to fix that gap. Routed experts and attention re-encoded from MXFP4/BF16 into compressed-tensors pack-quantized, served by vLLM's Marlin kernels. Activations stay BF16 (W4A16 / W8A16). Round-to-nearest only — no calibration data, so no dataset is baked into these weights. Errors were measured by round-tripping each tensor through compressed-tensors' compress()/decompress(). Weight error is a proxy, not a quality…
Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date.
Runs On
What it takes to serve Kimi-K3 (2.8T parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 5559.9 GB | 6671.8 GB | More than one server of any accelerator the SAVRN Index prices. | ||
| 8-bit | 2779.9 GB | 3335.9 GB | More than one server of any accelerator the SAVRN Index prices. | ||
| 4-bit | 1390.0 GB | 1668.0 GB | 7x MI325X (256 GB) Vultr |
$14.00 | 6x MI355X $15.54 · 7x B300 $46.20 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
SAVRN's Notes on Kimi-K3
Nobody runs this one on a single box. At 16-bit the weights are 5,559.9 GB with 6,671.8 GB needed, more than one server of anything we price, and 8-bit still spills over. 4-bit is the first fit: 1,390.0 GB of weights, 1,668.0 GB needed, on seven MI325X cards with 256 GB each at $14.00 per hour. That buys Moonshot AI's 2.8T-parameter model, image and text to text, with a 1,048,576-token context window built for long-horizon coding, knowledge work and reasoning.
The license field reads other, with no summary in our file, so we cannot say whether commercial use is permitted; read the publisher's text before planning a deployment. Then weigh the rent. Baseten, Fireworks and Together AI serve it at $3.00 per million input tokens and $15.00 per million output, DeepInfra at $2.85 and $14.25, and that bill is what your seven-card cluster has to beat at your volume.
Model Card
Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning. All Kimi K3 results are obtained with reasoning effort set to 'max' and temperature = 1.0. For single-step tasks, such as GPQA Diamond, HLE-Full, and vision benchmarks without tools, we set top-p = 0.95; for agentic tasks, we set top-p = 1.0. For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision…
Excerpt from the card by Moonshot AI, licensed other.
Configuration
- Architecture
- KimiK3ForConditionalGeneration
- Context length (tokens)
- 1,048,576
- Layers
- 93
- Hidden size
- 7,168
- Feed-forward size
- 33,792
- Attention heads
- 96
- Key/value heads
- 96
- Vocabulary size
- 163,840
- Experts
- 896
- Model type
- kimi_k3
Identity and Version
- Repository
- moonshotai/Kimi-K3
- Publisher
- Moonshot AI
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 2.8T parameters
- Languages
- Not stated by the source
- Revision
- f831ab66814297da540d832a5235f8e904f29d06
- First published
- 2026-06-13
- Last updated
- 2026-09-02
Files and Weights
119 files, 1.6 TB in total. The weights are 96 files totalling 1.6 TB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-000096.safetensors | Weights | 2.3 GB | 975584c00f85 |
| model-00002-of-000096.safetensors | Weights | 17.0 GB | 26a3284e1d2c |
| model-00003-of-000096.safetensors | Weights | 17.0 GB | e54af9de4c55 |
| model-00004-of-000096.safetensors | Weights | 16.6 GB | 5955fd8feda8 |
| model-00005-of-000096.safetensors | Weights | 17.0 GB | d60d68ad0381 |
| model-00006-of-000096.safetensors | Weights | 17.0 GB | b1d480576747 |
| model-00007-of-000096.safetensors | Weights | 17.0 GB | fb1120fef34c |
| model-00008-of-000096.safetensors | Weights | 16.6 GB | 2318dda54fc1 |
| model-00009-of-000096.safetensors | Weights | 17.0 GB | 8b66cdde34f5 |
| model-00010-of-000096.safetensors | Weights | 17.0 GB | d34f55f7b734 |
| model-00011-of-000096.safetensors | Weights | 17.0 GB | 1738856a4cca |
| model-00012-of-000096.safetensors | Weights | 16.6 GB | c6b9bef38415 |
| model-00013-of-000096.safetensors | Weights | 17.0 GB | 3cbf43d56d9c |
| model-00014-of-000096.safetensors | Weights | 17.0 GB | ce5f343f07d4 |
| model-00015-of-000096.safetensors | Weights | 17.0 GB | b55caa801334 |
| model-00016-of-000096.safetensors | Weights | 16.6 GB | 5a63e63ced65 |
| model-00017-of-000096.safetensors | Weights | 17.0 GB | 622bfa605205 |
| model-00018-of-000096.safetensors | Weights | 17.0 GB | 838b265ebb86 |
| model-00019-of-000096.safetensors | Weights | 17.0 GB | 599a8ecf88f2 |
| model-00020-of-000096.safetensors | Weights | 16.6 GB | 8e01b61ab766 |
| model-00021-of-000096.safetensors | Weights | 17.0 GB | 1944265bd024 |
| model-00022-of-000096.safetensors | Weights | 17.0 GB | 2d32d3e3c8da |
| model-00023-of-000096.safetensors | Weights | 17.0 GB | c291f2ec1578 |
| model-00024-of-000096.safetensors | Weights | 16.6 GB | 278855ab81d4 |
| model-00025-of-000096.safetensors | Weights | 17.0 GB | 6e3ae6ad868f |
| model-00026-of-000096.safetensors | Weights | 17.0 GB | cdd79fb52c7a |
| model-00027-of-000096.safetensors | Weights | 17.0 GB | 1d974b40c4ae |
| model-00028-of-000096.safetensors | Weights | 16.6 GB | 1bec58e89ab7 |
| model-00029-of-000096.safetensors | Weights | 17.0 GB | c3bf0e738aa5 |
| model-00030-of-000096.safetensors | Weights | 17.0 GB | 7b996410482a |
| model-00031-of-000096.safetensors | Weights | 17.0 GB | c26689540a24 |
| model-00032-of-000096.safetensors | Weights | 16.6 GB | f7dc9e726d46 |
| model-00033-of-000096.safetensors | Weights | 17.0 GB | 615afa33b69c |
| model-00034-of-000096.safetensors | Weights | 17.0 GB | a53b27fe92df |
| model-00035-of-000096.safetensors | Weights | 17.0 GB | 9f4b44c89e49 |
| model-00036-of-000096.safetensors | Weights | 16.6 GB | 55aef33fab36 |
| model-00037-of-000096.safetensors | Weights | 17.0 GB | e95dd3599d69 |
| model-00038-of-000096.safetensors | Weights | 17.0 GB | 1ce472771309 |
| model-00039-of-000096.safetensors | Weights | 17.0 GB | 9c75b18c0d3a |
| model-00040-of-000096.safetensors | Weights | 16.6 GB | 6f664598c00a |
| model-00041-of-000096.safetensors | Weights | 17.0 GB | 375f05f94a59 |
| model-00042-of-000096.safetensors | Weights | 17.0 GB | b65947611d6c |
| model-00043-of-000096.safetensors | Weights | 17.0 GB | b5a425f100bb |
| model-00044-of-000096.safetensors | Weights | 16.6 GB | 113ee0120442 |
| model-00045-of-000096.safetensors | Weights | 17.0 GB | eb0698659da5 |
| model-00046-of-000096.safetensors | Weights | 17.0 GB | 0270727a399c |
| model-00047-of-000096.safetensors | Weights | 17.0 GB | b38f63eb0803 |
| model-00048-of-000096.safetensors | Weights | 16.6 GB | 131e243c02cf |
| model-00049-of-000096.safetensors | Weights | 17.0 GB | 72c91dcf2909 |
| model-00050-of-000096.safetensors | Weights | 17.0 GB | 0a627a082cd3 |
| model-00051-of-000096.safetensors | Weights | 17.0 GB | 38a37bcec20a |
| model-00052-of-000096.safetensors | Weights | 16.6 GB | 9703b6321741 |
| model-00053-of-000096.safetensors | Weights | 17.0 GB | 413ec9ea0b69 |
| model-00054-of-000096.safetensors | Weights | 17.0 GB | 0e3da201b765 |
| model-00055-of-000096.safetensors | Weights | 17.0 GB | e7c9f4e44f8a |
| model-00056-of-000096.safetensors | Weights | 16.6 GB | efd6176016b9 |
| model-00057-of-000096.safetensors | Weights | 17.0 GB | e6982c96ac91 |
| model-00058-of-000096.safetensors | Weights | 17.0 GB | c93a23cbd653 |
| model-00059-of-000096.safetensors | Weights | 17.0 GB | a62eb8220710 |
| model-00060-of-000096.safetensors | Weights | 16.6 GB | 9bccbaa71b98 |
| model-00061-of-000096.safetensors | Weights | 17.0 GB | 1ae3969540fc |
| model-00062-of-000096.safetensors | Weights | 17.0 GB | 96babdc24f22 |
| model-00063-of-000096.safetensors | Weights | 17.0 GB | 8a81caa697a7 |
| model-00064-of-000096.safetensors | Weights | 16.6 GB | 325c72d6ca5a |
| model-00065-of-000096.safetensors | Weights | 17.0 GB | 276d1cce1d8d |
| model-00066-of-000096.safetensors | Weights | 17.0 GB | 2ebd83fea628 |
| model-00067-of-000096.safetensors | Weights | 17.0 GB | f0228892f819 |
| model-00068-of-000096.safetensors | Weights | 16.6 GB | fa75764056d1 |
| model-00069-of-000096.safetensors | Weights | 17.0 GB | 9375584663bd |
| model-00070-of-000096.safetensors | Weights | 17.0 GB | ab5346414872 |
| model-00071-of-000096.safetensors | Weights | 17.0 GB | 28ac0d3286ff |
| model-00072-of-000096.safetensors | Weights | 16.6 GB | 0a269faaf8ea |
| model-00073-of-000096.safetensors | Weights | 17.0 GB | a1b2e79e1bb7 |
| model-00074-of-000096.safetensors | Weights | 17.0 GB | df945022b493 |
| model-00075-of-000096.safetensors | Weights | 17.0 GB | 8c6d7cb12f7c |
| model-00076-of-000096.safetensors | Weights | 16.6 GB | d5174eb5de19 |
| model-00077-of-000096.safetensors | Weights | 17.0 GB | 8707eacfd69e |
| model-00078-of-000096.safetensors | Weights | 17.0 GB | 2773d41de168 |
| model-00079-of-000096.safetensors | Weights | 17.0 GB | da02c8a46b81 |
| model-00080-of-000096.safetensors | Weights | 16.6 GB | 96b5accdf3bc |
| model-00081-of-000096.safetensors | Weights | 17.0 GB | 01562aa616ea |
| model-00082-of-000096.safetensors | Weights | 17.0 GB | 8cba090734d9 |
| model-00083-of-000096.safetensors | Weights | 17.0 GB | 6ceebf8ce621 |
| model-00084-of-000096.safetensors | Weights | 16.6 GB | c3f1318e7e1c |
| model-00085-of-000096.safetensors | Weights | 17.0 GB | 633b2e3b86ec |
| model-00086-of-000096.safetensors | Weights | 17.0 GB | add4056ab3ec |
| model-00087-of-000096.safetensors | Weights | 17.0 GB | f6d608f2c40b |
| model-00088-of-000096.safetensors | Weights | 16.6 GB | 87afe43b8a71 |
| model-00089-of-000096.safetensors | Weights | 17.0 GB | 24016b28cfdf |
| model-00090-of-000096.safetensors | Weights | 17.0 GB | 1ecd85dbd77c |
| model-00091-of-000096.safetensors | Weights | 17.0 GB | a4e666132aed |
| model-00092-of-000096.safetensors | Weights | 16.6 GB | 359848294be5 |
| model-00093-of-000096.safetensors | Weights | 16.6 GB | d31d58d1bd3f |
| model-00094-of-000096.safetensors | Weights | 4.7 GB | ad66e1cb96b8 |
| model-00095-of-000096.safetensors | Weights | 92.3 MB | 01d41139abb8 |
| model-00096-of-000096.safetensors | Weights | 802.4 MB | 9d10c74fc101 |
| .eval_results/apex-agents.yaml | Configuration | 157 B | — |
| .eval_results/deep-swe.yaml | Configuration | 156 B | — |
| .eval_results/gpqa.yaml | Configuration | 152 B | — |
| .eval_results/hle.yaml | Configuration | 139 B | — |
| .eval_results/moonshotai__Kimi-K3.yaml | Configuration | 211 B | — |
| config.json | Configuration | 7.0 KB | — |
| configuration_kimi_k3.py | Configuration | 11.3 KB | — |
| encoding_k3.py | Configuration | 26.3 KB | — |
| generation_config.json | Configuration | 53 B | — |
| kimi_k3_processor.py | Configuration | 7.7 KB | — |
| kimi_k3_vision_processing.py | Configuration | 6.7 KB | — |
| media_utils.py | Configuration | 13.8 KB | — |
| model.safetensors.index.json | Configuration | 59.8 MB | a1c5210650ce |
| modeling_kimi_k3.py | Configuration | 53.4 KB | — |
| modeling_kimi_linear.py | Configuration | 51.5 KB | — |
| preprocessor_config.json | Configuration | 1.0 KB | — |
| tokenization_kimi.py | Configuration | 16.1 KB | — |
| LICENSE | Documentation | 3.1 KB | — |
| README.md | Documentation | 45.3 KB | — |
| assets/kimi-logo.png | Other | 88.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tiktoken.model | Tokenizer | 2.8 MB | b6c497a7469b |
| tokenizer_config.json | Tokenizer | 3.5 KB | — |
License and Download
- License
- other
- Access
- Open weights, no gate
- Download size
- 1.6 TB
Released by Moonshot AI through its official repository on Hugging Face.
Evaluations
Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Delores-Lin/MDPBench | Task arMetric arComparison conditions not established | 77.4 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task deMetric deComparison conditions not established | 89.1 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task digitalMetric digitalComparison conditions not established | 90.8 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task enMetric enComparison conditions not established | 87.2 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task esMetric esComparison conditions not established | 80.2 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task frMetric frComparison conditions not established | 80 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task hiMetric hiComparison conditions not established | 77.5 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task idMetric idComparison conditions not established | 86.9 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task itMetric itComparison conditions not established | 92.7 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task jpMetric jpComparison conditions not established | 74.9 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task koMetric koComparison conditions not established | 89.9 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task latinMetric latinComparison conditions not established | 86.2 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task nlMetric nlComparison conditions not established | 86 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task non_latinMetric non_latinComparison conditions not established | 80.7 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task overallMetric overallComparison conditions not established | 83.6 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task photographedMetric photographedComparison conditions not established | 81.2 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task privateMetric privateComparison conditions not established | 85.6 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task ptMetric ptComparison conditions not established | 88.9 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task ruMetric ruComparison conditions not established | 82.4 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task thMetric thComparison conditions not established | 72.1 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task viMetric viComparison conditions not established | 84.8 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task zhMetric zhComparison conditions not established | 89.5 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Delores-Lin/MDPBench | Task zh_tMetric zh_tComparison conditions not established | 81.9 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-08-15 |
| Idavidrein/gpqa | Task diamondMetric diamondComparison conditions not established | 93.5 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-07-27 |
| IntelligenceLab/Long-Horizon-Terminal-Bench | Task lhtb_solvedMetric lhtb_solvedSetup 6/46 tasks solved at reward >= 0.95 ([email protected]); mean reward x100 = 37.8; official LHTB Harbor harnessComparison conditions not established | 6 | LHTB leaderboard Reported by a third party |
Evaluated revision not stated | 2026-07-22 |
| cais/hle | Task hleMetric hleComparison conditions not established | 56 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-07-27 |
| crosbylegal/RedlineBench | Task redline_overallMetric redline_overallSetup agent=kimi-k3; 3-LLM judge panel (majority vote); turn-weighted weighted pass rate (0-100); post-publication runComparison conditions not established | 49.3 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-08-20 |
| datacurve/deep-swe | Task deep_sweMetric deep_sweComparison conditions not established | 67.5 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-07-27 |
| hkust-nlp/Toolathlon | Task toolathlon_verifiedMetric toolathlon_verifiedComparison conditions not established | 76.5 | moonshotai/Kimi-K3 model card Reported by a third party |
Evaluated revision not stated | 2026-08-20 |
| internlm/WildClawBench | Task avg_costMetric avg_costComparison conditions not established | 40.08 | WildClawBench Reported by a third party |
Evaluated revision not stated | 2026-08-11 |
| internlm/WildClawBench | Task avg_timeMetric avg_timeComparison conditions not established | 488 | WildClawBench Reported by a third party |
Evaluated revision not stated | 2026-08-11 |
| internlm/WildClawBench | Task overallMetric overallComparison conditions not established | 54.5 | WildClawBench Reported by a third party |
Evaluated revision not stated | 2026-08-11 |
| joelniklaus/LEXam-hard | Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established | 29.54 | SwissLegalEvals per-sample details (lighteval) Reported by a third party |
Evaluated revision not stated | 2026-08-03 |
| llamaindex/ExtractBench | Task longMetric longSetup Pipeline name: kimi_k3_extract_oneshot_structured_output_file; served via Fireworks hosted APIComparison conditions not established | 4.99 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-25 |
| llamaindex/ExtractBench | Task meanMetric meanSetup Pipeline name: kimi_k3_extract_oneshot_structured_output_file; served via Fireworks hosted APIComparison conditions not established | 83.17 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-25 |
| llamaindex/ExtractBench | Task mediumMetric mediumSetup Pipeline name: kimi_k3_extract_oneshot_structured_output_file; served via Fireworks hosted APIComparison conditions not established | 69.64 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-25 |
| llamaindex/ExtractBench | Task shortMetric shortSetup Pipeline name: kimi_k3_extract_oneshot_structured_output_file; served via Fireworks hosted APIComparison conditions not established | 94.64 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-25 |
| mercor/apex-agents | Task apex-agentsMetric apex-agentsComparison conditions not established | 41 | Model Card Reported by a third party |
Evaluated revision not stated | 2026-07-27 |
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 1.6 TB |
| 16-bit | 5559.9 GB |
| 8-bit | 2779.9 GB |
| 4-bit | 1390.0 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Hosted Prices
| Host | Input / output | Unit | Observed |
|---|---|---|---|
| Baseten | $3.00 / $15.00 | input / output, per million tokens | Sep 18, 2026 |
| DeepInfra | $2.85 / $14.25 | input / output, per million tokens | Sep 18, 2026 |
| Fireworks | $3.00 / $15.00 | input / output, per million tokens | Sep 18, 2026 |
| Fireworks | $3.00 / $15.00 | input / output, per million tokens | Sep 18, 2026 |
| Together AI | $3.00 / $15.00 | input / output, per million tokens | Sep 18, 2026 |
From the SAVRN Index.
Built on This Model
- Quantized fromKimi-K3-W4A16-RTN
- Derived fromKimi-K3-W4A16-RTN
Questions About Kimi-K3
How much GPU memory does Kimi-K3 need?
About 6671.8 GB at 16-bit and 1668 GB at 4-bit: the weights (2.8T parameters) plus a working margin. A long context needs more.
What license is Kimi-K3 released under?
other, as its publisher declares it. Read the license text before commercial use.
What is Kimi-K3's context length?
1,048,576 tokens, from the maximum position embeddings in its published configuration.
Similar Models
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. MathVision:our model’s score is…
Join our WeChat or Discord community. Check out the GLM-5.3-Flash blog and GLM-5 Technical report. Use GLM-5.3-Flash API services on Z.ai API Platform. We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear…
As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…
This model has had its safety alignment substantially removed via abliteration (orthogonalizing the refusal direction out of the residual stream). It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment. Use must comply with the Apache 2.0 License inherited from the base model and all applicable law. The authors accept no liability for misuse. - A Blackwell GPU…
A sensitivity-guided, mixed-precision MLX conversion of Qwen/Qwen3.8-Flash-Next, rebuilt directly from the official BF16 checkpoint with the model's matching native MTP block preserved. oQ3 uses a 3-bit affine base and spends additional precision on sensitive modules. Layer sensitivity was measured with a validated quantized calibration proxy, while every released weight was quantized from the official BF16 checkpoint. The result is a compact model with 746 higher-precision module overrides rather than a uniform 3-bit layout. The upstream tokenizer, current chat template, vision processor, generation configuration, licence, and native MTP configuration are retained. In a compatible oMLX…