SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Kimi-K3

by SAIFI INDUSTRIES SAIFIINDUSTRIES/Kimi-K3

Kimi-K3 is an open-weight model for image and text to text from SAIFI INDUSTRIES, released under other. It has 2.8T parameters and a 1,048,576-token context. Its published files total 1.6 TB.

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date.

Parameters2.8T
Context1,048,576
Weights1.6 TB
Licenseother
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Kimi-K3 (2.8T parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 5559.9 GB 6671.8 GB More than one server of any accelerator the SAVRN Index prices.
8-bit 2779.9 GB 3335.9 GB More than one server of any accelerator the SAVRN Index prices.
4-bit 1390.0 GB 1668.0 GB 7x MI325X (256 GB)
Vultr
$14.00 6x MI355X $15.54 · 7x B300 $46.20

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Kimi-K3 on every accelerator the SAVRN Index prices, at every precision

Model Card

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning. All Kimi K3 results are obtained with reasoning effort set to 'max' and temperature = 1.0. For single-step tasks, such as GPQA Diamond, HLE-Full, and vision benchmarks without tools, we set top-p = 0.95; for agentic tasks, we set top-p = 1.0. For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision…

Excerpt from the card by SAIFI INDUSTRIES, licensed other.

Configuration

Architecture
KimiK3ForConditionalGeneration
Context length (tokens)
1,048,576
Layers
93
Hidden size
7,168
Feed-forward size
33,792
Attention heads
96
Key/value heads
96
Vocabulary size
163,840
Experts
896
Model type
kimi_k3

Identity and Version

Repository
SAIFIINDUSTRIES/Kimi-K3
Publisher
SAIFI INDUSTRIES
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
2.8T parameters
Languages
Not stated by the source
Revision
2896233350d6856bc7f6cb478c1c4f706075c1ec
First published
2026-10-03
Last updated
2026-10-03

Files and Weights

119 files, 1.6 TB in total. The weights are 96 files totalling 1.6 TB in safetensors.

Weights96 files · 1.6 TB
Configuration17 files · 60.0 MB
Tokenizer2 files · 2.8 MB
Documentation2 files · 48.3 KB
Other1 file · 88.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-000096.safetensorsWeights2.3 GB 975584c00f85
model-00002-of-000096.safetensorsWeights17.0 GB 26a3284e1d2c
model-00003-of-000096.safetensorsWeights17.0 GB e54af9de4c55
model-00004-of-000096.safetensorsWeights16.6 GB 5955fd8feda8
model-00005-of-000096.safetensorsWeights17.0 GB d60d68ad0381
model-00006-of-000096.safetensorsWeights17.0 GB b1d480576747
model-00007-of-000096.safetensorsWeights17.0 GB fb1120fef34c
model-00008-of-000096.safetensorsWeights16.6 GB 2318dda54fc1
model-00009-of-000096.safetensorsWeights17.0 GB 8b66cdde34f5
model-00010-of-000096.safetensorsWeights17.0 GB d34f55f7b734
model-00011-of-000096.safetensorsWeights17.0 GB 1738856a4cca
model-00012-of-000096.safetensorsWeights16.6 GB c6b9bef38415
model-00013-of-000096.safetensorsWeights17.0 GB 3cbf43d56d9c
model-00014-of-000096.safetensorsWeights17.0 GB ce5f343f07d4
model-00015-of-000096.safetensorsWeights17.0 GB b55caa801334
model-00016-of-000096.safetensorsWeights16.6 GB 5a63e63ced65
model-00017-of-000096.safetensorsWeights17.0 GB 622bfa605205
model-00018-of-000096.safetensorsWeights17.0 GB 838b265ebb86
model-00019-of-000096.safetensorsWeights17.0 GB 599a8ecf88f2
model-00020-of-000096.safetensorsWeights16.6 GB 8e01b61ab766
model-00021-of-000096.safetensorsWeights17.0 GB 1944265bd024
model-00022-of-000096.safetensorsWeights17.0 GB 2d32d3e3c8da
model-00023-of-000096.safetensorsWeights17.0 GB c291f2ec1578
model-00024-of-000096.safetensorsWeights16.6 GB 278855ab81d4
model-00025-of-000096.safetensorsWeights17.0 GB 6e3ae6ad868f
model-00026-of-000096.safetensorsWeights17.0 GB cdd79fb52c7a
model-00027-of-000096.safetensorsWeights17.0 GB 1d974b40c4ae
model-00028-of-000096.safetensorsWeights16.6 GB 1bec58e89ab7
model-00029-of-000096.safetensorsWeights17.0 GB c3bf0e738aa5
model-00030-of-000096.safetensorsWeights17.0 GB 7b996410482a
model-00031-of-000096.safetensorsWeights17.0 GB c26689540a24
model-00032-of-000096.safetensorsWeights16.6 GB f7dc9e726d46
model-00033-of-000096.safetensorsWeights17.0 GB 615afa33b69c
model-00034-of-000096.safetensorsWeights17.0 GB a53b27fe92df
model-00035-of-000096.safetensorsWeights17.0 GB 9f4b44c89e49
model-00036-of-000096.safetensorsWeights16.6 GB 55aef33fab36
model-00037-of-000096.safetensorsWeights17.0 GB e95dd3599d69
model-00038-of-000096.safetensorsWeights17.0 GB 1ce472771309
model-00039-of-000096.safetensorsWeights17.0 GB 9c75b18c0d3a
model-00040-of-000096.safetensorsWeights16.6 GB 6f664598c00a
model-00041-of-000096.safetensorsWeights17.0 GB 375f05f94a59
model-00042-of-000096.safetensorsWeights17.0 GB b65947611d6c
model-00043-of-000096.safetensorsWeights17.0 GB b5a425f100bb
model-00044-of-000096.safetensorsWeights16.6 GB 113ee0120442
model-00045-of-000096.safetensorsWeights17.0 GB eb0698659da5
model-00046-of-000096.safetensorsWeights17.0 GB 0270727a399c
model-00047-of-000096.safetensorsWeights17.0 GB b38f63eb0803
model-00048-of-000096.safetensorsWeights16.6 GB 131e243c02cf
model-00049-of-000096.safetensorsWeights17.0 GB 72c91dcf2909
model-00050-of-000096.safetensorsWeights17.0 GB 0a627a082cd3
model-00051-of-000096.safetensorsWeights17.0 GB 38a37bcec20a
model-00052-of-000096.safetensorsWeights16.6 GB 9703b6321741
model-00053-of-000096.safetensorsWeights17.0 GB 413ec9ea0b69
model-00054-of-000096.safetensorsWeights17.0 GB 0e3da201b765
model-00055-of-000096.safetensorsWeights17.0 GB e7c9f4e44f8a
model-00056-of-000096.safetensorsWeights16.6 GB efd6176016b9
model-00057-of-000096.safetensorsWeights17.0 GB e6982c96ac91
model-00058-of-000096.safetensorsWeights17.0 GB c93a23cbd653
model-00059-of-000096.safetensorsWeights17.0 GB a62eb8220710
model-00060-of-000096.safetensorsWeights16.6 GB 9bccbaa71b98
model-00061-of-000096.safetensorsWeights17.0 GB 1ae3969540fc
model-00062-of-000096.safetensorsWeights17.0 GB 96babdc24f22
model-00063-of-000096.safetensorsWeights17.0 GB 8a81caa697a7
model-00064-of-000096.safetensorsWeights16.6 GB 325c72d6ca5a
model-00065-of-000096.safetensorsWeights17.0 GB 276d1cce1d8d
model-00066-of-000096.safetensorsWeights17.0 GB 2ebd83fea628
model-00067-of-000096.safetensorsWeights17.0 GB f0228892f819
model-00068-of-000096.safetensorsWeights16.6 GB fa75764056d1
model-00069-of-000096.safetensorsWeights17.0 GB 9375584663bd
model-00070-of-000096.safetensorsWeights17.0 GB ab5346414872
model-00071-of-000096.safetensorsWeights17.0 GB 28ac0d3286ff
model-00072-of-000096.safetensorsWeights16.6 GB 0a269faaf8ea
model-00073-of-000096.safetensorsWeights17.0 GB a1b2e79e1bb7
model-00074-of-000096.safetensorsWeights17.0 GB df945022b493
model-00075-of-000096.safetensorsWeights17.0 GB 8c6d7cb12f7c
model-00076-of-000096.safetensorsWeights16.6 GB d5174eb5de19
model-00077-of-000096.safetensorsWeights17.0 GB 8707eacfd69e
model-00078-of-000096.safetensorsWeights17.0 GB 2773d41de168
model-00079-of-000096.safetensorsWeights17.0 GB da02c8a46b81
model-00080-of-000096.safetensorsWeights16.6 GB 96b5accdf3bc
model-00081-of-000096.safetensorsWeights17.0 GB 01562aa616ea
model-00082-of-000096.safetensorsWeights17.0 GB 8cba090734d9
model-00083-of-000096.safetensorsWeights17.0 GB 6ceebf8ce621
model-00084-of-000096.safetensorsWeights16.6 GB c3f1318e7e1c
model-00085-of-000096.safetensorsWeights17.0 GB 633b2e3b86ec
model-00086-of-000096.safetensorsWeights17.0 GB add4056ab3ec
model-00087-of-000096.safetensorsWeights17.0 GB f6d608f2c40b
model-00088-of-000096.safetensorsWeights16.6 GB 87afe43b8a71
model-00089-of-000096.safetensorsWeights17.0 GB 24016b28cfdf
model-00090-of-000096.safetensorsWeights17.0 GB 1ecd85dbd77c
model-00091-of-000096.safetensorsWeights17.0 GB a4e666132aed
model-00092-of-000096.safetensorsWeights16.6 GB 359848294be5
model-00093-of-000096.safetensorsWeights16.6 GB d31d58d1bd3f
model-00094-of-000096.safetensorsWeights4.7 GB ad66e1cb96b8
model-00095-of-000096.safetensorsWeights92.3 MB 01d41139abb8
model-00096-of-000096.safetensorsWeights802.4 MB 9d10c74fc101
.eval_results/apex-agents.yamlConfiguration157 B —
.eval_results/deep-swe.yamlConfiguration156 B —
.eval_results/gpqa.yamlConfiguration152 B —
.eval_results/hle.yamlConfiguration139 B —
.eval_results/moonshotai__Kimi-K3.yamlConfiguration211 B —
config.jsonConfiguration7.0 KB —
configuration_kimi_k3.pyConfiguration11.3 KB —
encoding_k3.pyConfiguration26.3 KB —
generation_config.jsonConfiguration53 B —
kimi_k3_processor.pyConfiguration7.7 KB —
kimi_k3_vision_processing.pyConfiguration6.7 KB —
media_utils.pyConfiguration13.8 KB —
model.safetensors.index.jsonConfiguration59.8 MB a1c5210650ce
modeling_kimi_k3.pyConfiguration53.4 KB —
modeling_kimi_linear.pyConfiguration51.5 KB —
preprocessor_config.jsonConfiguration1.0 KB —
tokenization_kimi.pyConfiguration16.1 KB —
LICENSEDocumentation3.1 KB —
README.mdDocumentation45.3 KB —
assets/kimi-logo.pngOther88.0 KB —
.gitattributesRepository1.6 KB —
tiktoken.modelTokenizer2.8 MB b6c497a7469b
tokenizer_config.jsonTokenizer3.5 KB —

License and Download

License
other
Access
Open weights, no gate
Download size
1.6 TB
Download from SAIFI INDUSTRIES

Released by SAIFI INDUSTRIES through its official repository on Hugging Face.

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 93.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-10-03
cais/hle Task hleMetric hleComparison conditions not established 56 Model Card
Reported by a third party
Evaluated revision not stated 2026-10-03
datacurve/deep-swe Task deep_sweMetric deep_sweComparison conditions not established 67.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-10-03
hkust-nlp/Toolathlon Task toolathlon_verifiedMetric toolathlon_verifiedComparison conditions not established 76.5 moonshotai/Kimi-K3 model card
Reported by a third party
Evaluated revision not stated 2026-08-20
mercor/apex-agents Task apex-agentsMetric apex-agentsComparison conditions not established 41 Model Card
Reported by a third party
Evaluated revision not stated 2026-10-03

Memory Requirements

PrecisionWeights in memory
As published1.6 TB
16-bit5559.9 GB
8-bit2779.9 GB
4-bit1390.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Kimi-K3

How much GPU memory does Kimi-K3 need?

About 6671.8 GB at 16-bit and 1668 GB at 4-bit: the weights (2.8T parameters) plus a working margin. A long context needs more.

What license is Kimi-K3 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Kimi-K3's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Kimi-K3

Moonshot AI

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning. All Kimi K3 results are obtained with reasoning effort set to 'max' and temperature = 1.0. For single-step tasks, such as GPQA Diamond, HLE-Full, and vision benchmarks without tools, we set top-p = 0.95; for agentic tasks, we set top-p = 1.0. For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision…

Open weights other 2.8T parameters 1,048,576 tokens transformers

Model · Image and text to text

Kimi-K3-W4A16-RTN

Aaron Beckley

Kimi K3 on a single NVIDIA A100 80GB. A weight-only quantisation of Moonshot AI's Kimi K3 (2.8T total / 104B activated parameters) that loads and generates on one A100 80GB GPU, with the routed experts held in host RAM. No Ampere-targeted K3 build existed for vLLM, so this was made to fix that gap. Routed experts and attention re-encoded from MXFP4/BF16 into compressed-tensors pack-quantized, served by vLLM's Marlin kernels. Activations stay BF16 (W4A16 / W8A16). Round-to-nearest only — no calibration data, so no dataset is baked into these weights. Errors were measured by round-tripping each tensor through compressed-tensors' compress()/decompress(). Weight error is a proxy, not a quality…

Open weights other 2.7T parameters 1,048,576 tokens vllm

Model · Image and text to text

DeepSeek-V4.1-Flash

DeepSeek

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

DeepSeek-V4.1-Flash-UNCENSORED-FP8

SAIFI INDUSTRIES

with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence. Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py, no runtime hooks, no steering vectors — it's a standard checkpoint that loads exactly like the base model. The refusal circuitry is surgically removed while every capability-critical component (routed experts, Engram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved byte-identical to the base. Every response 4-tier graded (HARDREF / SOFTRED / HEDGE /…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

s

Lautaro Rodriguez

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Access requested at publisher mit 763.2B parameters transformers

Model · Image and text to text

Synin-V1.1-Flash

Synin AI Lab

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Open weights mit 763.2B parameters 1,048,576 tokens transformers