SAVRN
Search Contact SAVRN

Open-weight model · Text generation

K2-Horizon-375B-A23B

by Institute of Foundation Models IFM/K2-Horizon-375B-A23B

K2-Horizon-375B-A23B is an open-weight model for text generation from Institute of Foundation Models, released under Apache License 2.0. It has 379.2B parameters and a 524,288-token context. At 16-bit it needs about 910 GB of GPU memory, which fits on 4x MI325X from $8.00 an hour; at 4-bit, 227.5 GB on 1x MI325X from $2.00, at the lowest prices in the SAVRN Index. It draws 9k downloads a month.

K2-Horizon-375B-A23B is the flagship of the K2-Horizon family: a sparse Mixture-of-Experts model that stores 375B parameters and runs 23B per token, with a 512K context window.

Parameters379.2B
Context524,288
Weights758.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads9k

Runs On

What it takes to serve K2-Horizon-375B-A23B (379.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 758.3 GB 910.0 GB 4x MI325X (256 GB)
Vultr
$8.00 5x MI300X $9.25 · 4x MI355X $10.36
8-bit 379.2 GB 455.0 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55
4-bit 189.6 GB 227.5 GB 1x MI325X (256 GB)
Vultr
$2.00 1x MI355X $2.59 · 2x MI300X $3.70

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 23, 2026.

K2-Horizon-375B-A23B on every accelerator the SAVRN Index prices, at every precision

Model Card

By Institute of Foundation Models, published under apache-2.0, revision b55fa9affc85.

K2-Horizon-375B-A23B is the flagship of the K2-Horizon family: a sparse Mixture-of-Experts model that stores 375B parameters and runs 23B per token, with a 512K context window. We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released.

K2-Horizon-375B-A23B Highlights

  • Frontier-class agentic performance. On agentic tool use, terminal, and long-horizon workflow benchmarks it matches or beats open-weight MoE models up to 2.6× its size and is competitive with closed frontier models (see Benchmark Results).
  • 512K context. Native 524,288-token context from the midtraining stages onward.
  • Intermediate checkpoints. Intermediate checkpoints will be released so capability changes can be studied across training rather than at a single checkpoint.
  • Fully open. Training data/recipe and the training code will be made public.

Benchmark Results

Read the full model card (1,667 words)

Configuration

Architecture
K2HorizonForCausalLM
Context length (tokens)
524,288
Layers
61
Hidden size
6,144
Feed-forward size
16,384
Attention heads
48
Key/value heads
8
Head dimension
128
Vocabulary size
250,624
Experts
192
Experts active per token
8
Model type
k2_horizon

Identity and Version

Repository
IFM/K2-Horizon-375B-A23B
Publisher
Institute of Foundation Models
Task
Text generation
Modality
Text
Library
transformers
Parameters
379.2B parameters
Languages
en
Revision
b55fa9affc85fd11ec5f141ac863819a1f1d0cc9
First published
2026-09-01
Last updated
2026-09-23

Files and Weights

76 files, 758.4 GB in total. The weights are 61 files totalling 758.3 GB in safetensors.

Weights61 files · 758.3 GB
Configuration8 files · 3.1 MB
Tokenizer2 files · 29.2 MB
Documentation1 file · 49.1 KB
Other3 files · 1.6 MB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00061.safetensorsWeights780.2 MB 46a3a4798cae
model-00002-of-00061.safetensorsWeights780.2 MB 2f453ef706ba
model-00003-of-00061.safetensorsWeights780.2 MB 6e049caa67dd
model-00004-of-00061.safetensorsWeights12.9 GB 145c3df1fb55
model-00005-of-00061.safetensorsWeights12.9 GB 18eb1fa3153b
model-00006-of-00061.safetensorsWeights12.9 GB 74fc1103a0ea
model-00007-of-00061.safetensorsWeights12.9 GB 4a17e5428dc3
model-00008-of-00061.safetensorsWeights12.9 GB b7688833e747
model-00009-of-00061.safetensorsWeights12.9 GB 4107a4fe1fad
model-00010-of-00061.safetensorsWeights12.9 GB 6a53f1cd2969
model-00011-of-00061.safetensorsWeights12.9 GB 47950f553953
model-00012-of-00061.safetensorsWeights12.9 GB 565c01d77b1a
model-00013-of-00061.safetensorsWeights12.9 GB 26a14e92e122
model-00014-of-00061.safetensorsWeights12.9 GB eaf64e445652
model-00015-of-00061.safetensorsWeights12.9 GB c5b3516a67e3
model-00016-of-00061.safetensorsWeights12.9 GB dac5152478f9
model-00017-of-00061.safetensorsWeights12.9 GB f3ad09da47be
model-00018-of-00061.safetensorsWeights12.9 GB 48547c3050ba
model-00019-of-00061.safetensorsWeights12.9 GB 40df6daa3ca8
model-00020-of-00061.safetensorsWeights12.9 GB 411e7cf158a1
model-00021-of-00061.safetensorsWeights12.9 GB 202a9f3318cd
model-00022-of-00061.safetensorsWeights12.9 GB d55278742614
model-00023-of-00061.safetensorsWeights12.9 GB 53e63f907a9b
model-00024-of-00061.safetensorsWeights12.9 GB ecccaf226681
model-00025-of-00061.safetensorsWeights12.9 GB 3911247b805f
model-00026-of-00061.safetensorsWeights12.9 GB 87502e14ceb7
model-00027-of-00061.safetensorsWeights12.9 GB ead5f776c5a0
model-00028-of-00061.safetensorsWeights12.9 GB fcadacf04322
model-00029-of-00061.safetensorsWeights12.9 GB 9b15d0397ef6
model-00030-of-00061.safetensorsWeights12.9 GB f51c18a64e05
model-00031-of-00061.safetensorsWeights12.9 GB bb17f6e1e61e
model-00032-of-00061.safetensorsWeights12.9 GB 4d7258e143d3
model-00033-of-00061.safetensorsWeights12.9 GB e057796d35ac
model-00034-of-00061.safetensorsWeights12.9 GB bb62694c7810
model-00035-of-00061.safetensorsWeights12.9 GB c1ea36bb43aa
model-00036-of-00061.safetensorsWeights12.9 GB 18e93509072c
model-00037-of-00061.safetensorsWeights12.9 GB 35330b9861da
model-00038-of-00061.safetensorsWeights12.9 GB 4326b58af30a
model-00039-of-00061.safetensorsWeights12.9 GB 3a11f8928cb4
model-00040-of-00061.safetensorsWeights12.9 GB c094846af9a8
model-00041-of-00061.safetensorsWeights12.9 GB ff9d9d60b779
model-00042-of-00061.safetensorsWeights12.9 GB 66762c898816
model-00043-of-00061.safetensorsWeights12.9 GB 1d90f8fc24b2
model-00044-of-00061.safetensorsWeights12.9 GB 5bf95b695b65
model-00045-of-00061.safetensorsWeights12.9 GB 4cb35bf07b19
model-00046-of-00061.safetensorsWeights12.9 GB d30e887630f5
model-00047-of-00061.safetensorsWeights12.9 GB 4a1e32b0459e
model-00048-of-00061.safetensorsWeights12.9 GB 4f19c007fd57
model-00049-of-00061.safetensorsWeights12.9 GB a5dacb69ec66
model-00050-of-00061.safetensorsWeights12.9 GB 04ad142f7b7d
model-00051-of-00061.safetensorsWeights12.9 GB 0c0a9eb58af3
model-00052-of-00061.safetensorsWeights12.9 GB ef5c0c03dd9f
model-00053-of-00061.safetensorsWeights12.9 GB e0b40fc3b1b9
model-00054-of-00061.safetensorsWeights12.9 GB 77bb70c9b16b
model-00055-of-00061.safetensorsWeights12.9 GB 7ed4d7487438
model-00056-of-00061.safetensorsWeights12.9 GB 5252eb229fc0
model-00057-of-00061.safetensorsWeights12.9 GB 4c3b6983e0c9
model-00058-of-00061.safetensorsWeights12.9 GB 55764b8f3e5b
model-00059-of-00061.safetensorsWeights12.9 GB bfa82db7d7c8
model-00060-of-00061.safetensorsWeights12.9 GB 5faa6ae6c002
model-00061-of-00061.safetensorsWeights19.1 GB ecb06cb4715b
config.jsonConfiguration1.5 KB
configuration_k2_horizon.pyConfiguration3.5 KB
generation_config.jsonConfiguration81 B
migration_manifest.jsonConfiguration4.9 KB
model.safetensors.index.jsonConfiguration3.1 MB
modeling_k2_horizon.pyConfiguration49.5 KB
special_tokens_map.jsonConfiguration79 B
validation.jsonConfiguration957 B
README.mdDocumentation49.1 KB
assets/k2-horizon-375b-a23b-benchmarks.pngOther928.0 KB a48347a85c80
assets/k2-horizon-375b-training-loss-vs-tokens.pngOther586.7 KB e8a4c9d860a4
chat_template.jinjaOther51.0 KB
.gitattributesRepository1.7 KB
tokenizer.jsonTokenizer29.2 MB 2fa69519ff1e
tokenizer_config.jsonTokenizer253 B

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
758.3 GB
Download from Institute of Foundation Models

Released by Institute of Foundation Models through its official repository on Hugging Face. Read the license.

Built From

  • Trained on (disclosed) IFM/K2-Horizon-Midtrain-Data
  • Trained on (disclosed) IFM/K2-Horizon-Pretrain-Data

Memory Requirements

PrecisionWeights in memory
As published758.3 GB
16-bit758.3 GB
8-bit379.2 GB
4-bit189.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About K2-Horizon-375B-A23B

How much GPU memory does K2-Horizon-375B-A23B need?

About 910 GB at 16-bit and 227.5 GB at 4-bit: the weights (379.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run K2-Horizon-375B-A23B on?

At 16-bit, 4x MI325X from $8.00 an hour; at 4-bit, 1x MI325X from $2.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use K2-Horizon-375B-A23B commercially?

Yes. K2-Horizon-375B-A23B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is K2-Horizon-375B-A23B's context length?

524,288 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

GLM-5.2-NVFP4

NVIDIA

The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention (with an IndexShare indexer) to support a long context. For more information, please check here. The NVIDIA GLM-5.2 NVFP4 model is quantized with Model Optimizer. This model is ready for commercial or non-commercial use. GOVERNING TERMS: Use of the model is governed by the MIT License, same as the base model. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG…

Open weights mit 381B parameters 1,048,576 tokens Model Optimizer

Model · Text generation

Darwin-397B-JGOS

FINAL_Bench

Darwin-397B-JGOS is the largest and highest-scoring member of the Darwin family. Built on Qwen 3.5 397B as the base, it transplants the FFN (expert) strengths of multiple high-performance models through the Darwin V9 platform, producing a 397B-parameter Mixture-of-Experts model with ~17B active parameters per token. It reaches 90.9 % on GPQA Diamond with pure greedy decoding (single sample) — surpassing Darwin-28B-REASON (89.39 %, achieved with the Darwin-DELPHI test-time engine) without using any test-time engine at all. This is the highest GPQA Diamond score in the Darwin family to date. Darwin is VIDRAFT's measuring-result-driven reasoning model family — approximately 20 official models…

Open weights apache-2.0 403.4B parameters 262,144 tokens transformers

Model · Text generation

Darwin-397B-ZTC

FINAL_Bench

the footprint, GPQA Diamond 93.43 %. And this model stops itself before it acts on an answer it is about to get wrong. Darwin is VIDRAFT's measurement-driven reasoning model family — roughly 20 official models, 400+ community derivatives, and a standing place among the top open models on GPQA. A large MoE model is made of hundreds of experts. Darwin V9 selects the experts that perform best across several high-performing models, transplants them onto a base backbone, and fuses them with trust-weighted evolutionary merging. Nothing is trained from scratch — proven capability is grafted on. That is why the same method holds across every model size. - Darwin V9 — evolutionary FFN/expert…

Open weights apache-2.0 403.6B parameters 262,144 tokens transformers

Model · Text generation

DeepSeek-V4-Flash-0731

DeepSeek

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available. 1. For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort…

Open weights mit 304.2B parameters 1,048,576 tokens transformers

Model · Text generation

DeepSeek-V4-Flash

DeepSeek

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache…

Open weights mit 290.9B parameters 1,048,576 tokens transformers

Model · Text generation

Qwen3-Coder-480B-A35B-Instruct-FP8

Qwen

Today, we're announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we're excited to introduce its most powerful variant first: Qwen3-Coder-480B-A35B-Instruct. featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks, achieving results comparable to Claude Sonnet. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format.…

Open weights apache-2.0 480.2B parameters 262,144 tokens transformers