SAVRN
Search Contact SAVRN

Open-weight model

KAT-Coder-V2.5-Dev-APEX-GGUF

by Mudler mudler/KAT-Coder-V2.5-Dev-APEX-GGUF

APEX (Adaptive Precision for EXpert Models) quantizations of Kwaipilot/KAT-Coder-V2.5-Dev — Kwaipilot's Mixture-of-Experts model for agentic coding. Brought to you by is a quantization strategy for Mixture-of-Experts (MoE) models.

Parameters
Context
Weights142.7 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.4M

Model Card

By Mudler, published under apache-2.0, revision 4086d2602dd0.

APEX (Adaptive Precision for EXpert Models) quantizations of Kwaipilot/KAT-Coder-V2.5-Dev — Kwaipilot's Mixture-of-Experts model for agentic coding. Brought to you by is a quantization strategy for Mixture-of-Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention) and applies a layer-wise precision gradient — edge layers (first/last 5) get higher precision, middle layers compress more aggressively. I-variants use diverse imatrix calibration (chat, code, reasoning, tool-calling, agentic traces, Wikipedia). In MoE models the routed-expert FFN tensors dominate the weight budget but only ~8/256 experts activate per token, so APEX compresses middle-layer…

Read Mudler's full model card

Each donation = another big MoE quantized

I host 30+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory), enough for ~30-50B-class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20-100 per quant.
If APEX quants are useful to you, your support directly funds those bigger runs.

Patreon (Monthly)  |  Buy Me a Coffee  |  GitHub Sponsors

KAT-Coder-V2.5-Dev — APEX GGUF

APEX (Adaptive Precision for EXpert Models) quantizations of Kwaipilot/KAT-Coder-V2.5-Dev — Kwaipilot's Mixture-of-Experts model for agentic coding.

Brought to you by the LocalAI team | APEX Project | Technical Report

Available Files

File Profile Best For
KAT-Coder-V2.5-Dev-APEX-I-Balanced.gguf I-Balanced Best overall — imatrix-enhanced
KAT-Coder-V2.5-Dev-APEX-I-Quality.gguf I-Quality Highest quality with imatrix
KAT-Coder-V2.5-Dev-APEX-Quality.gguf Quality Highest quality (no imatrix)
KAT-Coder-V2.5-Dev-APEX-Balanced.gguf Balanced General purpose
KAT-Coder-V2.5-Dev-APEX-I-Compact.gguf I-Compact Consumer GPUs, imatrix-enhanced
KAT-Coder-V2.5-Dev-APEX-Compact.gguf Compact Consumer GPUs
KAT-Coder-V2.5-Dev-APEX-I-Mini.gguf I-Mini Smallest viable, fastest inference

What is APEX?

APEX is a quantization strategy for Mixture-of-Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention) and applies a layer-wise precision gradient — edge layers (first/last 5) get higher precision, middle layers compress more aggressively. I-variants use diverse imatrix calibration (chat, code, reasoning, tool-calling, agentic traces, Wikipedia).

In MoE models the routed-expert FFN tensors dominate the weight budget but only ~8/256 experts activate per token, so APEX compresses middle-layer experts hardest while preserving edge layers, attention, and the always-active shared expert.

See the APEX project for full details.

Architecture

  • Model: KAT-Coder-V2.5-Dev (Qwen3_5MoeForConditionalGeneration)
  • Layers: 40 · Experts: 256 routed + 1 shared (8 active per token)
  • Attention: 16 heads / 2 KV, hybrid (full attention every 4th layer)
  • Calibration: v1.3 diverse dataset

Note: the config advertises an image token, but the released checkpoint ships no vision encoder weights, so these are text-only GGUFs (no mmproj).

Run with LocalAI

local-ai run mudler/KAT-Coder-V2.5-Dev-APEX-GGUF@KAT-Coder-V2.5-Dev-APEX-I-Balanced.gguf

Credits

APEX is brought to you by the LocalAI team. Built on llama.cpp. Base model by Kwaipilot.

Identity and Version

Repository
mudler/KAT-Coder-V2.5-Dev-APEX-GGUF
Publisher
Mudler
Task
Not stated by the source
Modality
Other
Library
Not stated by the source
Parameters
Not stated by the source
Languages
en, zh
Revision
4086d2602dd09c003cbb4052178f03e4685a2f89
First published
2026-07-24
Last updated
2026-08-17

Files and Weights

9 files, 142.7 GB in total. The weights are 7 files totalling 142.7 GB in gguf.

Weights7 files · 142.7 GB
Documentation1 file · 4.0 KB
Repository1 file · 2.1 KB
Every file
FileTypeSizeSHA-256
KAT-Coder-V2.5-Dev-APEX-Balanced.ggufWeights25.3 GB 9a5d31b110a9
KAT-Coder-V2.5-Dev-APEX-Compact.ggufWeights16.5 GB 98dfee53102b
KAT-Coder-V2.5-Dev-APEX-I-Balanced.ggufWeights25.3 GB ee6e0ec15964
KAT-Coder-V2.5-Dev-APEX-I-Compact.ggufWeights16.5 GB 5235ac39e798
KAT-Coder-V2.5-Dev-APEX-I-Mini.ggufWeights13.5 GB 9d901bf1c449
KAT-Coder-V2.5-Dev-APEX-I-Quality.ggufWeights22.8 GB e1cf7f33e13e
KAT-Coder-V2.5-Dev-APEX-Quality.ggufWeights22.8 GB 849737baf591
README.mdDocumentation4.0 KB
.gitattributesRepository2.1 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
142.7 GB
Download from Mudler

Released by Mudler through its official repository on Hugging Face. Read the license.

Built From

  • Derived from Kwaipilot/KAT-Coder-V2.5-Dev
  • Quantized from Kwaipilot/KAT-Coder-V2.5-Dev

Memory Requirements

PrecisionWeights in memory
As published142.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About KAT-Coder-V2.5-Dev-APEX-GGUF

Can I use KAT-Coder-V2.5-Dev-APEX-GGUF commercially?

Yes. KAT-Coder-V2.5-Dev-APEX-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.