APEX (Adaptive Precision for EXpert Models) quantizations of Kwaipilot/KAT-Coder-V2.5-Dev — Kwaipilot's Mixture-of-Experts model for agentic coding. Brought to you by is a quantization strategy for Mixture-of-Experts (MoE) models.
Model Card
By Mudler, published under apache-2.0, revision 4086d2602dd0.
APEX (Adaptive Precision for EXpert Models) quantizations of Kwaipilot/KAT-Coder-V2.5-Dev — Kwaipilot's Mixture-of-Experts model for agentic coding. Brought to you by is a quantization strategy for Mixture-of-Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention) and applies a layer-wise precision gradient — edge layers (first/last 5) get higher precision, middle layers compress more aggressively. I-variants use diverse imatrix calibration (chat, code, reasoning, tool-calling, agentic traces, Wikipedia). In MoE models the routed-expert FFN tensors dominate the weight budget but only ~8/256 experts activate per token, so APEX compresses middle-layer…
Read Mudler's full model card
Each donation = another big MoE quantized
I host 30+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory), enough for ~30-50B-class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20-100 per quant.
If APEX quants are useful to you, your support directly funds those bigger runs.
KAT-Coder-V2.5-Dev — APEX GGUF
APEX (Adaptive Precision for EXpert Models) quantizations of Kwaipilot/KAT-Coder-V2.5-Dev — Kwaipilot's Mixture-of-Experts model for agentic coding.
Brought to you by the LocalAI team | APEX Project | Technical Report
Available Files
| File | Profile | Best For |
|---|---|---|
| KAT-Coder-V2.5-Dev-APEX-I-Balanced.gguf | I-Balanced | Best overall — imatrix-enhanced |
| KAT-Coder-V2.5-Dev-APEX-I-Quality.gguf | I-Quality | Highest quality with imatrix |
| KAT-Coder-V2.5-Dev-APEX-Quality.gguf | Quality | Highest quality (no imatrix) |
| KAT-Coder-V2.5-Dev-APEX-Balanced.gguf | Balanced | General purpose |
| KAT-Coder-V2.5-Dev-APEX-I-Compact.gguf | I-Compact | Consumer GPUs, imatrix-enhanced |
| KAT-Coder-V2.5-Dev-APEX-Compact.gguf | Compact | Consumer GPUs |
| KAT-Coder-V2.5-Dev-APEX-I-Mini.gguf | I-Mini | Smallest viable, fastest inference |
What is APEX?
APEX is a quantization strategy for Mixture-of-Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention) and applies a layer-wise precision gradient — edge layers (first/last 5) get higher precision, middle layers compress more aggressively. I-variants use diverse imatrix calibration (chat, code, reasoning, tool-calling, agentic traces, Wikipedia).
In MoE models the routed-expert FFN tensors dominate the weight budget but only ~8/256 experts activate per token, so APEX compresses middle-layer experts hardest while preserving edge layers, attention, and the always-active shared expert.
See the APEX project for full details.
Architecture
- Model: KAT-Coder-V2.5-Dev (Qwen3_5MoeForConditionalGeneration)
- Layers: 40 · Experts: 256 routed + 1 shared (8 active per token)
- Attention: 16 heads / 2 KV, hybrid (full attention every 4th layer)
- Calibration: v1.3 diverse dataset
Note: the config advertises an image token, but the released checkpoint ships no vision encoder weights, so these are text-only GGUFs (no mmproj).
Run with LocalAI
local-ai run mudler/KAT-Coder-V2.5-Dev-APEX-GGUF@KAT-Coder-V2.5-Dev-APEX-I-Balanced.gguf
Credits
APEX is brought to you by the LocalAI team. Built on llama.cpp. Base model by Kwaipilot.
Identity and Version
- Repository
- mudler/KAT-Coder-V2.5-Dev-APEX-GGUF
- Publisher
- Mudler
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- Not stated by the source
- Languages
- en, zh
- Revision
- 4086d2602dd09c003cbb4052178f03e4685a2f89
- First published
- 2026-07-24
- Last updated
- 2026-08-17
Files and Weights
9 files, 142.7 GB in total. The weights are 7 files totalling 142.7 GB in gguf.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| KAT-Coder-V2.5-Dev-APEX-Balanced.gguf | Weights | 25.3 GB | 9a5d31b110a9 |
| KAT-Coder-V2.5-Dev-APEX-Compact.gguf | Weights | 16.5 GB | 98dfee53102b |
| KAT-Coder-V2.5-Dev-APEX-I-Balanced.gguf | Weights | 25.3 GB | ee6e0ec15964 |
| KAT-Coder-V2.5-Dev-APEX-I-Compact.gguf | Weights | 16.5 GB | 5235ac39e798 |
| KAT-Coder-V2.5-Dev-APEX-I-Mini.gguf | Weights | 13.5 GB | 9d901bf1c449 |
| KAT-Coder-V2.5-Dev-APEX-I-Quality.gguf | Weights | 22.8 GB | e1cf7f33e13e |
| KAT-Coder-V2.5-Dev-APEX-Quality.gguf | Weights | 22.8 GB | 849737baf591 |
| README.md | Documentation | 4.0 KB | — |
| .gitattributes | Repository | 2.1 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 142.7 GB
Released by Mudler through its official repository on Hugging Face. Read the license.
Built From
- Derived from Kwaipilot/KAT-Coder-V2.5-Dev
- Quantized from Kwaipilot/KAT-Coder-V2.5-Dev
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 142.7 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About KAT-Coder-V2.5-Dev-APEX-GGUF
Can I use KAT-Coder-V2.5-Dev-APEX-GGUF commercially?
Yes. KAT-Coder-V2.5-Dev-APEX-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.