K2-Horizon-MoVA-36B-A4B — NanoSeedLM p4mx-q4v (22.76 GB)
An experimental NanoSeedLM (SeedLM) compression of
IFM/K2-Horizon-MoVA-36B-A4B for Apple Silicon Macs with 32 GB
of unified memory.
SeedLM keeps each block of 8 weights as a 16-bit seed of a linear feedback shift
register (LFSR), 4 coefficients and an exponent. The GPU makes the weights again from the seed at run time. The paper
used an FPGA for this. This model uses Metal kernels on the Mac GPU.
This is a research proof of concept. It is not a replacement for a calibrated quantization.
This upload
- Six safetensors shards (22.76 GB in total, at most 4.5 GB each) and
model.safetensors.index.json. The other files
are config.json, the tokenizer files of the base model and nanoseedlm_k2.py, the MLX loader.
- Routed FFN experts (100 per layer, layers 3-47): SeedLM P=4, 4.5 bits per weight. Data-free LFSR basis,
activation-weighted seed search over all 65,535 seeds.
- Value experts (MoVA, 64 per layer): affine Q4, group 64.
- Attention, shared experts, dense layers 0-2, embedding, LM head: affine Q8, group 64.
- Routers and norms: BF16.
- Seed tensors are stored as
NAME.seeds (U16), NAME.coefs (U16), NAME.codes (U8) and NAME.exp_bias (I32). Q8 and
Q4 tensors use MLX's layout (NAME.weight, NAME.scales, NAME.biases).
Results
Measured on one M4 Max (128 GB). KLD is against MLX BF16 logits (held-out text / chat).
| Model |
Size |
KLD held-out / chat |
Decode, C engine, 1k / 4k ctx |
| BF16 |
74.89 GB |
reference |
35.8 / 34.0 tok/s |
| Affine Q4 experts, Q8 rest |
26.53 GB |
0.0233 / 0.0110 |
60.2 / 55.9 tok/s |
| SeedLM P=4 experts, Q8 rest (p4mx) |
26.53 GB |
0.0195 / 0.0095 |
49.5 / 46.5 tok/s |
| This model (p4mx-q4v) |
22.76 GB |
0.0260 / 0.0126 |
50.9 / 48.0 tok/s |
- GPU memory in the C engine: 23.3 GB at 1k context, 23.9 GB at 4k. This fits the default GPU memory of a 32 GB Mac
to approximately 4k tokens of context.
- Generations: four blind judges compared 80 outputs against BF16. 27 of 28 checkable answers were correct. Two outputs
(one prompt, greedy and sampled) fell into a loop. The 26.53 GB p4mx model passes all checks; this smaller model
does not.
- In oMLX on the same Mac: decode 41 tok/s, against 52 tok/s for the Q4 model in the same path.
Use
NanoSeedLM engine (C99 and Metal)
git clone https://github.com/mbarnson/nanoseedlm && cd nanoseedlm && make
hf download txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v --local-dir K2-NSLM
out/bin/nslm-chat --model K2-NSLM "Why do tide pools matter?"
out/bin/nslm-serve --model K2-NSLM --port 8080 # OpenAI-compatible API with tool calling
oMLX
Put the folder in the oMLX model directory, or download it to the Hugging Face cache. Enable Trust Remote Code
for the model: config.json names nanoseedlm_k2.py (the MLX loader with the seed kernels) in model_file. oMLX
supplies the K2-Horizon model code.
Settings
The chat template has three reasoning levels: high (default), medium and low. Stop tokens are <|ifm|endoftext|> and
<|ifm|im_end|>. Native context is 524,288 tokens. On a 32 GB Mac, the KV cache limits context to approximately
4k tokens.
Why this model
- Mixture of Values. MoVA routes the attention value projection through 64 experts. Together with 100 FFN experts,
expert matrices hold most of the weights. Seeds go there.
- Open. IFM publishes the training data, the training code and the intermediate checkpoints.
- A quiet corner. The K2-Horizon family has few users. Experiments here do not disturb popular models.
How it was made
The full story is in NanoSeedLM. In short:
- All weights of K2-Horizon-0.9B as 4.0-bit seeds: fail.
- GPU seed search and fast decode kernels: search 25x faster.
- MoVA expert gate and up projections as seeds: outputs hold, KLD worse than Q4.
- All routed experts as seeds, the rest Q8: 24.87 GB, two failures.
- A C99 and Metal engine that matches MLX.
- Kernel work: seeds at 8-11% behind Q4 decode.
- 4.5-bit blocks with 4 coefficients: Q4 size, lower KLD than Q4.
- Value experts in Q4: 22.76 GB. This model.
Limits
- Decode is 14-21% slower than the affine Q4 model (C engine and oMLX).
- Prefill in oMLX decodes one layer's experts to BF16 for a short time: approximately 2 GB more peak memory.
- Calibrated quantizations (for example imatrix K-quants) are strong. On small dense models they beat seeds.
- The seed tensors run in the NanoSeedLM engine and in MLX through
nanoseedlm_k2.py (oMLX). Transformers cannot run
them.
License and attribution
The base model is licensed under Apache-2.0 by IFM. This model is a modified version: the weights are compressed as
described above. The license is in LICENSE.
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/}
}
@inproceedings{shafipour2025seedlm,
title = {SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators},
author = {Shafipour, Rasoul and Harrison, David and Horton, Maxwell and Marker, Jeffrey and Bedayat, Houman and
Mehta, Sachin and Rastegari, Mohammad and Najibi, Mahyar and Naderiparizi, Saman},
booktitle = {International Conference on Learning Representations},
year = {2025}
}