This repository contains a text-only Mini-K3-1H v2 pretraining checkpoint from a controlled 20-architecture comparison. The family retains Kimi-K3's KDA and Gated MLA operators, block Attention Residuals, Stable LatentMoE, SiTU activations, output gates, and Quantile Balancing at approximately one billion logical parameters. The exact architecture for this repository is listed below; some ablations deliberately replace the baseline KDA/MLA ratio, decay granularity, convolution length, or positional encoding. - Hidden width / attention heads / KDA head width: 1024 / 12 / 128 - Vocabulary / BOS / generation EOS / PAD: 163840 / 163584 / 163586 / 163839 control state retained in FP32 where…
Independent publisher
Nkkbr
nkkbr
Multimodal Large Language Model, Vision Language Model
Models in Library1
Datasets in Library0
Models on Hugging Face184
Followers5