Kimi K3 on a single NVIDIA A100 80GB. A weight-only quantisation of Moonshot AI's Kimi K3 (2.8T total / 104B activated parameters) that loads and generates on one A100 80GB GPU, with the routed experts held in host RAM. No Ampere-targeted K3 build existed for vLLM, so this was made to fix that gap. Routed experts and attention re-encoded from MXFP4/BF16 into compressed-tensors pack-quantized, served by vLLM's Marlin kernels. Activations stay BF16 (W4A16 / W8A16). Round-to-nearest only — no calibration data, so no dataset is baked into these weights. Errors were measured by round-tripping each tensor through compressed-tensors' compress()/decompress(). Weight error is a proxy, not a quality…
Independent publisher
Aaron Beckley
Fluffy
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—