This repository provides an optimized FP8 (float8e4m3fn) weight-only quantized version of the newly released Krea 2 OSS (Turbo) transformer. This optimization reduces the model size from the original 24.76 GiB (BF16) down to 12.01 GiB, making it highly accessible and runnable on standard consumer hardware (such as 16GB and 24GB GPUs) without sacrificing output quality. Unlike generic global quantization scripts that aggressively convert every parameter (which often degrades generation details or introduces NaN/promotion calculation errors in neural networks), this model was quantized using a selective weight-only strategy: 1. Targeted Quantization: Only 2D floating-point weight matrices…
Independent publisher
Alper
AlperKTS
Hiiii
Models in Library1
Datasets in Library0
Models on Hugging Face2
Followers22