SAVRN
Search Contact SAVRN

Independent publisher

Hung Cheung Chan

pt810

Models in Library1
Datasets in Library0
Models on Hugging Face9
Followers—

Models

Experimental calibration-based weight-only GPTQ variant of The Thinker transformer uses 4-bit weights for layers 0–26 and 8-bit weights for layer 27. Audio, vision, talker, token2wav, embeddings, and norms remain in the original precision. Calibration used eight short text samples with LLM Compressor 0.14.0 and compressed-tensors 0.19.0. This checkpoint includes a packaging repair: the compressor export contained invalid group scales, so scales were recomputed from the original BF16 weights per group before the vLLM test. Treat this as an experimental GPTQ-derived checkpoint and benchmark retrieval quality before production use. Tested with vLLM 0.30.0 on an 8-GiB RTX 3080 Laptop GPU: The…

Open weights 2,935 parameters transformers