The weights use IQ4NL mixed-precision quantization. The quantized GGUF file is approximately 18 GB and can run on a single consumer-grade GPU. Xing4.0-29B-A4B is a next-generation large language model in the Xing series (formerly TeleChat), developed by China Telecom Artificial Intelligence Technology Co., Ltd. With 29B total parameters and only 4B activated per token, it natively supports a 256K context length, extensible to 512K. It is the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework, and deeply optimized for complex engineering tasks. For more information, please refer to our GitHub repository. Xing4.0-29B-A4B can be accessed via an…