A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF (3.0-bit) made for faster single-stream decoding of this 117 GB model on consumer GPUs with an expert cache. Only the non-expert Q80 weights changed (attention, shared experts, dense FFN, output head: Q80 to Q4K, ~8.3 GB of the file). All 126 routed-expert tensors, tokenembd and every F32/BF16 tensor are copied bit-for-bit from the original. Those Q80 weights are streamed by the GPU for every token, so shrinking them speeds up decoding; the experts are unchanged. Stock llama.cpp can't load GLM-5.3-Flash yet. Use neurall/llama.cpp (GLM-5.3-Flash support plus a VRAM-filling MoE expert cache). It adds real GPU acceleration: on 2 GPUs, 2x the…
Independent publisher
Neuralll
neuralll
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—