Open-weight model
GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF
by Neuralll neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF
GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF is an open-weight model from Neuralll, released under MIT License. Its published files total 113.6 GB.
A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF (3.0-bit) made for faster single-stream decoding of this 117 GB model on consumer GPUs with an expert cache.
Model Card
By Neuralll, published under mit, revision ccdb8315ec08.
A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF (3.0-bit) made for faster single-stream decoding of this 117 GB model on consumer GPUs with an expert cache. Only the non-expert Q80 weights changed (attention, shared experts, dense FFN, output head: Q80 to Q4K, ~8.3 GB of the file). All 126 routed-expert tensors, tokenembd and every F32/BF16 tensor are copied bit-for-bit from the original. Those Q80 weights are streamed by the GPU for every token, so shrinking them speeds up decoding; the experts are unchanged. Stock llama.cpp can't load GLM-5.3-Flash yet. Use neurall/llama.cpp (GLM-5.3-Flash support plus a VRAM-filling MoE expert cache). It adds real GPU acceleration: on 2 GPUs, 2x the…
Read Neuralll's full model card
GLM-5.3-Flash GSQ-RCO 3.0-bit, Q4_K attention (GGUF)
Big thanks to @csantiago78: the expert cache here builds on their implementation in llama.cpp PR #27861 ("GPU-resident LRU cache for host-offloaded MoE expert weights"), the first to get a working hot-expert cache into llama.cpp. This fork takes it further: filling all free VRAM, smarter eviction, CPU/GPU overlap and fused kernels.
A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF (3.0-bit) made for faster single-stream decoding of this 117 GB model on consumer GPUs with an expert cache.
Only the non-expert Q8_0 weights changed (attention, shared experts, dense FFN,
output head: Q8_0 to Q4_K, ~8.3 GB of the file). All 126 routed-expert tensors,
token_embd and every F32/BF16 tensor are copied bit-for-bit from the original.
Those Q8_0 weights are streamed by the GPU for every token, so shrinking them
speeds up decoding; the experts are unchanged.
Requires a llama.cpp fork
Stock llama.cpp can't load GLM-5.3-Flash yet. Use neurall/llama.cpp (GLM-5.3-Flash support plus a VRAM-filling MoE expert cache). It adds real GPU acceleration: on 2 GPUs, 2x the tokens per second of stock llama.cpp (12.3 to 26-28 t/s):
llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf \
-np 1 -c 1024 -t 6 --cpu-moe -nr --moe-expert-cache -1
Set -c explicitly (without it autofit grows the context and takes the VRAM the
cache needs).
Why the fork matters: stock llama.cpp splits a model across GPUs by layer, so for a single request only one GPU works at a time: GPU1 idles while GPU2 runs its layers, and both idle while the CPU computes the experts that didn't fit in VRAM. A second GPU adds memory, not parallel compute. This fork turns that memory into a live cache of the experts actually being used, so most expert work runs on the GPUs, and in parallel with the CPU computing the rest. On 2x RTX 3090: 12.3 t/s stock vs 26-28 t/s with the fork, over 2x. The file also loads on any build with GLM-5.3-Flash support (PRs #27773 / #27917), at stock speed.
Same VRAM, different use. Stock llama.cpp and this fork get the same 48 GB; what differs is what it holds. Each token uses only 8 of the 288 experts in each layer.
- Stock places experts statically, whole layers at a time: ~36 GB fits all 288 experts of ~14 of the 42 MoE layers. Most of that VRAM holds experts the current token doesn't touch, so only ~33% of each token's expert work runs on GPU and the CPU does ~67%, one after the other.
- This fork fills the same VRAM with the ~100 most-used experts of every layer. Usage is skewed, so those cover ~85% of what tokens actually pick: ~85% of expert work runs on GPU and the CPU does ~15%, at the same time as the GPUs.
| expert work on GPU | expert work on CPU | decode t/s | |
|---|---|---|---|
| stock (static whole layers) | ~33% | ~67% | 12.3 |
| this fork (cache of hot experts) | ~85% | ~15%, in parallel | 26-28 |
So stock can't reach 2x on the same hardware: without an expert cache, extra VRAM mostly holds experts that aren't being used.
Results
2x RTX 3090 (48 GB VRAM) + 125 GB RAM, 8-core CPU, single stream. Decode: prompt "generate smallest html tetris game.", 1024 context, temperature 0. Perplexity: wikitext-2 test, 40 x 512-token chunks, same build for both files.
| decode t/s | PPL | |
|---|---|---|
| original 3.0-bit, stock llama.cpp (autofit) | 12.3 | 3.5534 |
| original 3.0-bit, fork with expert cache | ~25 | 3.5534 |
| this file, fork with expert cache | 27.72 | 3.5871 (+0.95%) |
Note: the original quant's RCO allocation chose each tensor's precision on purpose; overriding the Q8_0 tensors to Q4_K trades ~1% perplexity for ~10% speed. If you want the RCO allocation as designed, use the original file.
How it was made
llama-quantize --allow-requantize --tensor-type-file tensor-types-q4kattn.txt \
GLM-5.3-Flash-GSQ-RCO-3.0bit.gguf GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf Q4_K
tensor-types-q4kattn.txt (in this repo) lists every tensor with its target type:
Q8_0 non-expert tensors (except token_embd) to q4_k, everything else its
original type. Tensors whose type doesn't change are copied, not requantized.
Credits and license
- Base model: zai-org/GLM-5.3-Flash, MIT.
- Original quantization: pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF, a community reproduction of the GSQ and RCO methods by IST-DASLab (GSQ, RCO). Not an IST-DASLab release.
- Expert cache: builds on llama.cpp PR #27861 (csantiago78).
- GLM-5.3-Flash support: llama.cpp PRs #27773 and #27917 (timkhronos).
MIT license, see LICENSE.
Identity and Version
- Repository
- neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF
- Publisher
- Neuralll
- Task
- Not stated by the source
- Modality
- Other
- Library
- gguf
- Parameters
- Not stated by the source
- Languages
- moe
- Revision
- ccdb8315ec08cb747bb6acbd5bff370f1ad94c84
- First published
- 2026-09-24
- Last updated
- 2026-09-24
Files and Weights
5 files, 113.6 GB in total. The weights are 1 file totalling 113.6 GB in gguf.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf | Weights | 113.6 GB | 5f03b74ffe75 |
| LICENSE | Documentation | 1.1 KB | — |
| README.md | Documentation | 5.3 KB | — |
| tensor-types-q4kattn.txt | Other | 47.4 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 113.6 GB
Released by Neuralll through its official repository on Hugging Face. Read the license.
Built From
- Derived from pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF
- Described by arXiv:2604.18556
- Described by arXiv:2605.00649
- Quantized from pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 113.6 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF
Can I use GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF commercially?
Yes. GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.