SAVRN
Search Contact SAVRN

Open-weight model

GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF

by Neuralll neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF

GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF is an open-weight model from Neuralll, released under MIT License. Its published files total 113.6 GB.

A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF (3.0-bit) made for faster single-stream decoding of this 117 GB model on consumer GPUs with an expert cache.

Parameters—
Context—
Weights113.6 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Model Card

By Neuralll, published under mit, revision ccdb8315ec08.

A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF (3.0-bit) made for faster single-stream decoding of this 117 GB model on consumer GPUs with an expert cache. Only the non-expert Q80 weights changed (attention, shared experts, dense FFN, output head: Q80 to Q4K, ~8.3 GB of the file). All 126 routed-expert tensors, tokenembd and every F32/BF16 tensor are copied bit-for-bit from the original. Those Q80 weights are streamed by the GPU for every token, so shrinking them speeds up decoding; the experts are unchanged. Stock llama.cpp can't load GLM-5.3-Flash yet. Use neurall/llama.cpp (GLM-5.3-Flash support plus a VRAM-filling MoE expert cache). It adds real GPU acceleration: on 2 GPUs, 2x the…

Read Neuralll's full model card

GLM-5.3-Flash GSQ-RCO 3.0-bit, Q4_K attention (GGUF)

Big thanks to @csantiago78: the expert cache here builds on their implementation in llama.cpp PR #27861 ("GPU-resident LRU cache for host-offloaded MoE expert weights"), the first to get a working hot-expert cache into llama.cpp. This fork takes it further: filling all free VRAM, smarter eviction, CPU/GPU overlap and fused kernels.

A variant of pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF (3.0-bit) made for faster single-stream decoding of this 117 GB model on consumer GPUs with an expert cache.

Only the non-expert Q8_0 weights changed (attention, shared experts, dense FFN, output head: Q8_0 to Q4_K, ~8.3 GB of the file). All 126 routed-expert tensors, token_embd and every F32/BF16 tensor are copied bit-for-bit from the original. Those Q8_0 weights are streamed by the GPU for every token, so shrinking them speeds up decoding; the experts are unchanged.

Requires a llama.cpp fork

Stock llama.cpp can't load GLM-5.3-Flash yet. Use neurall/llama.cpp (GLM-5.3-Flash support plus a VRAM-filling MoE expert cache). It adds real GPU acceleration: on 2 GPUs, 2x the tokens per second of stock llama.cpp (12.3 to 26-28 t/s):

llama-server -m GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf \
    -np 1 -c 1024 -t 6 --cpu-moe -nr --moe-expert-cache -1

Set -c explicitly (without it autofit grows the context and takes the VRAM the cache needs).

Why the fork matters: stock llama.cpp splits a model across GPUs by layer, so for a single request only one GPU works at a time: GPU1 idles while GPU2 runs its layers, and both idle while the CPU computes the experts that didn't fit in VRAM. A second GPU adds memory, not parallel compute. This fork turns that memory into a live cache of the experts actually being used, so most expert work runs on the GPUs, and in parallel with the CPU computing the rest. On 2x RTX 3090: 12.3 t/s stock vs 26-28 t/s with the fork, over 2x. The file also loads on any build with GLM-5.3-Flash support (PRs #27773 / #27917), at stock speed.

Same VRAM, different use. Stock llama.cpp and this fork get the same 48 GB; what differs is what it holds. Each token uses only 8 of the 288 experts in each layer.

  • Stock places experts statically, whole layers at a time: ~36 GB fits all 288 experts of ~14 of the 42 MoE layers. Most of that VRAM holds experts the current token doesn't touch, so only ~33% of each token's expert work runs on GPU and the CPU does ~67%, one after the other.
  • This fork fills the same VRAM with the ~100 most-used experts of every layer. Usage is skewed, so those cover ~85% of what tokens actually pick: ~85% of expert work runs on GPU and the CPU does ~15%, at the same time as the GPUs.
expert work on GPU expert work on CPU decode t/s
stock (static whole layers) ~33% ~67% 12.3
this fork (cache of hot experts) ~85% ~15%, in parallel 26-28

So stock can't reach 2x on the same hardware: without an expert cache, extra VRAM mostly holds experts that aren't being used.

Results

2x RTX 3090 (48 GB VRAM) + 125 GB RAM, 8-core CPU, single stream. Decode: prompt "generate smallest html tetris game.", 1024 context, temperature 0. Perplexity: wikitext-2 test, 40 x 512-token chunks, same build for both files.

decode t/s PPL
original 3.0-bit, stock llama.cpp (autofit) 12.3 3.5534
original 3.0-bit, fork with expert cache ~25 3.5534
this file, fork with expert cache 27.72 3.5871 (+0.95%)

Note: the original quant's RCO allocation chose each tensor's precision on purpose; overriding the Q8_0 tensors to Q4_K trades ~1% perplexity for ~10% speed. If you want the RCO allocation as designed, use the original file.

How it was made

llama-quantize --allow-requantize --tensor-type-file tensor-types-q4kattn.txt \
    GLM-5.3-Flash-GSQ-RCO-3.0bit.gguf GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.gguf Q4_K

tensor-types-q4kattn.txt (in this repo) lists every tensor with its target type: Q8_0 non-expert tensors (except token_embd) to q4_k, everything else its original type. Tensors whose type doesn't change are copied, not requantized.

Credits and license

MIT license, see LICENSE.

Identity and Version

Repository
neuralll/GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF
Publisher
Neuralll
Task
Not stated by the source
Modality
Other
Library
gguf
Parameters
Not stated by the source
Languages
moe
Revision
ccdb8315ec08cb747bb6acbd5bff370f1ad94c84
First published
2026-09-24
Last updated
2026-09-24

Files and Weights

5 files, 113.6 GB in total. The weights are 1 file totalling 113.6 GB in gguf.

Weights1 file · 113.6 GB
Documentation2 files · 6.4 KB
Other1 file · 47.4 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
GLM-5.3-Flash-GSQ-RCO-3.0bit-q4kattn.ggufWeights113.6 GB 5f03b74ffe75
LICENSEDocumentation1.1 KB —
README.mdDocumentation5.3 KB —
tensor-types-q4kattn.txtOther47.4 KB —
.gitattributesRepository1.6 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
113.6 GB
Download from Neuralll

Released by Neuralll through its official repository on Hugging Face. Read the license.

Built From

  • Derived from pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF
  • Described by arXiv:2604.18556
  • Described by arXiv:2605.00649
  • Quantized from pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF

Memory Requirements

PrecisionWeights in memory
As published113.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF

Can I use GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF commercially?

Yes. GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.