Versions below for what a v2 would have to clear. A derivative of the official nvidia/GLM-5.3-Flash-NVFP4 checkpoint in which the layers NVIDIA's release left in BF16 — the attention linears (KDA fused inproj/out projections, MLA q/kv/o, indexer wqb), the shared experts (gate/up/down), and lmhead — are quantized to W4A16 NVFP4 (weight-only 4-bit, group size 16). Routers, norms, embeddings, the vision tower, and the MTP layer (layers.45) stay BF16; the routed experts keep the stock W4A4 NVFP4 quantization. enjoying the incomplete This checkpoint does not boot on the stock vLLM image. It needs the kda-quant and mla-quant source overlays from the companion repository (plus glm5next-mtp-bf16…
Independent publisher
Tenhkspark
tenhkspark
DGX Spark
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—