Quantized version of https://huggingface.co/Qwen/Qwen3.8-27B
Search public pages, research tools, and SAVRN solutions.
Open-weight model · Image and text to text
by Jarrel Seah jarrelscy/GLM-5.3-Flash-NestQuant-1.5-4bit
GLM-5.3-Flash-NestQuant-1.5-4bit is an open-weight model for image and text to text from Jarrel Seah, released under MIT License. It has 16.9B parameters and a 1,048,576-token context. At 16-bit it needs about 40.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
GLM-5.3-Flash with its routed experts quantized to NestQuant, a nested two-level format. Each expert has a 1.5-bit base plus a residual plane that lifts it to 4 bits.
What it takes to serve GLM-5.3-Flash-NestQuant-1.5-4bit (16.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 33.8 GB | 40.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 16.9 GB | 20.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 8.5 GB | 10.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
GLM-5.3-Flash-NestQuant-1.5-4bit on every accelerator the SAVRN Index prices, at every precision
By Jarrel Seah, published under mit, revision e0d0aa652d33.
GLM-5.3-Flash with its routed experts quantized to NestQuant, a nested two-level format. Each expert has a 1.5-bit base plus a residual plane that lifts it to 4 bits. At serving time an expert moves from 1.5 bit to 4 bit by loading its residual plane on top of the base, and the base bytes stay the same. This build targets a 128 GB Mac with 96 GB usable for weights, KV cache and runtime. The native vision tower is included, so the model still takes images. Status: TODO. Encoding in progress. This repo cannot be loaded with stock transformers, vLLM or MLX. The serving kernel and loader will be published separately. - 67 of 288 experts per layer (23%) run at 4 bit. The rest run at 1.5 bit.…
GLM-5.3-Flash with its routed experts quantized to NestQuant, a nested two-level format. Each expert has a 1.5-bit base plus a residual plane that lifts it to 4 bits. At serving time an expert moves from 1.5 bit to 4 bit by loading its residual plane on top of the base, and the base bytes stay the same.
This build targets a 128 GB Mac with 96 GB usable for weights, KV cache and runtime. The native vision tower is included, so the model still takes images.
Status: TODO. Encoding in progress. This repo cannot be loaded with stock transformers, vLLM or MLX. The serving kernel and loader will be published separately.
| Part | Format | Size in memory |
|---|---|---|
| Routed experts, layers 3-44 (42 x 288), 1.5-bit base | NestQuant base | ~58 GB (TODO: measured) |
| 4-bit residual planes for 67 experts per layer (19 fixed + 48 floating) | NestQuant residual | ~23 GB (TODO: measured) |
| Backbone: attention, dense MLP, shared experts, router, norms, mHC, embeddings, lm_head | FP8 at load (bf16 linears converted) | ~10 GB (15.5 GB as shipped) |
| Vision tower + merger | bf16 | 1.1 GB |
| KV cache at 128K | ~0.9 GB |
serving/predictor/ picks them ahead of routing.| File | What | Size |
|---|---|---|
layers/L{L}/tp{0..7}.safetensors, L = 3..44 |
Expert planes of layer L (1.5-bit base and 4-bit residual), split for tensor parallel 8 | TODO per layer |
layers/L{L}/manifest.json |
Layout of layer L, the fixed set and the floating default | |
nonexpert-0000{1..4}-of-00004.safetensors |
Everything that is not a routed expert: attention (Kimi KDA linear attention, plus DSA/MLA every 4th layer), mHC hyper-connections, dense MLP (layers 0-2), shared experts, router gates and e_score_correction_bias, norms, embeddings, lm_head, and the MTP layer 45 apart from its routed experts. Copied byte for byte from zai-org/GLM-5.3-Flash (FP8 where the source is FP8, bf16/fp32 elsewhere). |
15.5 GB |
vision_tower.safetensors |
The native vision tower and merger (model.visual.*), bf16, byte for byte |
1.1 GB |
mtp_experts-0000{1,2}-of-00002.safetensors |
The 288 routed experts of the MTP layer 45, FP8, byte for byte | 7.2 GB |
model.safetensors.index.json |
Index of all the above (3,532 tensors) | |
nonexpert_manifest.json, nonexpert_tensor_sha256.json |
Per-file and per-tensor sha256 | |
fixed_set.json |
The 19 fixed experts per layer, with their scores | TODO |
serving/predictor/ |
Floating-set predictor, trained on Flash routing | TODO |
config.json, generation_config.json, tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json |
From zai-org/GLM-5.3-Flash. config.json adds a quantization_config.nestquant block; all other keys are unchanged. |
The MTP layer is kept as the source FP8, as in the GLM-5.3 NestQuant releases. Its routed experts (7.25B
parameters, 7.2 GB) are in separate files. They are too large for a 96 GB Mac, so the Mac setup runs without MTP and
does not load mtp_experts-*. Machines with more memory can load them for speculative decoding.
TODO. KL divergence against the BF16 model's full logits on held-out windows, at the 96 GB setup (19 fixed + 48 floating):
| Setup | 4-bit experts per layer | KLD |
|---|---|---|
| FP8 reference | all | TODO |
| 1.5-4 bit (this repo), 96 GB | 67 | TODO |
nonexpert_tensor_sha256.json). None of the 72,576 routed-expert tensors of layers 3-44
are included.393 files, 176.8 GB in total. The weights are 321 files totalling 176.8 GB in pt, safetensors.
| File | Type | Size | SHA-256 |
|---|---|---|---|
| layers/L10/tp0.safetensors | Weights | 489.8 MB | d8bc1c11a410 |
| layers/L10/tp1.safetensors | Weights | 489.8 MB | ea65da8f0100 |
| layers/L10/tp2.safetensors | Weights | 489.8 MB | efa934d7543c |
| layers/L10/tp3.safetensors | Weights | 489.8 MB | edd7b5345c8a |
| layers/L10/tp4.safetensors | Weights | 489.8 MB | fbfa3453f417 |
| layers/L10/tp5.safetensors | Weights | 489.8 MB | 217412c8ac2c |
| layers/L10/tp6.safetensors | Weights | 489.8 MB | 573f787a1d4a |
| layers/L10/tp7.safetensors | Weights | 489.8 MB | c4e9a09bd836 |
| layers/L11/tp0.safetensors | Weights | 489.6 MB | fc30c4559bad |
| layers/L11/tp1.safetensors | Weights | 489.6 MB | 1d05fbdab2fe |
| layers/L11/tp2.safetensors | Weights | 489.6 MB | 7ec2a4ed3a4d |
| layers/L11/tp3.safetensors | Weights | 489.6 MB | 433268e04aa4 |
| layers/L11/tp4.safetensors | Weights | 489.6 MB | c8df323549a1 |
| layers/L11/tp5.safetensors | Weights | 489.6 MB | 0503eae05c75 |
| layers/L11/tp6.safetensors | Weights | 489.6 MB | 7ab3808eb33c |
| layers/L11/tp7.safetensors | Weights | 489.6 MB | f88307d36e7e |
| layers/L12/tp0.safetensors | Weights | 490.4 MB | d4417f3ee12e |
| layers/L12/tp1.safetensors | Weights | 490.4 MB | dcc6f0353a2a |
| layers/L12/tp2.safetensors | Weights | 490.4 MB | c82fee428422 |
| layers/L12/tp3.safetensors | Weights | 490.4 MB | 7a3e5d4b81b8 |
| layers/L12/tp4.safetensors | Weights | 490.4 MB | 2569d7d5cb82 |
| layers/L12/tp5.safetensors | Weights | 490.4 MB | 86523521b9fe |
| layers/L12/tp6.safetensors | Weights | 490.4 MB | 037286d6608a |
| layers/L12/tp7.safetensors | Weights | 490.4 MB | e749b6aa0d32 |
| layers/L13/tp0.safetensors | Weights | 491.5 MB | e1576628fed2 |
| layers/L13/tp1.safetensors | Weights | 491.5 MB | bbc29a0036b9 |
| layers/L13/tp2.safetensors | Weights | 491.5 MB | cb0b7b9a9f79 |
| layers/L13/tp3.safetensors | Weights | 491.5 MB | e29c707e5a40 |
| layers/L13/tp4.safetensors | Weights | 491.5 MB | 51b153e7bdf9 |
| layers/L13/tp5.safetensors | Weights | 491.5 MB | d83922ccdc61 |
| layers/L13/tp6.safetensors | Weights | 491.5 MB | e9b50b364600 |
| layers/L13/tp7.safetensors | Weights | 491.5 MB | ef44f3d3b69d |
| layers/L14/tp0.safetensors | Weights | 490.6 MB | 21eb313cc54f |
| layers/L14/tp1.safetensors | Weights | 490.6 MB | e2db698fb45d |
| layers/L14/tp2.safetensors | Weights | 490.6 MB | 90d8e0f95805 |
| layers/L14/tp3.safetensors | Weights | 490.6 MB | cebfde350f04 |
| layers/L14/tp4.safetensors | Weights | 490.6 MB | 0db02648fdf7 |
| layers/L14/tp5.safetensors | Weights | 490.6 MB | a0b8830c9105 |
| layers/L14/tp6.safetensors | Weights | 490.6 MB | cad990f66b44 |
| layers/L14/tp7.safetensors | Weights | 490.6 MB | 5f6de26e42be |
| layers/L15/tp0.safetensors | Weights | 490.4 MB | 330d82108a4d |
| layers/L15/tp1.safetensors | Weights | 490.4 MB | d20c7b1c896b |
| layers/L15/tp2.safetensors | Weights | 490.4 MB | 3a20e0e7a146 |
| layers/L15/tp3.safetensors | Weights | 490.4 MB | 7579a83c3dd6 |
| layers/L15/tp4.safetensors | Weights | 490.4 MB | 4998271493b4 |
| layers/L15/tp5.safetensors | Weights | 490.4 MB | 1f37781b900f |
| layers/L15/tp6.safetensors | Weights | 490.4 MB | a00c8773755d |
| layers/L15/tp7.safetensors | Weights | 490.4 MB | 6885c7efd8a2 |
| layers/L16/tp0.safetensors | Weights | 490.4 MB | 0606289459a2 |
| layers/L16/tp1.safetensors | Weights | 490.4 MB | 15e8693d48e2 |
| layers/L16/tp2.safetensors | Weights | 490.4 MB | c0392912bc9e |
| layers/L16/tp3.safetensors | Weights | 490.4 MB | eb7e9f6630e5 |
| layers/L16/tp4.safetensors | Weights | 490.4 MB | 1225ce57e5b9 |
| layers/L16/tp5.safetensors | Weights | 490.4 MB | 5003700242e6 |
| layers/L16/tp6.safetensors | Weights | 490.4 MB | 7e2857b625d9 |
| layers/L16/tp7.safetensors | Weights | 490.4 MB | f4ef21cdedde |
| layers/L17/tp0.safetensors | Weights | 490.3 MB | 8ec1503dceea |
| layers/L17/tp1.safetensors | Weights | 490.3 MB | 6773c0679794 |
| layers/L17/tp2.safetensors | Weights | 490.3 MB | 035e8b472d58 |
| layers/L17/tp3.safetensors | Weights | 490.3 MB | 1440c888b32e |
| layers/L17/tp4.safetensors | Weights | 490.3 MB | facdbdf311c6 |
| layers/L17/tp5.safetensors | Weights | 490.3 MB | 496e75579565 |
| layers/L17/tp6.safetensors | Weights | 490.3 MB | 165ebadd96b3 |
| layers/L17/tp7.safetensors | Weights | 490.3 MB | 1856fb66621c |
| layers/L18/tp0.safetensors | Weights | 491.0 MB | 26a5886750ef |
| layers/L18/tp1.safetensors | Weights | 491.0 MB | 4a1b1febfe15 |
| layers/L18/tp2.safetensors | Weights | 491.0 MB | bad5f3923f09 |
| layers/L18/tp3.safetensors | Weights | 491.0 MB | 505a871f37b6 |
| layers/L18/tp4.safetensors | Weights | 491.0 MB | 75f0dcac5093 |
| layers/L18/tp5.safetensors | Weights | 491.0 MB | 228225450d53 |
| layers/L18/tp6.safetensors | Weights | 491.0 MB | 868b818436d5 |
| layers/L18/tp7.safetensors | Weights | 491.0 MB | 81c4b8ad6b6e |
| layers/L19/tp0.safetensors | Weights | 490.3 MB | 03bbc31fb306 |
| layers/L19/tp1.safetensors | Weights | 490.3 MB | d9c214f4e211 |
| layers/L19/tp2.safetensors | Weights | 490.3 MB | f47f6d2100a1 |
| layers/L19/tp3.safetensors | Weights | 490.3 MB | d5baf876f42f |
| layers/L19/tp4.safetensors | Weights | 490.3 MB | 31dbc8e3c482 |
| layers/L19/tp5.safetensors | Weights | 490.3 MB | 7d4934a954ce |
| layers/L19/tp6.safetensors | Weights | 490.3 MB | 084dfbe19800 |
| layers/L19/tp7.safetensors | Weights | 490.3 MB | 35a74d67fecf |
| layers/L20/tp0.safetensors | Weights | 490.1 MB | 7cccb83665f3 |
| layers/L20/tp1.safetensors | Weights | 490.1 MB | 2a432b520256 |
| layers/L20/tp2.safetensors | Weights | 490.1 MB | 0147e278bfe7 |
| layers/L20/tp3.safetensors | Weights | 490.1 MB | 5723797201a8 |
| layers/L20/tp4.safetensors | Weights | 490.1 MB | c56c237244cc |
| layers/L20/tp5.safetensors | Weights | 490.1 MB | b87bc268f488 |
| layers/L20/tp6.safetensors | Weights | 490.1 MB | e389189fec83 |
| layers/L20/tp7.safetensors | Weights | 490.1 MB | d25d6860c978 |
| layers/L21/tp0.safetensors | Weights | 490.3 MB | 325c2aac0218 |
| layers/L21/tp1.safetensors | Weights | 490.3 MB | 3827a34d74f7 |
| layers/L21/tp2.safetensors | Weights | 490.3 MB | c41d6154ec27 |
| layers/L21/tp3.safetensors | Weights | 490.3 MB | 3e229f4967ea |
| layers/L21/tp4.safetensors | Weights | 490.3 MB | 84415f8c09f6 |
| layers/L21/tp5.safetensors | Weights | 490.3 MB | f1a145564b13 |
| layers/L21/tp6.safetensors | Weights | 490.3 MB | afe444e09ec5 |
| layers/L21/tp7.safetensors | Weights | 490.3 MB | fa04282cf8dd |
| layers/L22/tp0.safetensors | Weights | 489.7 MB | 78dee494d41c |
| layers/L22/tp1.safetensors | Weights | 489.7 MB | 216269f40326 |
| layers/L22/tp2.safetensors | Weights | 489.7 MB | ba5bf470b053 |
| layers/L22/tp3.safetensors | Weights | 489.7 MB | 39ec64b97851 |
| layers/L22/tp4.safetensors | Weights | 489.7 MB | 49a440b6e604 |
| layers/L22/tp5.safetensors | Weights | 489.7 MB | ca14fb5d559e |
| layers/L22/tp6.safetensors | Weights | 489.7 MB | 152cd00783b3 |
| layers/L22/tp7.safetensors | Weights | 489.7 MB | de1e500890ec |
| layers/L23/tp0.safetensors | Weights | 489.7 MB | 5e663f5447b4 |
| layers/L23/tp1.safetensors | Weights | 489.7 MB | 164a81f0ef27 |
| layers/L23/tp2.safetensors | Weights | 489.7 MB | 5075045f364f |
| layers/L23/tp3.safetensors | Weights | 489.7 MB | 78e188966077 |
| layers/L23/tp4.safetensors | Weights | 489.7 MB | 06d5290d84b7 |
| layers/L23/tp5.safetensors | Weights | 489.7 MB | 9650044f64d0 |
| layers/L23/tp6.safetensors | Weights | 489.7 MB | e720198c4373 |
| layers/L23/tp7.safetensors | Weights | 489.7 MB | e1c6570bbc3d |
| layers/L24/tp0.safetensors | Weights | 489.9 MB | 2c45cfda6dec |
| layers/L24/tp1.safetensors | Weights | 489.9 MB | b8f0605c703d |
| layers/L24/tp2.safetensors | Weights | 489.9 MB | 30c85bb86900 |
| layers/L24/tp3.safetensors | Weights | 489.9 MB | 48eb7fe5f9a8 |
| layers/L24/tp4.safetensors | Weights | 489.9 MB | 296f299d0af4 |
| layers/L24/tp5.safetensors | Weights | 489.9 MB | f0303fadd05a |
| layers/L24/tp6.safetensors | Weights | 489.9 MB | d7242ba65711 |
| layers/L24/tp7.safetensors | Weights | 489.9 MB | d6d02579470c |
| layers/L25/tp0.safetensors | Weights | 489.5 MB | 18ba78a53c5f |
| layers/L25/tp1.safetensors | Weights | 489.5 MB | 00ea664b7079 |
| layers/L25/tp2.safetensors | Weights | 489.5 MB | 2c38e305cca3 |
| layers/L25/tp3.safetensors | Weights | 489.5 MB | 9175d0e8bad5 |
| layers/L25/tp4.safetensors | Weights | 489.5 MB | acdf7d3ef0d0 |
| layers/L25/tp5.safetensors | Weights | 489.5 MB | d73661ec4944 |
| layers/L25/tp6.safetensors | Weights | 489.5 MB | e45b485eb1f6 |
| layers/L25/tp7.safetensors | Weights | 489.5 MB | eb12def583fe |
| layers/L26/tp0.safetensors | Weights | 489.4 MB | 39ce4e028b2c |
| layers/L26/tp1.safetensors | Weights | 489.4 MB | b8b040184259 |
| layers/L26/tp2.safetensors | Weights | 489.4 MB | cdb209007091 |
| layers/L26/tp3.safetensors | Weights | 489.4 MB | 72f0df9091e1 |
| layers/L26/tp4.safetensors | Weights | 489.4 MB | dd0fee174f98 |
| layers/L26/tp5.safetensors | Weights | 489.4 MB | fbecc06d3639 |
| layers/L26/tp6.safetensors | Weights | 489.4 MB | 48de8daed0b1 |
| layers/L26/tp7.safetensors | Weights | 489.4 MB | 9eae9d68ae89 |
| layers/L27/tp0.safetensors | Weights | 489.4 MB | 66263729cc0c |
| layers/L27/tp1.safetensors | Weights | 489.4 MB | 7feb48f3a511 |
| layers/L27/tp2.safetensors | Weights | 489.4 MB | de19feca1c5e |
| layers/L27/tp3.safetensors | Weights | 489.4 MB | 534ad0bb7a74 |
| layers/L27/tp4.safetensors | Weights | 489.4 MB | a250e6061b41 |
| layers/L27/tp5.safetensors | Weights | 489.4 MB | 4d6c87919a6a |
| layers/L27/tp6.safetensors | Weights | 489.4 MB | 137ed6360bb7 |
| layers/L27/tp7.safetensors | Weights | 489.4 MB | 79f0c05ca9d8 |
| layers/L28/tp0.safetensors | Weights | 489.4 MB | 5f160ef2a5ce |
| layers/L28/tp1.safetensors | Weights | 489.4 MB | e1edb67cdce5 |
| layers/L28/tp2.safetensors | Weights | 489.4 MB | 36b4285b4ab4 |
| layers/L28/tp3.safetensors | Weights | 489.4 MB | 22bae3f0a865 |
| layers/L28/tp4.safetensors | Weights | 489.4 MB | bd857a550bb6 |
| layers/L28/tp5.safetensors | Weights | 489.4 MB | f57db4819af7 |
| layers/L28/tp6.safetensors | Weights | 489.4 MB | f9814acf53d8 |
| layers/L28/tp7.safetensors | Weights | 489.4 MB | c23699dcf42f |
| layers/L29/tp0.safetensors | Weights | 489.8 MB | 1656b3bf38ea |
| layers/L29/tp1.safetensors | Weights | 489.8 MB | 5190c9afdeaa |
| layers/L29/tp2.safetensors | Weights | 489.8 MB | 3dc7e7a887d7 |
| layers/L29/tp3.safetensors | Weights | 489.8 MB | 5b3b6d2a3d43 |
| layers/L29/tp4.safetensors | Weights | 489.8 MB | 93f3e74f5427 |
| layers/L29/tp5.safetensors | Weights | 489.8 MB | 2a05e5bad199 |
| layers/L29/tp6.safetensors | Weights | 489.8 MB | c57b9628469d |
| layers/L29/tp7.safetensors | Weights | 489.8 MB | 5031f9511777 |
| layers/L3/tp0.safetensors | Weights | 495.6 MB | 530caf95403e |
| layers/L3/tp1.safetensors | Weights | 495.6 MB | 84d5ea9a44de |
| layers/L3/tp2.safetensors | Weights | 495.6 MB | 867326577890 |
| layers/L3/tp3.safetensors | Weights | 495.6 MB | 363e17109be9 |
| layers/L3/tp4.safetensors | Weights | 495.6 MB | 484b13bdd458 |
| layers/L3/tp5.safetensors | Weights | 495.6 MB | 2cf22c267217 |
| layers/L3/tp6.safetensors | Weights | 495.6 MB | 3c8df8e307d2 |
| layers/L3/tp7.safetensors | Weights | 495.6 MB | 5591d24443e4 |
| layers/L30/tp0.safetensors | Weights | 489.5 MB | 04f3db971b07 |
| layers/L30/tp1.safetensors | Weights | 489.5 MB | 40a1dad07a72 |
| layers/L30/tp2.safetensors | Weights | 489.5 MB | d219609f90ca |
| layers/L30/tp3.safetensors | Weights | 489.5 MB | 33855d3a03df |
| layers/L30/tp4.safetensors | Weights | 489.5 MB | bf580eaf3b75 |
| layers/L30/tp5.safetensors | Weights | 489.5 MB | 9f7ab2cabd7c |
| layers/L30/tp6.safetensors | Weights | 489.5 MB | 55ae1c50d967 |
| layers/L30/tp7.safetensors | Weights | 489.5 MB | aeb39b57578e |
| layers/L31/tp0.safetensors | Weights | 489.3 MB | 96d723e6e95d |
| layers/L31/tp1.safetensors | Weights | 489.3 MB | 867ad8b5ddff |
| layers/L31/tp2.safetensors | Weights | 489.3 MB | 203beb373ef4 |
| layers/L31/tp3.safetensors | Weights | 489.3 MB | f1ed86393200 |
| layers/L31/tp4.safetensors | Weights | 489.3 MB | 6e172c294cad |
| layers/L31/tp5.safetensors | Weights | 489.3 MB | ab7a354df76b |
| layers/L31/tp6.safetensors | Weights | 489.3 MB | 203ae951e345 |
| layers/L31/tp7.safetensors | Weights | 489.3 MB | 314d2e444bd4 |
| layers/L32/tp0.safetensors | Weights | 489.5 MB | ba4e97138fbd |
| layers/L32/tp1.safetensors | Weights | 489.5 MB | e7eec356243f |
| layers/L32/tp2.safetensors | Weights | 489.5 MB | b8e0fe9d4ebc |
| layers/L32/tp3.safetensors | Weights | 489.5 MB | 1037c9a7e161 |
| layers/L32/tp4.safetensors | Weights | 489.5 MB | b3062db88412 |
| layers/L32/tp5.safetensors | Weights | 489.5 MB | ec3adc44d4b7 |
| layers/L32/tp6.safetensors | Weights | 489.5 MB | 357e16cc06f5 |
| layers/L32/tp7.safetensors | Weights | 489.5 MB | daaeedd492c3 |
| layers/L33/tp0.safetensors | Weights | 489.5 MB | fc6917093537 |
| layers/L33/tp1.safetensors | Weights | 489.5 MB | 944d9f79683b |
| layers/L33/tp2.safetensors | Weights | 489.5 MB | 00e6d529756a |
| layers/L33/tp3.safetensors | Weights | 489.5 MB | 0581b1f54190 |
| layers/L33/tp4.safetensors | Weights | 489.5 MB | afc262112081 |
| layers/L33/tp5.safetensors | Weights | 489.5 MB | 99ef7d91761c |
| layers/L33/tp6.safetensors | Weights | 489.5 MB | 2cf797c464e1 |
| layers/L33/tp7.safetensors | Weights | 489.5 MB | 21e84329a9bf |
| layers/L34/tp0.safetensors | Weights | 489.5 MB | 3d73e99b7370 |
| layers/L34/tp1.safetensors | Weights | 489.5 MB | 6d1dd5a09074 |
| layers/L34/tp2.safetensors | Weights | 489.5 MB | 07ca8b7f53a8 |
| layers/L34/tp3.safetensors | Weights | 489.5 MB | 695dbf1164d0 |
| layers/L34/tp4.safetensors | Weights | 489.5 MB | 9e9915ba8e6a |
| layers/L34/tp5.safetensors | Weights | 489.5 MB | ba69061276f7 |
| layers/L34/tp6.safetensors | Weights | 489.5 MB | c8632513413b |
| layers/L34/tp7.safetensors | Weights | 489.5 MB | 9629e3b11610 |
| layers/L35/tp0.safetensors | Weights | 489.5 MB | b67ec8230b30 |
| layers/L35/tp1.safetensors | Weights | 489.5 MB | 49130d2a55c0 |
| layers/L35/tp2.safetensors | Weights | 489.5 MB | 1723d0a59225 |
| layers/L35/tp3.safetensors | Weights | 489.5 MB | 0efeaa2fe6f4 |
| layers/L35/tp4.safetensors | Weights | 489.5 MB | 798a95df39ed |
| layers/L35/tp5.safetensors | Weights | 489.5 MB | fc90b5a4d0d6 |
| layers/L35/tp6.safetensors | Weights | 489.5 MB | 7293a1e30718 |
| layers/L35/tp7.safetensors | Weights | 489.5 MB | 176db27dd4cf |
| layers/L36/tp0.safetensors | Weights | 489.4 MB | 50ac10366943 |
| layers/L36/tp1.safetensors | Weights | 489.4 MB | 01b49da9a365 |
| layers/L36/tp2.safetensors | Weights | 489.4 MB | 4f0fa4c64e68 |
| layers/L36/tp3.safetensors | Weights | 489.4 MB | 397a93e86de3 |
| layers/L36/tp4.safetensors | Weights | 489.4 MB | 7febc5da0013 |
| layers/L36/tp5.safetensors | Weights | 489.4 MB | 804bf8d34e55 |
| layers/L36/tp6.safetensors | Weights | 489.4 MB | 3c34b9b0f9f6 |
| layers/L36/tp7.safetensors | Weights | 489.4 MB | 95a52d36c899 |
| layers/L37/tp0.safetensors | Weights | 489.9 MB | 09cfa502d3b1 |
| layers/L37/tp1.safetensors | Weights | 489.9 MB | 1d7366e9d2f2 |
| layers/L37/tp2.safetensors | Weights | 489.9 MB | 2fd22335b64f |
| layers/L37/tp3.safetensors | Weights | 489.9 MB | a05756408999 |
| layers/L37/tp4.safetensors | Weights | 489.9 MB | be93107d0304 |
| layers/L37/tp5.safetensors | Weights | 489.9 MB | 63b8100ed09a |
| layers/L37/tp6.safetensors | Weights | 489.9 MB | 46ff9e102ca7 |
| layers/L37/tp7.safetensors | Weights | 489.9 MB | 9a91881d9f10 |
| layers/L38/tp0.safetensors | Weights | 489.8 MB | 1fcfd9e54bed |
| layers/L38/tp1.safetensors | Weights | 489.8 MB | 4ec556a81a6c |
| layers/L38/tp2.safetensors | Weights | 489.8 MB | 29884f00576c |
| layers/L38/tp3.safetensors | Weights | 489.8 MB | f0119c155ff1 |
| layers/L38/tp4.safetensors | Weights | 489.8 MB | 6be0e14692bc |
| layers/L38/tp5.safetensors | Weights | 489.8 MB | c562e486b7e6 |
| layers/L38/tp6.safetensors | Weights | 489.8 MB | 3feec42f7af7 |
| layers/L38/tp7.safetensors | Weights | 489.8 MB | ee71273548a7 |
| layers/L39/tp0.safetensors | Weights | 489.9 MB | 86cca04087c5 |
| layers/L39/tp1.safetensors | Weights | 489.9 MB | a6cb69b639ff |
| layers/L39/tp2.safetensors | Weights | 489.9 MB | e483b9a40ee2 |
| layers/L39/tp3.safetensors | Weights | 489.9 MB | 118fda9ab204 |
| layers/L39/tp4.safetensors | Weights | 489.9 MB | d0cbcf19d61f |
| layers/L39/tp5.safetensors | Weights | 489.9 MB | b090382c959a |
| layers/L39/tp6.safetensors | Weights | 489.9 MB | a9aa922d4203 |
| layers/L39/tp7.safetensors | Weights | 489.9 MB | 91448feee427 |
| layers/L4/tp0.safetensors | Weights | 489.5 MB | c38a698a0aaa |
| layers/L4/tp1.safetensors | Weights | 489.5 MB | a1f9140beeae |
| layers/L4/tp2.safetensors | Weights | 489.5 MB | e71cb614a938 |
| layers/L4/tp3.safetensors | Weights | 489.5 MB | e3b5b61490b8 |
| layers/L4/tp4.safetensors | Weights | 489.5 MB | 062682389e36 |
| layers/L4/tp5.safetensors | Weights | 489.5 MB | 9736f1e9bdb3 |
| layers/L4/tp6.safetensors | Weights | 489.5 MB | 8598cd7032db |
| layers/L4/tp7.safetensors | Weights | 489.5 MB | 3ef79233661f |
| layers/L40/tp0.safetensors | Weights | 491.2 MB | d3c8e65f57eb |
| layers/L40/tp1.safetensors | Weights | 491.2 MB | 7075c603d2cf |
| layers/L40/tp2.safetensors | Weights | 491.2 MB | 41e152282796 |
| layers/L40/tp3.safetensors | Weights | 491.2 MB | 1d24e16fe880 |
| layers/L40/tp4.safetensors | Weights | 491.2 MB | 72588c0850b4 |
| layers/L40/tp5.safetensors | Weights | 491.2 MB | 2beccdc62126 |
| layers/L40/tp6.safetensors | Weights | 491.2 MB | 6fb2d859d131 |
| layers/L40/tp7.safetensors | Weights | 491.2 MB | eb0d4cae4980 |
| layers/L41/tp0.safetensors | Weights | 493.9 MB | 55ff46f3dd90 |
| layers/L41/tp1.safetensors | Weights | 493.9 MB | f50128114b8d |
| layers/L41/tp2.safetensors | Weights | 493.9 MB | bfcf68e51fdb |
| layers/L41/tp3.safetensors | Weights | 493.9 MB | 9a5af93f207f |
| layers/L41/tp4.safetensors | Weights | 493.9 MB | 837fc6810f93 |
| layers/L41/tp5.safetensors | Weights | 493.9 MB | 298b439d22a3 |
| layers/L41/tp6.safetensors | Weights | 493.9 MB | fd7ce3e3f5a8 |
| layers/L41/tp7.safetensors | Weights | 493.9 MB | 8fe91d0912cf |
| layers/L5/tp0.safetensors | Weights | 489.3 MB | b3515163f6d2 |
| layers/L5/tp1.safetensors | Weights | 489.3 MB | d812187a2ba6 |
| layers/L5/tp2.safetensors | Weights | 489.3 MB | 23668a7e1e32 |
| layers/L5/tp3.safetensors | Weights | 489.3 MB | de28e64fcda0 |
| layers/L5/tp4.safetensors | Weights | 489.3 MB | 151423c84b20 |
| layers/L5/tp5.safetensors | Weights | 489.3 MB | 7d79d9a9f91a |
| layers/L5/tp6.safetensors | Weights | 489.3 MB | 68ec9d67a8d8 |
| layers/L5/tp7.safetensors | Weights | 489.3 MB | dcff95721731 |
| layers/L6/tp0.safetensors | Weights | 489.4 MB | 264f194f4b83 |
| layers/L6/tp1.safetensors | Weights | 489.4 MB | 7ecac3627105 |
| layers/L6/tp2.safetensors | Weights | 489.4 MB | baef7d07264c |
| layers/L6/tp3.safetensors | Weights | 489.4 MB | d2f1b6701013 |
| layers/L6/tp4.safetensors | Weights | 489.4 MB | a4de4dd56e5b |
| layers/L6/tp5.safetensors | Weights | 489.4 MB | fe680cc61056 |
| layers/L6/tp6.safetensors | Weights | 489.4 MB | c8cd28228ba5 |
| layers/L6/tp7.safetensors | Weights | 489.4 MB | f9a3a9d17819 |
| layers/L7/tp0.safetensors | Weights | 490.0 MB | 27bc76c0a927 |
| layers/L7/tp1.safetensors | Weights | 490.0 MB | ead23ab4038f |
| layers/L7/tp2.safetensors | Weights | 490.0 MB | 377e43f41c5f |
| layers/L7/tp3.safetensors | Weights | 490.0 MB | eb0ddee209c2 |
| layers/L7/tp4.safetensors | Weights | 490.0 MB | 24b96484812d |
| layers/L7/tp5.safetensors | Weights | 490.0 MB | d49bc05e5c43 |
| layers/L7/tp6.safetensors | Weights | 490.0 MB | 9cfaf234fdcb |
| layers/L7/tp7.safetensors | Weights | 490.0 MB | 7ef0dff39b64 |
| layers/L8/tp0.safetensors | Weights | 490.2 MB | ae25d602168f |
| layers/L8/tp1.safetensors | Weights | 490.2 MB | 42bec7dc8fbb |
| layers/L8/tp2.safetensors | Weights | 490.2 MB | 55230d46aba1 |
| layers/L8/tp3.safetensors | Weights | 490.2 MB | f8b44d40795e |
| layers/L8/tp4.safetensors | Weights | 490.2 MB | ac76f16c034c |
| layers/L8/tp5.safetensors | Weights | 490.2 MB | eb02d603cb98 |
| layers/L8/tp6.safetensors | Weights | 490.2 MB | bf70548f7faf |
| layers/L8/tp7.safetensors | Weights | 490.2 MB | cd3694d0368c |
| layers/L9/tp0.safetensors | Weights | 489.7 MB | 4b20460c86b5 |
| layers/L9/tp1.safetensors | Weights | 489.7 MB | fccb681374fe |
| layers/L9/tp2.safetensors | Weights | 489.7 MB | e72738c07454 |
| layers/L9/tp3.safetensors | Weights | 489.7 MB | 9080ec0b37ce |
| layers/L9/tp4.safetensors | Weights | 489.7 MB | 8913a55df0d5 |
| layers/L9/tp5.safetensors | Weights | 489.7 MB | 9327273b52ec |
| layers/L9/tp6.safetensors | Weights | 489.7 MB | 8b325c5a5afa |
| layers/L9/tp7.safetensors | Weights | 489.7 MB | 5d41eb8b6a4e |
| mtp_experts-00001-of-00002.safetensors | Weights | 5.0 GB | 32defe995062 |
| mtp_experts-00002-of-00002.safetensors | Weights | 2.3 GB | 28181b15ba7b |
| nonexpert-00001-of-00004.safetensors | Weights | 5.0 GB | 5c261c79d37d |
| nonexpert-00002-of-00004.safetensors | Weights | 5.0 GB | 7fc6df7f9832 |
| nonexpert-00003-of-00004.safetensors | Weights | 4.3 GB | ec23074b4c30 |
| nonexpert-00004-of-00004.safetensors | Weights | 1.3 GB | 2f1c6ac8756b |
| serving/predictor/joint/jF.pt | Weights | 5.3 MB | fe587d4b1b9f |
| serving/predictor_nf48/joint/jF.pt | Weights | 5.3 MB | 400613e5f0f0 |
| vision_tower.safetensors | Weights | 1.1 GB | e9f5cbb64bdd |
| config.json | Configuration | 70.9 KB | — |
| fixed_set.json | Configuration | 835.4 KB | — |
| generation_config.json | Configuration | 194 B | — |
| layers/L10/manifest.json | Configuration | 364.1 KB | — |
| layers/L11/manifest.json | Configuration | 363.5 KB | — |
| layers/L12/manifest.json | Configuration | 366.0 KB | — |
| layers/L13/manifest.json | Configuration | 369.4 KB | — |
| layers/L14/manifest.json | Configuration | 366.6 KB | — |
| layers/L15/manifest.json | Configuration | 365.8 KB | — |
| layers/L16/manifest.json | Configuration | 365.7 KB | — |
| layers/L17/manifest.json | Configuration | 365.6 KB | — |
| layers/L18/manifest.json | Configuration | 367.6 KB | — |
| layers/L19/manifest.json | Configuration | 365.7 KB | — |
| layers/L20/manifest.json | Configuration | 364.9 KB | — |
| layers/L21/manifest.json | Configuration | 365.3 KB | — |
| layers/L22/manifest.json | Configuration | 363.6 KB | — |
| layers/L23/manifest.json | Configuration | 363.9 KB | — |
| layers/L24/manifest.json | Configuration | 364.4 KB | — |
| layers/L25/manifest.json | Configuration | 363.1 KB | — |
| layers/L26/manifest.json | Configuration | 362.7 KB | — |
| layers/L27/manifest.json | Configuration | 363.1 KB | — |
| layers/L28/manifest.json | Configuration | 362.7 KB | — |
| layers/L29/manifest.json | Configuration | 364.0 KB | — |
| layers/L3/manifest.json | Configuration | 376.1 KB | — |
| layers/L30/manifest.json | Configuration | 362.9 KB | — |
| layers/L31/manifest.json | Configuration | 362.6 KB | — |
| layers/L32/manifest.json | Configuration | 363.0 KB | — |
| layers/L33/manifest.json | Configuration | 363.2 KB | — |
| layers/L34/manifest.json | Configuration | 362.9 KB | — |
| layers/L35/manifest.json | Configuration | 363.0 KB | — |
| layers/L36/manifest.json | Configuration | 362.8 KB | — |
| layers/L37/manifest.json | Configuration | 364.4 KB | — |
| layers/L38/manifest.json | Configuration | 364.1 KB | — |
| layers/L39/manifest.json | Configuration | 364.0 KB | — |
| layers/L4/manifest.json | Configuration | 363.2 KB | — |
| layers/L40/manifest.json | Configuration | 368.3 KB | — |
| layers/L41/manifest.json | Configuration | 376.1 KB | — |
| layers/L5/manifest.json | Configuration | 362.6 KB | — |
| layers/L6/manifest.json | Configuration | 362.9 KB | — |
| layers/L7/manifest.json | Configuration | 364.4 KB | — |
| layers/L8/manifest.json | Configuration | 365.2 KB | — |
| layers/L9/manifest.json | Configuration | 363.6 KB | — |
| model.safetensors.index.json | Configuration | 379.7 KB | — |
| nonexpert_manifest.json | Configuration | 4.9 KB | — |
| nonexpert_tensor_sha256.json | Configuration | 465.1 KB | — |
| processor_config.json | Configuration | 909 B | — |
| serving/predictor/joint/gpu_predictor.py | Configuration | 14.0 KB | — |
| serving/predictor/joint/jF.json | Configuration | 10.8 KB | — |
| serving/predictor/joint/jlib.py | Configuration | 15.5 KB | — |
| serving/predictor/joint/joint_predictor.py | Configuration | 11.5 KB | — |
| serving/predictor/joint/parity_stream.py | Configuration | 7.5 KB | — |
| serving/predictor/joint/train.py | Configuration | 19.8 KB | — |
| serving/predictor/predictor.json | Configuration | 981 B | — |
| serving/predictor_nf48/joint/gpu_predictor.py | Configuration | 14.0 KB | — |
| serving/predictor_nf48/joint/jF.json | Configuration | 10.8 KB | — |
| serving/predictor_nf48/joint/jlib.py | Configuration | 15.5 KB | — |
| serving/predictor_nf48/joint/joint_predictor.py | Configuration | 11.5 KB | — |
| serving/predictor_nf48/joint/parity_stream.py | Configuration | 7.5 KB | — |
| serving/predictor_nf48/joint/train.py | Configuration | 19.8 KB | — |
| serving/predictor_nf48/predictor.json | Configuration | 981 B | — |
| LICENSE | Documentation | 1.1 KB | — |
| README.md | Documentation | 5.5 KB | — |
| serving/predictor/joint/README.md | Documentation | 3.4 KB | — |
| serving/predictor/joint/SERVE.md | Documentation | 2.3 KB | — |
| serving/predictor_nf48/joint/README.md | Documentation | 3.4 KB | — |
| serving/predictor_nf48/joint/SERVE.md | Documentation | 2.3 KB | — |
| chat_template.jinja | Other | 10.9 KB | — |
| serving/predictor/joint/v2_sal_tweedie1.5.txt | Other | 117.6 KB | — |
| serving/predictor_nf48/joint/v2_sal_tweedie1.5.txt | Other | 117.6 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.2 MB | 19e773648cb4 |
| tokenizer_config.json | Tokenizer | 761 B | — |
Released by Jarrel Seah through its official repository on Hugging Face. Read the license.
| Precision | Weights in memory |
|---|---|
| As published | 176.8 GB |
| 16-bit | 33.8 GB |
| 8-bit | 16.9 GB |
| 4-bit | 8.5 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
About 40.6 GB at 16-bit and 10.2 GB at 4-bit: the weights (16.9B parameters) plus a working margin. A long context needs more.
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Yes. GLM-5.3-Flash-NestQuant-1.5-4bit is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
1,048,576 tokens, from the maximum position embeddings in its published configuration.
Quantized version of https://huggingface.co/Qwen/Qwen3.8-27B
The RadixArk Qwen3.8-27B-NVFP4 model is the quantized version of Qwen/Qwen3.8-27B. The quantization was produced at RadixArk using NVIDIA Model Optimizer, following a mixed NVFP4 W4A4 recipe. Run on SGLang: launch command and per-platform recipes in the Qwen3.8-27B cookbook. This model is not owned or developed by RadixArk. It is a quantized derivative of Qwen's model; see the upstream Qwen3.8-27B model card for the source model's capabilities, training information, limitations, and license. Global Developers looking to deploy an off-the-shelf, pre-quantized model in AI agent systems, chatbots, RAG systems, and other AI-powered applications. Hugging Face 08/14/2026 via…
NVFP4 checkpoint of an abliterated Swift-Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and SGLang. GGUFs for llama.cpp: The Swift 1.5 version is source). - Swift's own NVFP4 recipe, unmodified, from ukisai/Swift-Qwen3.8-27B-NVFP4, calibrated with NVIDIA ModelOpt. - MTP head and vision tower in BF16, bit-identical to the source. 21.9 GB, NVIDIA ModelOpt mixed-precision format. Needs a vLLM with ModelOpt mixed-precision support (tested on 0.29.0). No --quantization flag. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default. Tested on an RTX 5090 (32 GB) with vLLM 0.29.0: NVFP4 layers on…
NVFP4 checkpoint of an abliterated Swift 1.5 Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and SGLang. GGUFs for llama.cpp: (measured on the BF16 source). - Swift's own NVFP4 recipe, unmodified, from ukisai/Swift-Qwen3.8-27B-NVFP4, calibrated with NVIDIA ModelOpt. The module split matches UkisAI's Swift 1.5 NVFP4 exactly. - MTP head and vision tower in BF16, bit-identical to the source. 21.9 GB, NVIDIA ModelOpt mixed-precision format. Needs a vLLM with ModelOpt mixed-precision support. No --quantization flag. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default. Same format, recipe, module…
Mixed-precision quantization of prepared for dnv2003/v100-skinny. - MLP gateproj, upproj, downproj, and lmhead: NVFP4, group size 16 87c9f8cf83021957d1a1a575c90c9a4eaaf7ef0c See quantization-audit.json and hfquantconfig.json for the complete machine-readable layout. This model has had safety alignment substantially removed. Use it only for lawful, controlled research and add appropriate safeguards before deployment.
quantized to OCP MXFP4 4-bit weights in the same Quark checkpoint container as The weights are plain RTN, not AWQ: see the note below. The 15 MTP tensors are kept in BF16; the radiance runtime loads them with RADIANCEQUARKBF16MTP=1. Why RTN instead of AWQ. The first build of this checkpoint used AWQ with the same smoothing recipe as AMD's release. The AWQ fold itself was mathematically consistent, but on this model a few layers converged to weights to 32-element MXFP4 blocks then destroyed those layers on real, outlier-carrying inputs (layer 7 output relative error ~324), and repairing the worst layers individually was not enough; the remaining smoothed layers still accumulated too much…