SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

GLM-5.3-Flash-NestQuant-1.5-4bit

by Jarrel Seah jarrelscy/GLM-5.3-Flash-NestQuant-1.5-4bit

GLM-5.3-Flash-NestQuant-1.5-4bit is an open-weight model for image and text to text from Jarrel Seah, released under MIT License. It has 16.9B parameters and a 1,048,576-token context. At 16-bit it needs about 40.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

GLM-5.3-Flash with its routed experts quantized to NestQuant, a nested two-level format. Each expert has a 1.5-bit base plus a residual plane that lifts it to 4 bits.

Parameters16.9B
Context1,048,576
Weights176.8 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve GLM-5.3-Flash-NestQuant-1.5-4bit (16.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 33.8 GB 40.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 16.9 GB 20.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 8.5 GB 10.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

GLM-5.3-Flash-NestQuant-1.5-4bit on every accelerator the SAVRN Index prices, at every precision

Model Card

By Jarrel Seah, published under mit, revision e0d0aa652d33.

GLM-5.3-Flash with its routed experts quantized to NestQuant, a nested two-level format. Each expert has a 1.5-bit base plus a residual plane that lifts it to 4 bits. At serving time an expert moves from 1.5 bit to 4 bit by loading its residual plane on top of the base, and the base bytes stay the same. This build targets a 128 GB Mac with 96 GB usable for weights, KV cache and runtime. The native vision tower is included, so the model still takes images. Status: TODO. Encoding in progress. This repo cannot be loaded with stock transformers, vLLM or MLX. The serving kernel and loader will be published separately. - 67 of 288 experts per layer (23%) run at 4 bit. The rest run at 1.5 bit.…

Read Jarrel Seah's full model card

GLM-5.3-Flash with its routed experts quantized to NestQuant, a nested two-level format. Each expert has a 1.5-bit base plus a residual plane that lifts it to 4 bits. At serving time an expert moves from 1.5 bit to 4 bit by loading its residual plane on top of the base, and the base bytes stay the same.

This build targets a 128 GB Mac with 96 GB usable for weights, KV cache and runtime. The native vision tower is included, so the model still takes images.

Status: TODO. Encoding in progress. This repo cannot be loaded with stock transformers, vLLM or MLX. The serving kernel and loader will be published separately.

Sizing at 96 GB

Part Format Size in memory
Routed experts, layers 3-44 (42 x 288), 1.5-bit base NestQuant base ~58 GB (TODO: measured)
4-bit residual planes for 67 experts per layer (19 fixed + 48 floating) NestQuant residual ~23 GB (TODO: measured)
Backbone: attention, dense MLP, shared experts, router, norms, mHC, embeddings, lm_head FP8 at load (bf16 linears converted) ~10 GB (15.5 GB as shipped)
Vision tower + merger bf16 1.1 GB
KV cache at 128K ~0.9 GB
  • 67 of 288 experts per layer (23%) run at 4 bit. The rest run at 1.5 bit.
  • 19 per layer are a fixed set that always stays at 4 bit.
  • 48 per layer are floating. The predictor in serving/predictor/ picks them ahead of routing.
  • The hot pool is sized at launch, so the same files run with more 4-bit experts on a larger machine.

Contents

File What Size
layers/L{L}/tp{0..7}.safetensors, L = 3..44 Expert planes of layer L (1.5-bit base and 4-bit residual), split for tensor parallel 8 TODO per layer
layers/L{L}/manifest.json Layout of layer L, the fixed set and the floating default
nonexpert-0000{1..4}-of-00004.safetensors Everything that is not a routed expert: attention (Kimi KDA linear attention, plus DSA/MLA every 4th layer), mHC hyper-connections, dense MLP (layers 0-2), shared experts, router gates and e_score_correction_bias, norms, embeddings, lm_head, and the MTP layer 45 apart from its routed experts. Copied byte for byte from zai-org/GLM-5.3-Flash (FP8 where the source is FP8, bf16/fp32 elsewhere). 15.5 GB
vision_tower.safetensors The native vision tower and merger (model.visual.*), bf16, byte for byte 1.1 GB
mtp_experts-0000{1,2}-of-00002.safetensors The 288 routed experts of the MTP layer 45, FP8, byte for byte 7.2 GB
model.safetensors.index.json Index of all the above (3,532 tensors)
nonexpert_manifest.json, nonexpert_tensor_sha256.json Per-file and per-tensor sha256
fixed_set.json The 19 fixed experts per layer, with their scores TODO
serving/predictor/ Floating-set predictor, trained on Flash routing TODO
config.json, generation_config.json, tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json From zai-org/GLM-5.3-Flash. config.json adds a quantization_config.nestquant block; all other keys are unchanged.

MTP layer

The MTP layer is kept as the source FP8, as in the GLM-5.3 NestQuant releases. Its routed experts (7.25B parameters, 7.2 GB) are in separate files. They are too large for a 96 GB Mac, so the Mac setup runs without MTP and does not load mtp_experts-*. Machines with more memory can load them for speculative decoding.

Method

  • Rotation: random signs + Hadamard-128 on both sides of each weight matrix.
  • 1.5-bit base: a pattern-rate bitshift trellis code. Step p of each 16-step window uses 1 + ((0xAAAA >> (p % 16)) & 1) bits, so 24 bits per 16 weights. It has a per-tile sign and is fitted with LDLQ.
  • 4-bit residual: a second trellis code on the rotated residual, fitted jointly with the base. It is 2.5 bits per weight on gate/up and 2.8125 on down.
  • Calibration:
  • 4.0M text tokens.
  • 1,200 images (radiology, web screenshots, natural images, OCR), run through Flash's own vision tower.
  • Per-expert Hessians blend 75% text and 25% vision.
  • Fixed set: the 19 experts per layer with the highest boundary-weighted REAP score (routing weight times expert output norm, summed over tokens).
  • Tokens just before the end of reasoning and the end of each turn get more weight: 50 for the last token, 20 for 2-4 tokens before, 5 for 5-16 and 2 for 17-32. This keeps the experts that decide when to stop at 4 bit.
  • Text and vision scores are blended 75/25.
  • The weights of Flash's experts are close to iid Gaussian after rotation (kurtosis 2.996-3.004 on 60 matrices), as in GLM-5.3. The 1.5-bit base error is about 2x the 2-bit error.

Quality

TODO. KL divergence against the BF16 model's full logits on held-out windows, at the 96 GB setup (19 fixed + 48 floating):

Setup 4-bit experts per layer KLD
FP8 reference all TODO
1.5-4 bit (this repo), 96 GB 67 TODO

Verification

  • All 3,532 source tensors outside the routed experts of layers 3-44 are here exactly once. Each one's dtype, shape and sha256 match the source (nonexpert_tensor_sha256.json). None of the 72,576 routed-expert tensors of layers 3-44 are included.
  • TODO: per-layer round trip of the expert files and decode checks at both levels.

Configuration

Architecture
Glm5NextForConditionalGeneration
Context length (tokens)
1,048,576
Layers
45
Hidden size
4,096
Feed-forward size
12,288
Attention heads
64
Key/value heads
64
Head dimension
0
Vocabulary size
154,880
Routed experts
288
Experts active per token
8
Model type
glm5_next
Quantization
fp8

Identity and Version

Repository
jarrelscy/GLM-5.3-Flash-NestQuant-1.5-4bit
Publisher
Jarrel Seah
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
16.9B parameters
Languages
moe, glm
Revision
e0d0aa652d3323399adf5ee454f820835a9c5506
First published
2026-10-05
Last updated
2026-10-06

Files and Weights

393 files, 176.8 GB in total. The weights are 321 files totalling 176.8 GB in pt, safetensors.

Weights321 files · 176.8 GB
Configuration60 files · 16.2 MB
Tokenizer2 files · 20.2 MB
Documentation6 files · 18.0 KB
Other3 files · 246.2 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
layers/L10/tp0.safetensorsWeights489.8 MB d8bc1c11a410
layers/L10/tp1.safetensorsWeights489.8 MB ea65da8f0100
layers/L10/tp2.safetensorsWeights489.8 MB efa934d7543c
layers/L10/tp3.safetensorsWeights489.8 MB edd7b5345c8a
layers/L10/tp4.safetensorsWeights489.8 MB fbfa3453f417
layers/L10/tp5.safetensorsWeights489.8 MB 217412c8ac2c
layers/L10/tp6.safetensorsWeights489.8 MB 573f787a1d4a
layers/L10/tp7.safetensorsWeights489.8 MB c4e9a09bd836
layers/L11/tp0.safetensorsWeights489.6 MB fc30c4559bad
layers/L11/tp1.safetensorsWeights489.6 MB 1d05fbdab2fe
layers/L11/tp2.safetensorsWeights489.6 MB 7ec2a4ed3a4d
layers/L11/tp3.safetensorsWeights489.6 MB 433268e04aa4
layers/L11/tp4.safetensorsWeights489.6 MB c8df323549a1
layers/L11/tp5.safetensorsWeights489.6 MB 0503eae05c75
layers/L11/tp6.safetensorsWeights489.6 MB 7ab3808eb33c
layers/L11/tp7.safetensorsWeights489.6 MB f88307d36e7e
layers/L12/tp0.safetensorsWeights490.4 MB d4417f3ee12e
layers/L12/tp1.safetensorsWeights490.4 MB dcc6f0353a2a
layers/L12/tp2.safetensorsWeights490.4 MB c82fee428422
layers/L12/tp3.safetensorsWeights490.4 MB 7a3e5d4b81b8
layers/L12/tp4.safetensorsWeights490.4 MB 2569d7d5cb82
layers/L12/tp5.safetensorsWeights490.4 MB 86523521b9fe
layers/L12/tp6.safetensorsWeights490.4 MB 037286d6608a
layers/L12/tp7.safetensorsWeights490.4 MB e749b6aa0d32
layers/L13/tp0.safetensorsWeights491.5 MB e1576628fed2
layers/L13/tp1.safetensorsWeights491.5 MB bbc29a0036b9
layers/L13/tp2.safetensorsWeights491.5 MB cb0b7b9a9f79
layers/L13/tp3.safetensorsWeights491.5 MB e29c707e5a40
layers/L13/tp4.safetensorsWeights491.5 MB 51b153e7bdf9
layers/L13/tp5.safetensorsWeights491.5 MB d83922ccdc61
layers/L13/tp6.safetensorsWeights491.5 MB e9b50b364600
layers/L13/tp7.safetensorsWeights491.5 MB ef44f3d3b69d
layers/L14/tp0.safetensorsWeights490.6 MB 21eb313cc54f
layers/L14/tp1.safetensorsWeights490.6 MB e2db698fb45d
layers/L14/tp2.safetensorsWeights490.6 MB 90d8e0f95805
layers/L14/tp3.safetensorsWeights490.6 MB cebfde350f04
layers/L14/tp4.safetensorsWeights490.6 MB 0db02648fdf7
layers/L14/tp5.safetensorsWeights490.6 MB a0b8830c9105
layers/L14/tp6.safetensorsWeights490.6 MB cad990f66b44
layers/L14/tp7.safetensorsWeights490.6 MB 5f6de26e42be
layers/L15/tp0.safetensorsWeights490.4 MB 330d82108a4d
layers/L15/tp1.safetensorsWeights490.4 MB d20c7b1c896b
layers/L15/tp2.safetensorsWeights490.4 MB 3a20e0e7a146
layers/L15/tp3.safetensorsWeights490.4 MB 7579a83c3dd6
layers/L15/tp4.safetensorsWeights490.4 MB 4998271493b4
layers/L15/tp5.safetensorsWeights490.4 MB 1f37781b900f
layers/L15/tp6.safetensorsWeights490.4 MB a00c8773755d
layers/L15/tp7.safetensorsWeights490.4 MB 6885c7efd8a2
layers/L16/tp0.safetensorsWeights490.4 MB 0606289459a2
layers/L16/tp1.safetensorsWeights490.4 MB 15e8693d48e2
layers/L16/tp2.safetensorsWeights490.4 MB c0392912bc9e
layers/L16/tp3.safetensorsWeights490.4 MB eb7e9f6630e5
layers/L16/tp4.safetensorsWeights490.4 MB 1225ce57e5b9
layers/L16/tp5.safetensorsWeights490.4 MB 5003700242e6
layers/L16/tp6.safetensorsWeights490.4 MB 7e2857b625d9
layers/L16/tp7.safetensorsWeights490.4 MB f4ef21cdedde
layers/L17/tp0.safetensorsWeights490.3 MB 8ec1503dceea
layers/L17/tp1.safetensorsWeights490.3 MB 6773c0679794
layers/L17/tp2.safetensorsWeights490.3 MB 035e8b472d58
layers/L17/tp3.safetensorsWeights490.3 MB 1440c888b32e
layers/L17/tp4.safetensorsWeights490.3 MB facdbdf311c6
layers/L17/tp5.safetensorsWeights490.3 MB 496e75579565
layers/L17/tp6.safetensorsWeights490.3 MB 165ebadd96b3
layers/L17/tp7.safetensorsWeights490.3 MB 1856fb66621c
layers/L18/tp0.safetensorsWeights491.0 MB 26a5886750ef
layers/L18/tp1.safetensorsWeights491.0 MB 4a1b1febfe15
layers/L18/tp2.safetensorsWeights491.0 MB bad5f3923f09
layers/L18/tp3.safetensorsWeights491.0 MB 505a871f37b6
layers/L18/tp4.safetensorsWeights491.0 MB 75f0dcac5093
layers/L18/tp5.safetensorsWeights491.0 MB 228225450d53
layers/L18/tp6.safetensorsWeights491.0 MB 868b818436d5
layers/L18/tp7.safetensorsWeights491.0 MB 81c4b8ad6b6e
layers/L19/tp0.safetensorsWeights490.3 MB 03bbc31fb306
layers/L19/tp1.safetensorsWeights490.3 MB d9c214f4e211
layers/L19/tp2.safetensorsWeights490.3 MB f47f6d2100a1
layers/L19/tp3.safetensorsWeights490.3 MB d5baf876f42f
layers/L19/tp4.safetensorsWeights490.3 MB 31dbc8e3c482
layers/L19/tp5.safetensorsWeights490.3 MB 7d4934a954ce
layers/L19/tp6.safetensorsWeights490.3 MB 084dfbe19800
layers/L19/tp7.safetensorsWeights490.3 MB 35a74d67fecf
layers/L20/tp0.safetensorsWeights490.1 MB 7cccb83665f3
layers/L20/tp1.safetensorsWeights490.1 MB 2a432b520256
layers/L20/tp2.safetensorsWeights490.1 MB 0147e278bfe7
layers/L20/tp3.safetensorsWeights490.1 MB 5723797201a8
layers/L20/tp4.safetensorsWeights490.1 MB c56c237244cc
layers/L20/tp5.safetensorsWeights490.1 MB b87bc268f488
layers/L20/tp6.safetensorsWeights490.1 MB e389189fec83
layers/L20/tp7.safetensorsWeights490.1 MB d25d6860c978
layers/L21/tp0.safetensorsWeights490.3 MB 325c2aac0218
layers/L21/tp1.safetensorsWeights490.3 MB 3827a34d74f7
layers/L21/tp2.safetensorsWeights490.3 MB c41d6154ec27
layers/L21/tp3.safetensorsWeights490.3 MB 3e229f4967ea
layers/L21/tp4.safetensorsWeights490.3 MB 84415f8c09f6
layers/L21/tp5.safetensorsWeights490.3 MB f1a145564b13
layers/L21/tp6.safetensorsWeights490.3 MB afe444e09ec5
layers/L21/tp7.safetensorsWeights490.3 MB fa04282cf8dd
layers/L22/tp0.safetensorsWeights489.7 MB 78dee494d41c
layers/L22/tp1.safetensorsWeights489.7 MB 216269f40326
layers/L22/tp2.safetensorsWeights489.7 MB ba5bf470b053
layers/L22/tp3.safetensorsWeights489.7 MB 39ec64b97851
layers/L22/tp4.safetensorsWeights489.7 MB 49a440b6e604
layers/L22/tp5.safetensorsWeights489.7 MB ca14fb5d559e
layers/L22/tp6.safetensorsWeights489.7 MB 152cd00783b3
layers/L22/tp7.safetensorsWeights489.7 MB de1e500890ec
layers/L23/tp0.safetensorsWeights489.7 MB 5e663f5447b4
layers/L23/tp1.safetensorsWeights489.7 MB 164a81f0ef27
layers/L23/tp2.safetensorsWeights489.7 MB 5075045f364f
layers/L23/tp3.safetensorsWeights489.7 MB 78e188966077
layers/L23/tp4.safetensorsWeights489.7 MB 06d5290d84b7
layers/L23/tp5.safetensorsWeights489.7 MB 9650044f64d0
layers/L23/tp6.safetensorsWeights489.7 MB e720198c4373
layers/L23/tp7.safetensorsWeights489.7 MB e1c6570bbc3d
layers/L24/tp0.safetensorsWeights489.9 MB 2c45cfda6dec
layers/L24/tp1.safetensorsWeights489.9 MB b8f0605c703d
layers/L24/tp2.safetensorsWeights489.9 MB 30c85bb86900
layers/L24/tp3.safetensorsWeights489.9 MB 48eb7fe5f9a8
layers/L24/tp4.safetensorsWeights489.9 MB 296f299d0af4
layers/L24/tp5.safetensorsWeights489.9 MB f0303fadd05a
layers/L24/tp6.safetensorsWeights489.9 MB d7242ba65711
layers/L24/tp7.safetensorsWeights489.9 MB d6d02579470c
layers/L25/tp0.safetensorsWeights489.5 MB 18ba78a53c5f
layers/L25/tp1.safetensorsWeights489.5 MB 00ea664b7079
layers/L25/tp2.safetensorsWeights489.5 MB 2c38e305cca3
layers/L25/tp3.safetensorsWeights489.5 MB 9175d0e8bad5
layers/L25/tp4.safetensorsWeights489.5 MB acdf7d3ef0d0
layers/L25/tp5.safetensorsWeights489.5 MB d73661ec4944
layers/L25/tp6.safetensorsWeights489.5 MB e45b485eb1f6
layers/L25/tp7.safetensorsWeights489.5 MB eb12def583fe
layers/L26/tp0.safetensorsWeights489.4 MB 39ce4e028b2c
layers/L26/tp1.safetensorsWeights489.4 MB b8b040184259
layers/L26/tp2.safetensorsWeights489.4 MB cdb209007091
layers/L26/tp3.safetensorsWeights489.4 MB 72f0df9091e1
layers/L26/tp4.safetensorsWeights489.4 MB dd0fee174f98
layers/L26/tp5.safetensorsWeights489.4 MB fbecc06d3639
layers/L26/tp6.safetensorsWeights489.4 MB 48de8daed0b1
layers/L26/tp7.safetensorsWeights489.4 MB 9eae9d68ae89
layers/L27/tp0.safetensorsWeights489.4 MB 66263729cc0c
layers/L27/tp1.safetensorsWeights489.4 MB 7feb48f3a511
layers/L27/tp2.safetensorsWeights489.4 MB de19feca1c5e
layers/L27/tp3.safetensorsWeights489.4 MB 534ad0bb7a74
layers/L27/tp4.safetensorsWeights489.4 MB a250e6061b41
layers/L27/tp5.safetensorsWeights489.4 MB 4d6c87919a6a
layers/L27/tp6.safetensorsWeights489.4 MB 137ed6360bb7
layers/L27/tp7.safetensorsWeights489.4 MB 79f0c05ca9d8
layers/L28/tp0.safetensorsWeights489.4 MB 5f160ef2a5ce
layers/L28/tp1.safetensorsWeights489.4 MB e1edb67cdce5
layers/L28/tp2.safetensorsWeights489.4 MB 36b4285b4ab4
layers/L28/tp3.safetensorsWeights489.4 MB 22bae3f0a865
layers/L28/tp4.safetensorsWeights489.4 MB bd857a550bb6
layers/L28/tp5.safetensorsWeights489.4 MB f57db4819af7
layers/L28/tp6.safetensorsWeights489.4 MB f9814acf53d8
layers/L28/tp7.safetensorsWeights489.4 MB c23699dcf42f
layers/L29/tp0.safetensorsWeights489.8 MB 1656b3bf38ea
layers/L29/tp1.safetensorsWeights489.8 MB 5190c9afdeaa
layers/L29/tp2.safetensorsWeights489.8 MB 3dc7e7a887d7
layers/L29/tp3.safetensorsWeights489.8 MB 5b3b6d2a3d43
layers/L29/tp4.safetensorsWeights489.8 MB 93f3e74f5427
layers/L29/tp5.safetensorsWeights489.8 MB 2a05e5bad199
layers/L29/tp6.safetensorsWeights489.8 MB c57b9628469d
layers/L29/tp7.safetensorsWeights489.8 MB 5031f9511777
layers/L3/tp0.safetensorsWeights495.6 MB 530caf95403e
layers/L3/tp1.safetensorsWeights495.6 MB 84d5ea9a44de
layers/L3/tp2.safetensorsWeights495.6 MB 867326577890
layers/L3/tp3.safetensorsWeights495.6 MB 363e17109be9
layers/L3/tp4.safetensorsWeights495.6 MB 484b13bdd458
layers/L3/tp5.safetensorsWeights495.6 MB 2cf22c267217
layers/L3/tp6.safetensorsWeights495.6 MB 3c8df8e307d2
layers/L3/tp7.safetensorsWeights495.6 MB 5591d24443e4
layers/L30/tp0.safetensorsWeights489.5 MB 04f3db971b07
layers/L30/tp1.safetensorsWeights489.5 MB 40a1dad07a72
layers/L30/tp2.safetensorsWeights489.5 MB d219609f90ca
layers/L30/tp3.safetensorsWeights489.5 MB 33855d3a03df
layers/L30/tp4.safetensorsWeights489.5 MB bf580eaf3b75
layers/L30/tp5.safetensorsWeights489.5 MB 9f7ab2cabd7c
layers/L30/tp6.safetensorsWeights489.5 MB 55ae1c50d967
layers/L30/tp7.safetensorsWeights489.5 MB aeb39b57578e
layers/L31/tp0.safetensorsWeights489.3 MB 96d723e6e95d
layers/L31/tp1.safetensorsWeights489.3 MB 867ad8b5ddff
layers/L31/tp2.safetensorsWeights489.3 MB 203beb373ef4
layers/L31/tp3.safetensorsWeights489.3 MB f1ed86393200
layers/L31/tp4.safetensorsWeights489.3 MB 6e172c294cad
layers/L31/tp5.safetensorsWeights489.3 MB ab7a354df76b
layers/L31/tp6.safetensorsWeights489.3 MB 203ae951e345
layers/L31/tp7.safetensorsWeights489.3 MB 314d2e444bd4
layers/L32/tp0.safetensorsWeights489.5 MB ba4e97138fbd
layers/L32/tp1.safetensorsWeights489.5 MB e7eec356243f
layers/L32/tp2.safetensorsWeights489.5 MB b8e0fe9d4ebc
layers/L32/tp3.safetensorsWeights489.5 MB 1037c9a7e161
layers/L32/tp4.safetensorsWeights489.5 MB b3062db88412
layers/L32/tp5.safetensorsWeights489.5 MB ec3adc44d4b7
layers/L32/tp6.safetensorsWeights489.5 MB 357e16cc06f5
layers/L32/tp7.safetensorsWeights489.5 MB daaeedd492c3
layers/L33/tp0.safetensorsWeights489.5 MB fc6917093537
layers/L33/tp1.safetensorsWeights489.5 MB 944d9f79683b
layers/L33/tp2.safetensorsWeights489.5 MB 00e6d529756a
layers/L33/tp3.safetensorsWeights489.5 MB 0581b1f54190
layers/L33/tp4.safetensorsWeights489.5 MB afc262112081
layers/L33/tp5.safetensorsWeights489.5 MB 99ef7d91761c
layers/L33/tp6.safetensorsWeights489.5 MB 2cf797c464e1
layers/L33/tp7.safetensorsWeights489.5 MB 21e84329a9bf
layers/L34/tp0.safetensorsWeights489.5 MB 3d73e99b7370
layers/L34/tp1.safetensorsWeights489.5 MB 6d1dd5a09074
layers/L34/tp2.safetensorsWeights489.5 MB 07ca8b7f53a8
layers/L34/tp3.safetensorsWeights489.5 MB 695dbf1164d0
layers/L34/tp4.safetensorsWeights489.5 MB 9e9915ba8e6a
layers/L34/tp5.safetensorsWeights489.5 MB ba69061276f7
layers/L34/tp6.safetensorsWeights489.5 MB c8632513413b
layers/L34/tp7.safetensorsWeights489.5 MB 9629e3b11610
layers/L35/tp0.safetensorsWeights489.5 MB b67ec8230b30
layers/L35/tp1.safetensorsWeights489.5 MB 49130d2a55c0
layers/L35/tp2.safetensorsWeights489.5 MB 1723d0a59225
layers/L35/tp3.safetensorsWeights489.5 MB 0efeaa2fe6f4
layers/L35/tp4.safetensorsWeights489.5 MB 798a95df39ed
layers/L35/tp5.safetensorsWeights489.5 MB fc90b5a4d0d6
layers/L35/tp6.safetensorsWeights489.5 MB 7293a1e30718
layers/L35/tp7.safetensorsWeights489.5 MB 176db27dd4cf
layers/L36/tp0.safetensorsWeights489.4 MB 50ac10366943
layers/L36/tp1.safetensorsWeights489.4 MB 01b49da9a365
layers/L36/tp2.safetensorsWeights489.4 MB 4f0fa4c64e68
layers/L36/tp3.safetensorsWeights489.4 MB 397a93e86de3
layers/L36/tp4.safetensorsWeights489.4 MB 7febc5da0013
layers/L36/tp5.safetensorsWeights489.4 MB 804bf8d34e55
layers/L36/tp6.safetensorsWeights489.4 MB 3c34b9b0f9f6
layers/L36/tp7.safetensorsWeights489.4 MB 95a52d36c899
layers/L37/tp0.safetensorsWeights489.9 MB 09cfa502d3b1
layers/L37/tp1.safetensorsWeights489.9 MB 1d7366e9d2f2
layers/L37/tp2.safetensorsWeights489.9 MB 2fd22335b64f
layers/L37/tp3.safetensorsWeights489.9 MB a05756408999
layers/L37/tp4.safetensorsWeights489.9 MB be93107d0304
layers/L37/tp5.safetensorsWeights489.9 MB 63b8100ed09a
layers/L37/tp6.safetensorsWeights489.9 MB 46ff9e102ca7
layers/L37/tp7.safetensorsWeights489.9 MB 9a91881d9f10
layers/L38/tp0.safetensorsWeights489.8 MB 1fcfd9e54bed
layers/L38/tp1.safetensorsWeights489.8 MB 4ec556a81a6c
layers/L38/tp2.safetensorsWeights489.8 MB 29884f00576c
layers/L38/tp3.safetensorsWeights489.8 MB f0119c155ff1
layers/L38/tp4.safetensorsWeights489.8 MB 6be0e14692bc
layers/L38/tp5.safetensorsWeights489.8 MB c562e486b7e6
layers/L38/tp6.safetensorsWeights489.8 MB 3feec42f7af7
layers/L38/tp7.safetensorsWeights489.8 MB ee71273548a7
layers/L39/tp0.safetensorsWeights489.9 MB 86cca04087c5
layers/L39/tp1.safetensorsWeights489.9 MB a6cb69b639ff
layers/L39/tp2.safetensorsWeights489.9 MB e483b9a40ee2
layers/L39/tp3.safetensorsWeights489.9 MB 118fda9ab204
layers/L39/tp4.safetensorsWeights489.9 MB d0cbcf19d61f
layers/L39/tp5.safetensorsWeights489.9 MB b090382c959a
layers/L39/tp6.safetensorsWeights489.9 MB a9aa922d4203
layers/L39/tp7.safetensorsWeights489.9 MB 91448feee427
layers/L4/tp0.safetensorsWeights489.5 MB c38a698a0aaa
layers/L4/tp1.safetensorsWeights489.5 MB a1f9140beeae
layers/L4/tp2.safetensorsWeights489.5 MB e71cb614a938
layers/L4/tp3.safetensorsWeights489.5 MB e3b5b61490b8
layers/L4/tp4.safetensorsWeights489.5 MB 062682389e36
layers/L4/tp5.safetensorsWeights489.5 MB 9736f1e9bdb3
layers/L4/tp6.safetensorsWeights489.5 MB 8598cd7032db
layers/L4/tp7.safetensorsWeights489.5 MB 3ef79233661f
layers/L40/tp0.safetensorsWeights491.2 MB d3c8e65f57eb
layers/L40/tp1.safetensorsWeights491.2 MB 7075c603d2cf
layers/L40/tp2.safetensorsWeights491.2 MB 41e152282796
layers/L40/tp3.safetensorsWeights491.2 MB 1d24e16fe880
layers/L40/tp4.safetensorsWeights491.2 MB 72588c0850b4
layers/L40/tp5.safetensorsWeights491.2 MB 2beccdc62126
layers/L40/tp6.safetensorsWeights491.2 MB 6fb2d859d131
layers/L40/tp7.safetensorsWeights491.2 MB eb0d4cae4980
layers/L41/tp0.safetensorsWeights493.9 MB 55ff46f3dd90
layers/L41/tp1.safetensorsWeights493.9 MB f50128114b8d
layers/L41/tp2.safetensorsWeights493.9 MB bfcf68e51fdb
layers/L41/tp3.safetensorsWeights493.9 MB 9a5af93f207f
layers/L41/tp4.safetensorsWeights493.9 MB 837fc6810f93
layers/L41/tp5.safetensorsWeights493.9 MB 298b439d22a3
layers/L41/tp6.safetensorsWeights493.9 MB fd7ce3e3f5a8
layers/L41/tp7.safetensorsWeights493.9 MB 8fe91d0912cf
layers/L5/tp0.safetensorsWeights489.3 MB b3515163f6d2
layers/L5/tp1.safetensorsWeights489.3 MB d812187a2ba6
layers/L5/tp2.safetensorsWeights489.3 MB 23668a7e1e32
layers/L5/tp3.safetensorsWeights489.3 MB de28e64fcda0
layers/L5/tp4.safetensorsWeights489.3 MB 151423c84b20
layers/L5/tp5.safetensorsWeights489.3 MB 7d79d9a9f91a
layers/L5/tp6.safetensorsWeights489.3 MB 68ec9d67a8d8
layers/L5/tp7.safetensorsWeights489.3 MB dcff95721731
layers/L6/tp0.safetensorsWeights489.4 MB 264f194f4b83
layers/L6/tp1.safetensorsWeights489.4 MB 7ecac3627105
layers/L6/tp2.safetensorsWeights489.4 MB baef7d07264c
layers/L6/tp3.safetensorsWeights489.4 MB d2f1b6701013
layers/L6/tp4.safetensorsWeights489.4 MB a4de4dd56e5b
layers/L6/tp5.safetensorsWeights489.4 MB fe680cc61056
layers/L6/tp6.safetensorsWeights489.4 MB c8cd28228ba5
layers/L6/tp7.safetensorsWeights489.4 MB f9a3a9d17819
layers/L7/tp0.safetensorsWeights490.0 MB 27bc76c0a927
layers/L7/tp1.safetensorsWeights490.0 MB ead23ab4038f
layers/L7/tp2.safetensorsWeights490.0 MB 377e43f41c5f
layers/L7/tp3.safetensorsWeights490.0 MB eb0ddee209c2
layers/L7/tp4.safetensorsWeights490.0 MB 24b96484812d
layers/L7/tp5.safetensorsWeights490.0 MB d49bc05e5c43
layers/L7/tp6.safetensorsWeights490.0 MB 9cfaf234fdcb
layers/L7/tp7.safetensorsWeights490.0 MB 7ef0dff39b64
layers/L8/tp0.safetensorsWeights490.2 MB ae25d602168f
layers/L8/tp1.safetensorsWeights490.2 MB 42bec7dc8fbb
layers/L8/tp2.safetensorsWeights490.2 MB 55230d46aba1
layers/L8/tp3.safetensorsWeights490.2 MB f8b44d40795e
layers/L8/tp4.safetensorsWeights490.2 MB ac76f16c034c
layers/L8/tp5.safetensorsWeights490.2 MB eb02d603cb98
layers/L8/tp6.safetensorsWeights490.2 MB bf70548f7faf
layers/L8/tp7.safetensorsWeights490.2 MB cd3694d0368c
layers/L9/tp0.safetensorsWeights489.7 MB 4b20460c86b5
layers/L9/tp1.safetensorsWeights489.7 MB fccb681374fe
layers/L9/tp2.safetensorsWeights489.7 MB e72738c07454
layers/L9/tp3.safetensorsWeights489.7 MB 9080ec0b37ce
layers/L9/tp4.safetensorsWeights489.7 MB 8913a55df0d5
layers/L9/tp5.safetensorsWeights489.7 MB 9327273b52ec
layers/L9/tp6.safetensorsWeights489.7 MB 8b325c5a5afa
layers/L9/tp7.safetensorsWeights489.7 MB 5d41eb8b6a4e
mtp_experts-00001-of-00002.safetensorsWeights5.0 GB 32defe995062
mtp_experts-00002-of-00002.safetensorsWeights2.3 GB 28181b15ba7b
nonexpert-00001-of-00004.safetensorsWeights5.0 GB 5c261c79d37d
nonexpert-00002-of-00004.safetensorsWeights5.0 GB 7fc6df7f9832
nonexpert-00003-of-00004.safetensorsWeights4.3 GB ec23074b4c30
nonexpert-00004-of-00004.safetensorsWeights1.3 GB 2f1c6ac8756b
serving/predictor/joint/jF.ptWeights5.3 MB fe587d4b1b9f
serving/predictor_nf48/joint/jF.ptWeights5.3 MB 400613e5f0f0
vision_tower.safetensorsWeights1.1 GB e9f5cbb64bdd
config.jsonConfiguration70.9 KB —
fixed_set.jsonConfiguration835.4 KB —
generation_config.jsonConfiguration194 B —
layers/L10/manifest.jsonConfiguration364.1 KB —
layers/L11/manifest.jsonConfiguration363.5 KB —
layers/L12/manifest.jsonConfiguration366.0 KB —
layers/L13/manifest.jsonConfiguration369.4 KB —
layers/L14/manifest.jsonConfiguration366.6 KB —
layers/L15/manifest.jsonConfiguration365.8 KB —
layers/L16/manifest.jsonConfiguration365.7 KB —
layers/L17/manifest.jsonConfiguration365.6 KB —
layers/L18/manifest.jsonConfiguration367.6 KB —
layers/L19/manifest.jsonConfiguration365.7 KB —
layers/L20/manifest.jsonConfiguration364.9 KB —
layers/L21/manifest.jsonConfiguration365.3 KB —
layers/L22/manifest.jsonConfiguration363.6 KB —
layers/L23/manifest.jsonConfiguration363.9 KB —
layers/L24/manifest.jsonConfiguration364.4 KB —
layers/L25/manifest.jsonConfiguration363.1 KB —
layers/L26/manifest.jsonConfiguration362.7 KB —
layers/L27/manifest.jsonConfiguration363.1 KB —
layers/L28/manifest.jsonConfiguration362.7 KB —
layers/L29/manifest.jsonConfiguration364.0 KB —
layers/L3/manifest.jsonConfiguration376.1 KB —
layers/L30/manifest.jsonConfiguration362.9 KB —
layers/L31/manifest.jsonConfiguration362.6 KB —
layers/L32/manifest.jsonConfiguration363.0 KB —
layers/L33/manifest.jsonConfiguration363.2 KB —
layers/L34/manifest.jsonConfiguration362.9 KB —
layers/L35/manifest.jsonConfiguration363.0 KB —
layers/L36/manifest.jsonConfiguration362.8 KB —
layers/L37/manifest.jsonConfiguration364.4 KB —
layers/L38/manifest.jsonConfiguration364.1 KB —
layers/L39/manifest.jsonConfiguration364.0 KB —
layers/L4/manifest.jsonConfiguration363.2 KB —
layers/L40/manifest.jsonConfiguration368.3 KB —
layers/L41/manifest.jsonConfiguration376.1 KB —
layers/L5/manifest.jsonConfiguration362.6 KB —
layers/L6/manifest.jsonConfiguration362.9 KB —
layers/L7/manifest.jsonConfiguration364.4 KB —
layers/L8/manifest.jsonConfiguration365.2 KB —
layers/L9/manifest.jsonConfiguration363.6 KB —
model.safetensors.index.jsonConfiguration379.7 KB —
nonexpert_manifest.jsonConfiguration4.9 KB —
nonexpert_tensor_sha256.jsonConfiguration465.1 KB —
processor_config.jsonConfiguration909 B —
serving/predictor/joint/gpu_predictor.pyConfiguration14.0 KB —
serving/predictor/joint/jF.jsonConfiguration10.8 KB —
serving/predictor/joint/jlib.pyConfiguration15.5 KB —
serving/predictor/joint/joint_predictor.pyConfiguration11.5 KB —
serving/predictor/joint/parity_stream.pyConfiguration7.5 KB —
serving/predictor/joint/train.pyConfiguration19.8 KB —
serving/predictor/predictor.jsonConfiguration981 B —
serving/predictor_nf48/joint/gpu_predictor.pyConfiguration14.0 KB —
serving/predictor_nf48/joint/jF.jsonConfiguration10.8 KB —
serving/predictor_nf48/joint/jlib.pyConfiguration15.5 KB —
serving/predictor_nf48/joint/joint_predictor.pyConfiguration11.5 KB —
serving/predictor_nf48/joint/parity_stream.pyConfiguration7.5 KB —
serving/predictor_nf48/joint/train.pyConfiguration19.8 KB —
serving/predictor_nf48/predictor.jsonConfiguration981 B —
LICENSEDocumentation1.1 KB —
README.mdDocumentation5.5 KB —
serving/predictor/joint/README.mdDocumentation3.4 KB —
serving/predictor/joint/SERVE.mdDocumentation2.3 KB —
serving/predictor_nf48/joint/README.mdDocumentation3.4 KB —
serving/predictor_nf48/joint/SERVE.mdDocumentation2.3 KB —
chat_template.jinjaOther10.9 KB —
serving/predictor/joint/v2_sal_tweedie1.5.txtOther117.6 KB —
serving/predictor_nf48/joint/v2_sal_tweedie1.5.txtOther117.6 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.2 MB 19e773648cb4
tokenizer_config.jsonTokenizer761 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
176.8 GB
Download from Jarrel Seah

Released by Jarrel Seah through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published176.8 GB
16-bit33.8 GB
8-bit16.9 GB
4-bit8.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About GLM-5.3-Flash-NestQuant-1.5-4bit

How much GPU memory does GLM-5.3-Flash-NestQuant-1.5-4bit need?

About 40.6 GB at 16-bit and 10.2 GB at 4-bit: the weights (16.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run GLM-5.3-Flash-NestQuant-1.5-4bit on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use GLM-5.3-Flash-NestQuant-1.5-4bit commercially?

Yes. GLM-5.3-Flash-NestQuant-1.5-4bit is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is GLM-5.3-Flash-NestQuant-1.5-4bit's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B-NVFP4

RadixArk

The RadixArk Qwen3.8-27B-NVFP4 model is the quantized version of Qwen/Qwen3.8-27B. The quantization was produced at RadixArk using NVIDIA Model Optimizer, following a mixed NVFP4 W4A4 recipe. Run on SGLang: launch command and per-platform recipes in the Qwen3.8-27B cookbook. This model is not owned or developed by RadixArk. It is a quantized derivative of Qwen's model; see the upstream Qwen3.8-27B model card for the source model's capabilities, training information, limitations, and license. Global Developers looking to deploy an off-the-shelf, pre-quantized model in AI agent systems, chatbots, RAG systems, and other AI-powered applications. Hugging Face 08/14/2026 via…

Open weights apache-2.0 18.2B parameters 262,144 tokens Model Optimizer

Model · Image and text to text

Swift-Qwen3.8-27B-Uncensored-NVFP4

AJ Gazin

NVFP4 checkpoint of an abliterated Swift-Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and SGLang. GGUFs for llama.cpp: The Swift 1.5 version is source). - Swift's own NVFP4 recipe, unmodified, from ukisai/Swift-Qwen3.8-27B-NVFP4, calibrated with NVIDIA ModelOpt. - MTP head and vision tower in BF16, bit-identical to the source. 21.9 GB, NVIDIA ModelOpt mixed-precision format. Needs a vLLM with ModelOpt mixed-precision support (tested on 0.29.0). No --quantization flag. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default. Tested on an RTX 5090 (32 GB) with vLLM 0.29.0: NVFP4 layers on…

Open weights other 18.2B parameters 262,144 tokens vllm

Model · Image and text to text

Swift-1.5-Qwen3.8-27B-Uncensored-NVFP4

AJ Gazin

NVFP4 checkpoint of an abliterated Swift 1.5 Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and SGLang. GGUFs for llama.cpp: (measured on the BF16 source). - Swift's own NVFP4 recipe, unmodified, from ukisai/Swift-Qwen3.8-27B-NVFP4, calibrated with NVIDIA ModelOpt. The module split matches UkisAI's Swift 1.5 NVFP4 exactly. - MTP head and vision tower in BF16, bit-identical to the source. 21.9 GB, NVIDIA ModelOpt mixed-precision format. Needs a vLLM with ModelOpt mixed-precision support. No --quantization flag. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default. Same format, recipe, module…

Open weights other 18.2B parameters 262,144 tokens vllm

Model · Image and text to text

Qwen3.8-27B-Uncensored-NVFP4-v100-skinny

Tim Eastwood

Mixed-precision quantization of prepared for dnv2003/v100-skinny. - MLP gateproj, upproj, downproj, and lmhead: NVFP4, group size 16 87c9f8cf83021957d1a1a575c90c9a4eaaf7ef0c See quantization-audit.json and hfquantconfig.json for the complete machine-readable layout. This model has had safety alignment substantially removed. Use it only for lawful, controlled research and add appropriate safeguards before deployment.

Open weights apache-2.0 18.2B parameters 262,144 tokens Model Optimizer

Model · Image and text to text

Swift-1.5-Qwen3.8-27b-Quark-RTN-MXFP4

Ethan Todd

quantized to OCP MXFP4 4-bit weights in the same Quark checkpoint container as The weights are plain RTN, not AWQ: see the note below. The 15 MTP tensors are kept in BF16; the radiance runtime loads them with RADIANCEQUARKBF16MTP=1. Why RTN instead of AWQ. The first build of this checkpoint used AWQ with the same smoothing recipe as AMD's release. The AWQ fold itself was mathematically consistent, but on this model a few layers converged to weights to 32-element MXFP4 blocks then destroyed those layers on real, outlier-carrying inputs (layer 7 output relative error ~324), and repairing the worst layers individually was not enough; the remaining smoothed layers still accumulated too much…

Open weights other 15.6B parameters 262,144 tokens transformers