SAVRN Model Hub · Comparisons
GLM-5.3-Flash vs Qwen3.8-Flash-Next-Uncensored-NVFP4
GLM-5.3-Flash has 321.3B parameters and Qwen3.8-Flash-Next-Uncensored-NVFP4 has 180B parameters; GLM-5.3-Flash is released under MIT License and Qwen3.8-Flash-Next-Uncensored-NVFP4 under Apache License 2.0; at 16-bit, GLM-5.3-Flash needs about 771.2 GB (3x MI355X from $7.77 an hour) and Qwen3.8-Flash-Next-Uncensored-NVFP4 about 432 GB (2x MI325X from $4.00 an hour).
| Field | GLM-5.3-Flash zai-org/GLM-5.3-Flash | Qwen3.8-Flash-Next-Uncensored-NVFP4 orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 |
|---|---|---|
| Publisher | Z.ai | OrcaRouter |
| Task | Image and text to text | Image and text to text |
| Modality | Image and text | Image and text |
| Parameters, as reported | 321.3B parameters | 180B parameters |
| Architecture | Glm5NextForConditionalGeneration | Qwen4ExpForConditionalGeneration |
| Library | transformers | transformers |
| Context length | 1,048,576 tokens | Not stated |
| Repository size | 328.4 GB | 183.5 GB |
| Artifact formats | safetensors | safetensors |
| License | mit | apache-2.0 |
| Access | Open weights, no gate | Access requested at publisher |
| Memory at 16-bit (weights and margin) | 771.2 GB | 432 GB |
| Cheapest GPUs at 16-bit, per hour | 3x MI355X, $7.77 | 2x MI325X, $4.00 |
| Memory at 4-bit (weights and margin) | 192.8 GB | 108 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI325X, $2.00 | 1x MI300X, $1.85 |
| Revision viewed | eb9eb208eb0d | 38efbffeb215 |
| Downloads reported by the hub | 3.1M | 13.2k |
| Last observed | 2026-09-21 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
Other Reported Results
These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.
GLM-5.3-Flash
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| PaddlePaddle/Real5-OmniDocBench | Task illuminationMetric illuminationComparison conditions not established | 91.4 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task overallMetric overallComparison conditions not established | 90.76 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task scanningMetric scanningComparison conditions not established | 91.43 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task screen_photographyMetric screen_photographyComparison conditions not established | 90.36 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task skewMetric skewComparison conditions not established | 90.76 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task warpingMetric warpingComparison conditions not established | 89.84 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| cais/hle | Task hleMetric hleSetup HLE with tools (full set) and a 300K-context management strategy, not the no-tools default; judged by GPT-5.6-luna (medium).Comparison conditions not established | 55.3 | GLM-5.3-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| datacurve/deep-swe | Task deep_sweMetric deep_sweSetup Reported as DeepSWE v1.1 on the model card, run via the mini-swe-agent harness with 400K context.Comparison conditions not established | 63.4 | GLM-5.3-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| harborframework/terminal-bench-2.1 | Task terminalbench_2_1Metric terminalbench_2_1Comparison conditions not established | 84.3 | GLM-5.3-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ExtractBench | Task longMetric longSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established | 27.83 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ExtractBench | Task meanMetric meanSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established | 80.75 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ExtractBench | Task mediumMetric mediumSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established | 51.56 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ExtractBench | Task shortMetric shortSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established | 96.3 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task chartMetric chartSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 54.96 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task layoutMetric layoutSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 41.78 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task meanMetric meanSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 70.75 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task tableMetric tableSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 89.18 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task text_contentMetric text_contentSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 88.34 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task text_formattingMetric text_formattingSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 79.49 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
SAVRN's Notes on GLM-5.3-Flash
Only 18B of the 321.3B parameters fire on any token, but all 288 routed experts sit in memory regardless, so the 16-bit footprint is 771.2 GB on three 288 GB MI355X cards at $7.77 per hour, 8-bit is 385.6 GB on two MI325X at $4.00, and 4-bit is 192.8 GB on one 256 GB MI325X at $2.00. Z.ai's first natively multimodal GLM-5 model: image and text in, text out, 1,048,576-token context.
Under MIT the obligation is one line, keep the copyright and permission notice with the files; commercial use, modification and redistribution are allowed. That million-token window needs memory left after the weights; on one 256 GB card at 4-bit little is. Baseten, DeepInfra, Fireworks, Novita and Together AI list it on the Index at $0.15 in and $0.50 out per million tokens, so set your tokens per hour against $2.00, $4.00 or $7.77 and read the ledger.
Questions
Which is larger, GLM-5.3-Flash or Qwen3.8-Flash-Next-Uncensored-NVFP4?
GLM-5.3-Flash (321.3B parameters) is larger than Qwen3.8-Flash-Next-Uncensored-NVFP4 (180B parameters), by the parameter counts their publishers report.
Which is cheaper to run, GLM-5.3-Flash or Qwen3.8-Flash-Next-Uncensored-NVFP4?
At 4-bit, GLM-5.3-Flash fits on 1x MI325X from $2.00 an hour and Qwen3.8-Flash-Next-Uncensored-NVFP4 on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use GLM-5.3-Flash commercially?
Yes. GLM-5.3-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
Can I use Qwen3.8-Flash-Next-Uncensored-NVFP4 commercially?
Yes. Qwen3.8-Flash-Next-Uncensored-NVFP4 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.