SAVRN Model Hub · Comparisons
GLM-5.3-Flash vs Qwen3.8-Flash-Next
GLM-5.3-Flash has 321.3B parameters and Qwen3.8-Flash-Next has 180B parameters; GLM-5.3-Flash is released under MIT License and Qwen3.8-Flash-Next under other; at 16-bit, GLM-5.3-Flash needs about 771.2 GB (3x MI355X from $7.77 an hour) and Qwen3.8-Flash-Next about 432 GB (2x MI325X from $4.00 an hour).
| Field | GLM-5.3-Flash zai-org/GLM-5.3-Flash | Qwen3.8-Flash-Next Qwen/Qwen3.8-Flash-Next |
|---|---|---|
| Publisher | Z.ai | Qwen |
| Task | Image and text to text | Image and text to text |
| Modality | Image and text | Image and text |
| Parameters, as reported | 321.3B parameters | 180B parameters |
| Architecture | Glm5NextForConditionalGeneration | Qwen4ExpForConditionalGeneration |
| Library | transformers | transformers |
| Context length | 1,048,576 tokens | 262,144 tokens |
| Repository size | 328.4 GB | 360.0 GB |
| Artifact formats | safetensors | safetensors |
| License | mit | other |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 771.2 GB | 432 GB |
| Cheapest GPUs at 16-bit, per hour | 3x MI355X, $7.77 | 2x MI325X, $4.00 |
| Memory at 4-bit (weights and margin) | 192.8 GB | 108 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI325X, $2.00 | 1x MI300X, $1.85 |
| Revision viewed | eb9eb208eb0d | de4b8e4d43b9 |
| Downloads reported by the hub | 2.7M | 706.1k |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
Other Reported Results
These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.
GLM-5.3-Flash
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| PaddlePaddle/Real5-OmniDocBench | Task illuminationMetric illuminationComparison conditions not established | 91.4 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task overallMetric overallComparison conditions not established | 90.76 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task scanningMetric scanningComparison conditions not established | 91.43 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task screen_photographyMetric screen_photographyComparison conditions not established | 90.36 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task skewMetric skewComparison conditions not established | 90.76 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| PaddlePaddle/Real5-OmniDocBench | Task warpingMetric warpingComparison conditions not established | 89.84 | Real5-OmniDocBench Leaderboard Reported by a third party |
Evaluated revision not stated | 2026-09-12 |
| cais/hle | Task hleMetric hleSetup HLE with tools (full set) and a 300K-context management strategy, not the no-tools default; judged by GPT-5.6-luna (medium).Comparison conditions not established | 55.3 | GLM-5.3-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| datacurve/deep-swe | Task deep_sweMetric deep_sweSetup Reported as DeepSWE v1.1 on the model card, run via the mini-swe-agent harness with 400K context.Comparison conditions not established | 63.4 | GLM-5.3-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| harborframework/terminal-bench-2.1 | Task terminalbench_2_1Metric terminalbench_2_1Comparison conditions not established | 84.3 | GLM-5.3-Flash model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ExtractBench | Task longMetric longSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established | 27.83 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ExtractBench | Task meanMetric meanSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established | 80.75 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ExtractBench | Task mediumMetric mediumSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established | 51.56 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ExtractBench | Task shortMetric shortSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established | 96.3 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task chartMetric chartSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 54.96 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task layoutMetric layoutSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 41.78 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task meanMetric meanSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 70.75 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task tableMetric tableSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 89.18 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task text_contentMetric text_contentSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 88.34 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| llamaindex/ParseBench | Task text_formattingMetric text_formattingSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established | 79.49 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
Qwen3.8-Flash-Next
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Idavidrein/gpqa | Task diamondMetric diamondComparison conditions not established | 91.7 | Qwen3.8-Flash-Next model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| SWE-bench/SWE-bench_Multilingual | Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedSetup harness: mini-SWE-agent. 256K context.Comparison conditions not established | 81 | Qwen3.8-Flash-Next model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| ScaleAI/SWE-bench_Pro | Task SWE_Bench_ProMetric SWE_Bench_ProSetup harness: Claude Code. 256K context; evaluated on the card's own refined version of the benchmark (problematic tasks corrected).Comparison conditions not established | 62.5 | Qwen3.8-Flash-Next model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| cais/hle | Task hleMetric hleSetup Judged by GPT-4o per the card, not the task's default o3-mini grader.Comparison conditions not established | 35.9 | Qwen3.8-Flash-Next model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| claw-eval/Claw-Eval | Task multimodalMetric multimodalSetup Reported as 'ClawEval-MM' on the card; Pass@3 metric (average-score alternative was 60.4).Comparison conditions not established | 64.4 | Qwen3.8-Flash-Next model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| datacurve/deep-swe | Task deep_sweMetric deep_sweSetup harness: Claude Code / mini-SWE-agent (best of two). Reported as 'DeepSWE 1.1' on the card; 256K context.Comparison conditions not established | 58.7 | Qwen3.8-Flash-Next model card Reported by a third party |
Evaluated revision not stated | 2026-08-26 |
| hkust-nlp/Toolathlon | Task toolathlon_verifiedMetric toolathlon_verifiedComparison conditions not established | 73.5 | Qwen/Qwen3.8-Flash-Next model card Reported by a third party |
Evaluated revision not stated | 2026-08-27 |
| llamaindex/ExtractBench | Task longMetric longSetup Pipeline name: qwen3_8_flash_next_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-Flash-Next-FP8Comparison conditions not established | 37.74 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-29 |
| llamaindex/ExtractBench | Task meanMetric meanSetup Pipeline name: qwen3_8_flash_next_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-Flash-Next-FP8Comparison conditions not established | 89.88 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-29 |
| llamaindex/ExtractBench | Task mediumMetric mediumSetup Pipeline name: qwen3_8_flash_next_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-Flash-Next-FP8Comparison conditions not established | 87.81 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-29 |
| llamaindex/ExtractBench | Task shortMetric shortSetup Pipeline name: qwen3_8_flash_next_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-Flash-Next-FP8Comparison conditions not established | 94.82 | ExtractBench Reported by a third party |
Evaluated revision not stated | 2026-08-29 |
SAVRN's Notes on GLM-5.3-Flash
Only 18B of the 321.3B parameters fire on any token, but all 288 routed experts sit in memory regardless, so the 16-bit footprint is 771.2 GB on three 288 GB MI355X cards at $7.77 per hour, 8-bit is 385.6 GB on two MI325X at $4.00, and 4-bit is 192.8 GB on one 256 GB MI325X at $2.00. Z.ai's first natively multimodal GLM-5 model: image and text in, text out, 1,048,576-token context.
Under MIT the obligation is one line, keep the copyright and permission notice with the files; commercial use, modification and redistribution are allowed. That million-token window needs memory left after the weights; on one 256 GB card at 4-bit little is. Baseten, DeepInfra, Fireworks, Novita and Together AI list it on the Index at $0.15 in and $0.50 out per million tokens, so set your tokens per hour against $2.00, $4.00 or $7.77 and read the ledger.
SAVRN's Notes on Qwen3.8-Flash-Next
Precision decides the hardware on this one. At 16-bit, 360 GB of weights need a 432 GB working footprint, and the cheapest way we price that is two MI325X cards with 256 GB each at $4.00 per hour on-demand. At 8-bit the footprint drops to 216 GB and fits one MI325X at $2.00 per hour; at 4-bit it is 108 GB and a single MI300X at $1.85 per hour carries it. It takes image and text in and answers in text across a 262,144-token context, with 180B parameters spread over 512 experts, 10 active per token.
The license is listed as other with no summary on file, so read the publisher's terms before any commercial deployment. Access is open: 144 files, 360 GB, safetensors. The publisher calls it an experimental preview of the architecture behind Qwen4, so pin the revision you test; the files were last updated 2026-08-27.
Questions
Which is larger, GLM-5.3-Flash or Qwen3.8-Flash-Next?
GLM-5.3-Flash (321.3B parameters) is larger than Qwen3.8-Flash-Next (180B parameters), by the parameter counts their publishers report.
Which is cheaper to run, GLM-5.3-Flash or Qwen3.8-Flash-Next?
At 4-bit, GLM-5.3-Flash fits on 1x MI325X from $2.00 an hour and Qwen3.8-Flash-Next on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use GLM-5.3-Flash commercially?
Yes. GLM-5.3-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.