SAVRN
Search Contact SAVRN

SAVRN Model Hub · Comparisons

GLM-5.3-Flash vs Qwen3.8-Flash-Next

GLM-5.3-Flash has 321.3B parameters and Qwen3.8-Flash-Next has 180B parameters; GLM-5.3-Flash is released under MIT License and Qwen3.8-Flash-Next under other; at 16-bit, GLM-5.3-Flash needs about 771.2 GB (3x MI355X from $7.77 an hour) and Qwen3.8-Flash-Next about 432 GB (2x MI325X from $4.00 an hour).

Published metadata for 2 models, each read from its own repository.
Field GLM-5.3-Flash
zai-org/GLM-5.3-Flash
Qwen3.8-Flash-Next
Qwen/Qwen3.8-Flash-Next
Publisher Z.ai Qwen
Task Image and text to text Image and text to text
Modality Image and text Image and text
Parameters, as reported 321.3B parameters 180B parameters
Architecture Glm5NextForConditionalGeneration Qwen4ExpForConditionalGeneration
Library transformers transformers
Context length 1,048,576 tokens 262,144 tokens
Repository size 328.4 GB 360.0 GB
Artifact formats safetensors safetensors
License mit other
Access Open weights, no gate Open weights, no gate
Memory at 16-bit (weights and margin) 771.2 GB 432 GB
Cheapest GPUs at 16-bit, per hour 3x MI355X, $7.77 2x MI325X, $4.00
Memory at 4-bit (weights and margin) 192.8 GB 108 GB
Cheapest GPUs at 4-bit, per hour 1x MI325X, $2.00 1x MI300X, $1.85
Revision viewed eb9eb208eb0d de4b8e4d43b9
Downloads reported by the hub 2.7M 706.1k
Last observed 2026-09-18 2026-09-18

An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.

Other Reported Results

These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.

GLM-5.3-Flash

BenchmarkConditionsResultReported byRevisionDate
PaddlePaddle/Real5-OmniDocBench Task illuminationMetric illuminationComparison conditions not established 91.4 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task overallMetric overallComparison conditions not established 90.76 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task scanningMetric scanningComparison conditions not established 91.43 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task screen_photographyMetric screen_photographyComparison conditions not established 90.36 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task skewMetric skewComparison conditions not established 90.76 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task warpingMetric warpingComparison conditions not established 89.84 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
cais/hle Task hleMetric hleSetup HLE with tools (full set) and a 300K-context management strategy, not the no-tools default; judged by GPT-5.6-luna (medium).Comparison conditions not established 55.3 GLM-5.3-Flash model card
Reported by a third party
Evaluated revision not stated 2026-08-26
datacurve/deep-swe Task deep_sweMetric deep_sweSetup Reported as DeepSWE v1.1 on the model card, run via the mini-swe-agent harness with 400K context.Comparison conditions not established 63.4 GLM-5.3-Flash model card
Reported by a third party
Evaluated revision not stated 2026-08-26
harborframework/terminal-bench-2.1 Task terminalbench_2_1Metric terminalbench_2_1Comparison conditions not established 84.3 GLM-5.3-Flash model card
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task longMetric longSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established 27.83 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task meanMetric meanSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established 80.75 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task mediumMetric mediumSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established 51.56 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task shortMetric shortSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established 96.3 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task chartMetric chartSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 54.96 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task layoutMetric layoutSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 41.78 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task meanMetric meanSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 70.75 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task tableMetric tableSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 89.18 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task text_contentMetric text_contentSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 88.34 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task text_formattingMetric text_formattingSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 79.49 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26

Qwen3.8-Flash-Next

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 91.7 Qwen3.8-Flash-Next model card
Reported by a third party
Evaluated revision not stated 2026-08-26
SWE-bench/SWE-bench_Multilingual Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedSetup harness: mini-SWE-agent. 256K context.Comparison conditions not established 81 Qwen3.8-Flash-Next model card
Reported by a third party
Evaluated revision not stated 2026-08-26
ScaleAI/SWE-bench_Pro Task SWE_Bench_ProMetric SWE_Bench_ProSetup harness: Claude Code. 256K context; evaluated on the card's own refined version of the benchmark (problematic tasks corrected).Comparison conditions not established 62.5 Qwen3.8-Flash-Next model card
Reported by a third party
Evaluated revision not stated 2026-08-26
cais/hle Task hleMetric hleSetup Judged by GPT-4o per the card, not the task's default o3-mini grader.Comparison conditions not established 35.9 Qwen3.8-Flash-Next model card
Reported by a third party
Evaluated revision not stated 2026-08-26
claw-eval/Claw-Eval Task multimodalMetric multimodalSetup Reported as 'ClawEval-MM' on the card; Pass@3 metric (average-score alternative was 60.4).Comparison conditions not established 64.4 Qwen3.8-Flash-Next model card
Reported by a third party
Evaluated revision not stated 2026-08-26
datacurve/deep-swe Task deep_sweMetric deep_sweSetup harness: Claude Code / mini-SWE-agent (best of two). Reported as 'DeepSWE 1.1' on the card; 256K context.Comparison conditions not established 58.7 Qwen3.8-Flash-Next model card
Reported by a third party
Evaluated revision not stated 2026-08-26
hkust-nlp/Toolathlon Task toolathlon_verifiedMetric toolathlon_verifiedComparison conditions not established 73.5 Qwen/Qwen3.8-Flash-Next model card
Reported by a third party
Evaluated revision not stated 2026-08-27
llamaindex/ExtractBench Task longMetric longSetup Pipeline name: qwen3_8_flash_next_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-Flash-Next-FP8Comparison conditions not established 37.74 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-29
llamaindex/ExtractBench Task meanMetric meanSetup Pipeline name: qwen3_8_flash_next_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-Flash-Next-FP8Comparison conditions not established 89.88 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-29
llamaindex/ExtractBench Task mediumMetric mediumSetup Pipeline name: qwen3_8_flash_next_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-Flash-Next-FP8Comparison conditions not established 87.81 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-29
llamaindex/ExtractBench Task shortMetric shortSetup Pipeline name: qwen3_8_flash_next_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-Flash-Next-FP8Comparison conditions not established 94.82 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-29

SAVRN's Notes on GLM-5.3-Flash

Only 18B of the 321.3B parameters fire on any token, but all 288 routed experts sit in memory regardless, so the 16-bit footprint is 771.2 GB on three 288 GB MI355X cards at $7.77 per hour, 8-bit is 385.6 GB on two MI325X at $4.00, and 4-bit is 192.8 GB on one 256 GB MI325X at $2.00. Z.ai's first natively multimodal GLM-5 model: image and text in, text out, 1,048,576-token context.

Under MIT the obligation is one line, keep the copyright and permission notice with the files; commercial use, modification and redistribution are allowed. That million-token window needs memory left after the weights; on one 256 GB card at 4-bit little is. Baseten, DeepInfra, Fireworks, Novita and Together AI list it on the Index at $0.15 in and $0.50 out per million tokens, so set your tokens per hour against $2.00, $4.00 or $7.77 and read the ledger.

SAVRN's Notes on Qwen3.8-Flash-Next

Precision decides the hardware on this one. At 16-bit, 360 GB of weights need a 432 GB working footprint, and the cheapest way we price that is two MI325X cards with 256 GB each at $4.00 per hour on-demand. At 8-bit the footprint drops to 216 GB and fits one MI325X at $2.00 per hour; at 4-bit it is 108 GB and a single MI300X at $1.85 per hour carries it. It takes image and text in and answers in text across a 262,144-token context, with 180B parameters spread over 512 experts, 10 active per token.

The license is listed as other with no summary on file, so read the publisher's terms before any commercial deployment. Access is open: 144 files, 360 GB, safetensors. The publisher calls it an experimental preview of the architecture behind Qwen4, so pin the revision you test; the files were last updated 2026-08-27.

Questions

Which is larger, GLM-5.3-Flash or Qwen3.8-Flash-Next?

GLM-5.3-Flash (321.3B parameters) is larger than Qwen3.8-Flash-Next (180B parameters), by the parameter counts their publishers report.

Which is cheaper to run, GLM-5.3-Flash or Qwen3.8-Flash-Next?

At 4-bit, GLM-5.3-Flash fits on 1x MI325X from $2.00 an hour and Qwen3.8-Flash-Next on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use GLM-5.3-Flash commercially?

Yes. GLM-5.3-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Related Comparisons