SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

GLM-5.3-Flash

by Z.ai zai-org/GLM-5.3-Flash

Join our WeChat or Discord community. Check out the GLM-5.3-Flash blog and GLM-5 Technical report. Use GLM-5.3-Flash API services on Z.ai API Platform. We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series.

Parameters321.3B
Context1,048,576
Weights328.3 GB
Licensemit
AccessOpen weights
Monthly Downloads2.7M

Runs On

What it takes to serve GLM-5.3-Flash (321.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 642.6 GB 771.2 GB 3x MI355X (288 GB)
Vultr
$7.77 4x MI325X $8.00 · 5x MI300X $9.25
8-bit 321.3 GB 385.6 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55
4-bit 160.7 GB 192.8 GB 1x MI325X (256 GB)
Vultr
$2.00 1x MI355X $2.59 · 2x MI300X $3.70

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on GLM-5.3-Flash

Only 18B of the 321.3B parameters fire on any token, but all 288 routed experts sit in memory regardless, so the 16-bit footprint is 771.2 GB on three 288 GB MI355X cards at $7.77 per hour, 8-bit is 385.6 GB on two MI325X at $4.00, and 4-bit is 192.8 GB on one 256 GB MI325X at $2.00. Z.ai's first natively multimodal GLM-5 model: image and text in, text out, 1,048,576-token context.

Under MIT the obligation is one line, keep the copyright and permission notice with the files; commercial use, modification and redistribution are allowed. That million-token window needs memory left after the weights; on one 256 GB card at 4-bit little is. Baseten, DeepInfra, Fireworks, Novita and Together AI list it on the Index at $0.15 in and $0.50 out per million tokens, so set your tokens per hour against $2.00, $4.00 or $7.77 and read the ledger.

Model Card

By Z.ai, published under mit, revision eb9eb208eb0d.

Join ourWeChat or Discord community.
Check out the GLM-5.3-Flashblog and GLM-5 Technical report.
Use GLM-5.3-Flash API services onZ.ai API Platform.

Introduction

We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.

Serve GLM-5.3-Flash Locally

Read the full model card (1,052 words)

Configuration

Architecture
Glm5NextForConditionalGeneration
Context length (tokens)
1,048,576
Layers
45
Hidden size
4,096
Feed-forward size
12,288
Attention heads
64
Key/value heads
64
Head dimension
0
Vocabulary size
154,880
Routed experts
288
Experts active per token
8
Model type
glm5_next
Quantization
fp8

Identity and Version

Repository
zai-org/GLM-5.3-Flash
Publisher
Z.ai
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
321.3B parameters
Languages
en, zh
Revision
eb9eb208eb0d988989d07a6a12d0fdeb5f52574a
First published
2026-08-25
Last updated
2026-09-07

Files and Weights

73 files, 328.4 GB in total. The weights are 62 files totalling 328.3 GB in safetensors.

Weights62 files · 328.3 GB
Configuration5 files · 8.5 MB
Tokenizer2 files · 20.2 MB
Documentation2 files · 9.1 KB
Other1 file · 10.9 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00062.safetensorsWeights5.4 GB 9ff3c9397e75
model-00002-of-00062.safetensorsWeights5.3 GB 328225cabdc9
model-00003-of-00062.safetensorsWeights5.4 GB e0fc42c270a8
model-00004-of-00062.safetensorsWeights5.4 GB ea66d9139b41
model-00005-of-00062.safetensorsWeights5.4 GB d1da1e65d621
model-00006-of-00062.safetensorsWeights5.4 GB 1653c5e3d9af
model-00007-of-00062.safetensorsWeights5.4 GB 82396e0f55ab
model-00008-of-00062.safetensorsWeights5.4 GB 91f5cb329f6d
model-00009-of-00062.safetensorsWeights5.4 GB 69cbb58b336f
model-00010-of-00062.safetensorsWeights5.4 GB db5ea14fed53
model-00011-of-00062.safetensorsWeights5.4 GB 408046268e30
model-00012-of-00062.safetensorsWeights5.4 GB dc61a5684e5f
model-00013-of-00062.safetensorsWeights5.4 GB bfecdc0fd8d9
model-00014-of-00062.safetensorsWeights5.4 GB bab323a2bb17
model-00015-of-00062.safetensorsWeights5.4 GB e18e404b26f8
model-00016-of-00062.safetensorsWeights5.4 GB 6cc97434dfac
model-00017-of-00062.safetensorsWeights5.4 GB 530a4034da64
model-00018-of-00062.safetensorsWeights5.4 GB 035ee33c3279
model-00019-of-00062.safetensorsWeights5.4 GB 28b540f19710
model-00020-of-00062.safetensorsWeights5.4 GB 67a86a8bedaa
model-00021-of-00062.safetensorsWeights5.4 GB 48c8bbefadcf
model-00022-of-00062.safetensorsWeights5.4 GB 3b0232707a75
model-00023-of-00062.safetensorsWeights5.4 GB 1ea8f1bbc7e1
model-00024-of-00062.safetensorsWeights5.4 GB 4cda3a7dcfbf
model-00025-of-00062.safetensorsWeights5.4 GB 7527b7182a3f
model-00026-of-00062.safetensorsWeights5.4 GB 29aa015bbd8d
model-00027-of-00062.safetensorsWeights5.4 GB fdafd6e70645
model-00028-of-00062.safetensorsWeights5.4 GB 07089bdfe63c
model-00029-of-00062.safetensorsWeights5.4 GB 4013ec692173
model-00030-of-00062.safetensorsWeights5.4 GB a28adbc99919
model-00031-of-00062.safetensorsWeights5.4 GB e2f31cb37dcf
model-00032-of-00062.safetensorsWeights5.4 GB f33b53853ea5
model-00033-of-00062.safetensorsWeights5.4 GB a090ee4e5aee
model-00034-of-00062.safetensorsWeights5.4 GB 451bb0368e52
model-00035-of-00062.safetensorsWeights5.4 GB bc98812806e9
model-00036-of-00062.safetensorsWeights5.4 GB 19803f680722
model-00037-of-00062.safetensorsWeights5.4 GB 0e7c51bd69d8
model-00038-of-00062.safetensorsWeights5.4 GB e903f70ca5e9
model-00039-of-00062.safetensorsWeights5.4 GB eacebc6414c3
model-00040-of-00062.safetensorsWeights5.4 GB 6f7c05711e1d
model-00041-of-00062.safetensorsWeights5.4 GB 9ddfde38b0f0
model-00042-of-00062.safetensorsWeights5.4 GB 2b70ffbe9d82
model-00043-of-00062.safetensorsWeights5.4 GB f5fede3a7555
model-00044-of-00062.safetensorsWeights5.4 GB ed57a93696b5
model-00045-of-00062.safetensorsWeights5.4 GB 7f280939f4c8
model-00046-of-00062.safetensorsWeights5.4 GB 9e748de0fd06
model-00047-of-00062.safetensorsWeights5.4 GB ddf0691ca8aa
model-00048-of-00062.safetensorsWeights5.4 GB 85c512009854
model-00049-of-00062.safetensorsWeights5.4 GB 14ba26c9f6b0
model-00050-of-00062.safetensorsWeights5.4 GB 7a00d5221b41
model-00051-of-00062.safetensorsWeights5.4 GB b3423a93d9d4
model-00052-of-00062.safetensorsWeights5.4 GB 9162deb48ee6
model-00053-of-00062.safetensorsWeights5.4 GB 2f5ecc0c3f96
model-00054-of-00062.safetensorsWeights5.4 GB 1ae7351f071a
model-00055-of-00062.safetensorsWeights5.4 GB 8d64c91ff600
model-00056-of-00062.safetensorsWeights5.4 GB 748dda91c9b6
model-00057-of-00062.safetensorsWeights5.4 GB e128c2da3d3e
model-00058-of-00062.safetensorsWeights5.4 GB ac4e58d16001
model-00059-of-00062.safetensorsWeights5.4 GB bc1731273adc
model-00060-of-00062.safetensorsWeights5.4 GB 2cf6465538f8
model-00061-of-00062.safetensorsWeights5.3 GB 4c29e5654cab
model-00062-of-00062.safetensorsWeights1.3 GB d3087816db95
.eval_results/GLM-5.3-Flash.yamlConfiguration845 B
config.jsonConfiguration69.4 KB
generation_config.jsonConfiguration194 B
model.safetensors.index.jsonConfiguration8.4 MB
processor_config.jsonConfiguration909 B
LICENSEDocumentation1.1 KB
README.mdDocumentation8.0 KB
chat_template.jinjaOther10.9 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer20.2 MB 19e773648cb4
tokenizer_config.jsonTokenizer761 B

License and Download

License
mit
Access
Open weights, no gate
Download size
328.3 GB
Download from Z.ai

Released by Z.ai through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
PaddlePaddle/Real5-OmniDocBench Task illuminationMetric illuminationComparison conditions not established 91.4 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task overallMetric overallComparison conditions not established 90.76 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task scanningMetric scanningComparison conditions not established 91.43 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task screen_photographyMetric screen_photographyComparison conditions not established 90.36 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task skewMetric skewComparison conditions not established 90.76 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
PaddlePaddle/Real5-OmniDocBench Task warpingMetric warpingComparison conditions not established 89.84 Real5-OmniDocBench Leaderboard
Reported by a third party
Evaluated revision not stated 2026-09-12
cais/hle Task hleMetric hleSetup HLE with tools (full set) and a 300K-context management strategy, not the no-tools default; judged by GPT-5.6-luna (medium).Comparison conditions not established 55.3 GLM-5.3-Flash model card
Reported by a third party
Evaluated revision not stated 2026-08-26
datacurve/deep-swe Task deep_sweMetric deep_sweSetup Reported as DeepSWE v1.1 on the model card, run via the mini-swe-agent harness with 400K context.Comparison conditions not established 63.4 GLM-5.3-Flash model card
Reported by a third party
Evaluated revision not stated 2026-08-26
harborframework/terminal-bench-2.1 Task terminalbench_2_1Metric terminalbench_2_1Comparison conditions not established 84.3 GLM-5.3-Flash model card
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task longMetric longSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established 27.83 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task meanMetric meanSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established 80.75 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task mediumMetric mediumSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established 51.56 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task shortMetric shortSetup Pipeline name: glm_5_3_flash_extract_oneshot_structured_output_fileComparison conditions not established 96.3 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task chartMetric chartSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 54.96 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task layoutMetric layoutSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 41.78 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task meanMetric meanSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 70.75 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task tableMetric tableSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 89.18 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task text_contentMetric text_contentSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 88.34 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task text_formattingMetric text_formattingSetup Pipeline name: glm_5_3_flash_parse_with_layout_fileComparison conditions not established 79.49 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-26

Memory Requirements

PrecisionWeights in memory
As published328.3 GB
16-bit642.6 GB
8-bit321.3 GB
4-bit160.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
Baseten$0.15 / $0.50input / output, per million tokensSep 18, 2026
DeepInfra$0.15 / $0.50input / output, per million tokensSep 18, 2026
Fireworks$0.15 / $0.50input / output, per million tokensSep 18, 2026
Novita$0.15 / $0.50input / output, per million tokensSep 18, 2026
Together AI$0.15 / $0.50input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Built on This Model

Compare GLM-5.3-Flash

Questions About GLM-5.3-Flash

How much GPU memory does GLM-5.3-Flash need?

About 771.2 GB at 16-bit and 192.8 GB at 4-bit: the weights (321.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run GLM-5.3-Flash on?

At 16-bit, 3x MI355X from $7.77 an hour; at 4-bit, 1x MI325X from $2.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use GLM-5.3-Flash commercially?

Yes. GLM-5.3-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is GLM-5.3-Flash's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.5-397B-A17B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. MathVision:our model’s score is…

Open weights apache-2.0 403.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next

Qwen

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next-Uncensored-NVFP4

OrcaRouter

This model has had its safety alignment substantially removed via abliteration (orthogonalizing the refusal direction out of the residual stream). It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment. Use must comply with the Apache 2.0 License inherited from the base model and all applicable law. The authors accept no liability for misuse. - A Blackwell GPU…

Access requested at publisher apache-2.0 180B parameters transformers

Model · Image and text to text

Qwen3.8-Flash-Next-MLX-oQ3-MTP

Robot Haus

A sensitivity-guided, mixed-precision MLX conversion of Qwen/Qwen3.8-Flash-Next, rebuilt directly from the official BF16 checkpoint with the model's matching native MTP block preserved. oQ3 uses a 3-bit affine base and spends additional precision on sensitive modules. Layer sensitivity was measured with a validated quantized calibration proxy, while every released weight was quantized from the official BF16 checkpoint. The result is a compact model with 746 higher-precision module overrides rather than a uniform 3-bit layout. The upstream tokenizer, current chat template, vision processor, generation configuration, licence, and native MTP configuration are retained. In a compatible oMLX…

Open weights other 180B parameters 262,144 tokens mlx

Model · Image and text to text

Qwen3.8-Flash-Next

Tai Hua

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-122B-A10B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 125.1B parameters 262,144 tokens transformers