SAVRN
Search Contact SAVRN

Open-weight model · Text generation

GLM-5.3

by Z.ai zai-org/GLM-5.3

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks.

Parameters753.3B
Context1,048,576
Weights755.6 GB
Licenseother
AccessOpen weights
Monthly Downloads838.4k

Runs On

What it takes to serve GLM-5.3 (753.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1506.7 GB 1808.0 GB 8x MI325X (256 GB)
Vultr
$16.00 7x MI355X $18.13 · 7x B300 $46.20
8-bit 753.3 GB 904.0 GB 4x MI325X (256 GB)
Vultr
$8.00 5x MI300X $9.25 · 4x MI355X $10.36
4-bit 376.7 GB 452.0 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on GLM-5.3

Four hundred fifty-two gigabytes is the floor for GLM-5.3 at 4-bit, and the cheapest way there is two MI325X cards for $4.00 per hour. At 8-bit it wants 904 GB and four cards at $8.00; at 16-bit, 1,808 GB and eight at $16.00. Z.ai says the base is the same one under GLM-5.2, with the gains from post-training for complex coding and long-horizon work across a 1,048,576-token context. With a multi-GPU node running SGLang or vLLM, this comes in-house; otherwise you are renting.

The license is listed as other with no summary, so get the actual text before deciding. Weigh that against the Index hosts: DeepInfra at $1.20 in and $4.00 out per million tokens, and Baseten, Fireworks, Novita and Together AI at $1.40 and $4.40. Test the thinking budget control first, since output tokens are the expensive side on every host.

Model Card

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the Ascend NPU platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. - GLM-5.3 supports controlling the thinking budget through the reasoningeffort parameter, which accepts three levels: low, high, and max. It defaults to max…

Excerpt from the card by Z.ai, licensed other.

Configuration

Architecture
GlmMoeDsaForCausalLM
Context length (tokens)
1,048,576
Layers
78
Hidden size
6,144
Feed-forward size
12,288
Attention heads
64
Key/value heads
64
Head dimension
192
Vocabulary size
154,880
Routed experts
256
Experts active per token
8
Model type
glm_moe_dsa
Quantization
fp8

Identity and Version

Repository
zai-org/GLM-5.3
Publisher
Z.ai
Task
Text generation
Modality
Text
Library
transformers
Parameters
753.3B parameters
Languages
en, zh
Revision
aca966e4e02791568aa6a4ced368624b3d897f42
First published
2026-08-25
Last updated
2026-09-04

Files and Weights

155 files, 755.7 GB in total. The weights are 141 files totalling 755.6 GB in safetensors.

Weights141 files · 755.6 GB
Configuration8 files · 11.4 MB
Tokenizer2 files · 20.2 MB
Documentation2 files · 18.5 KB
Other1 file · 10.7 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00141.safetensorsWeights5.4 GB 29c537abddf4
model-00002-of-00141.safetensorsWeights5.4 GB dedd05754d90
model-00003-of-00141.safetensorsWeights5.4 GB a8b0aac7fdc8
model-00004-of-00141.safetensorsWeights5.4 GB 43155b1b8fd8
model-00005-of-00141.safetensorsWeights5.4 GB 7ef7a74c4d73
model-00006-of-00141.safetensorsWeights5.4 GB d6d12ad601a6
model-00007-of-00141.safetensorsWeights5.4 GB a5930833775b
model-00008-of-00141.safetensorsWeights5.4 GB c70bd3904324
model-00009-of-00141.safetensorsWeights5.4 GB c3174f361cc9
model-00010-of-00141.safetensorsWeights5.4 GB 52fc7f5a00da
model-00011-of-00141.safetensorsWeights5.4 GB b0fb8066c93b
model-00012-of-00141.safetensorsWeights5.4 GB b00d5dc53dc7
model-00013-of-00141.safetensorsWeights5.4 GB b5c86c0a3691
model-00014-of-00141.safetensorsWeights5.4 GB 97c98b26dd67
model-00015-of-00141.safetensorsWeights5.4 GB 08c11bb21650
model-00016-of-00141.safetensorsWeights5.4 GB 2496b285f611
model-00017-of-00141.safetensorsWeights5.4 GB 0642d34c281b
model-00018-of-00141.safetensorsWeights5.4 GB 1755443ec0ac
model-00019-of-00141.safetensorsWeights5.4 GB 4bd772671cce
model-00020-of-00141.safetensorsWeights5.4 GB b29f71186de8
model-00021-of-00141.safetensorsWeights5.4 GB 52a10d45444b
model-00022-of-00141.safetensorsWeights5.4 GB 51e81bc5beaa
model-00023-of-00141.safetensorsWeights5.4 GB 7b5bee00ed2e
model-00024-of-00141.safetensorsWeights5.4 GB 3aa762520255
model-00025-of-00141.safetensorsWeights5.4 GB 29031699298a
model-00026-of-00141.safetensorsWeights5.4 GB d4d5c305f39b
model-00027-of-00141.safetensorsWeights5.4 GB 2652b27a61aa
model-00028-of-00141.safetensorsWeights5.4 GB 2aadd7705973
model-00029-of-00141.safetensorsWeights5.4 GB 8f7f418937f8
model-00030-of-00141.safetensorsWeights5.4 GB 1d9c553bc388
model-00031-of-00141.safetensorsWeights5.4 GB d16a2c2d98a4
model-00032-of-00141.safetensorsWeights5.4 GB 3d890300d1c4
model-00033-of-00141.safetensorsWeights5.4 GB 503a2745947e
model-00034-of-00141.safetensorsWeights5.4 GB 436240414a63
model-00035-of-00141.safetensorsWeights5.4 GB fda066767b80
model-00036-of-00141.safetensorsWeights5.4 GB 43294735fb32
model-00037-of-00141.safetensorsWeights5.4 GB 7c9e2ecb1c45
model-00038-of-00141.safetensorsWeights5.4 GB e97f6e122331
model-00039-of-00141.safetensorsWeights5.4 GB fb6f9c9987cf
model-00040-of-00141.safetensorsWeights5.4 GB 77db405ecf43
model-00041-of-00141.safetensorsWeights5.4 GB bb4f9019124d
model-00042-of-00141.safetensorsWeights5.4 GB 16a7cdda8b21
model-00043-of-00141.safetensorsWeights5.4 GB 3f8bf5154605
model-00044-of-00141.safetensorsWeights5.4 GB 1d1645412f18
model-00045-of-00141.safetensorsWeights5.4 GB 88226cffca5b
model-00046-of-00141.safetensorsWeights5.4 GB 7efae0d6d515
model-00047-of-00141.safetensorsWeights5.4 GB b377da726711
model-00048-of-00141.safetensorsWeights5.4 GB 8dfedd9a4abf
model-00049-of-00141.safetensorsWeights5.4 GB 5af58c8a0043
model-00050-of-00141.safetensorsWeights5.4 GB 132bc9c74406
model-00051-of-00141.safetensorsWeights5.4 GB c6c2b274fdad
model-00052-of-00141.safetensorsWeights5.4 GB 5cb7eaaeca53
model-00053-of-00141.safetensorsWeights5.4 GB 54972554c009
model-00054-of-00141.safetensorsWeights5.4 GB 03fa148543df
model-00055-of-00141.safetensorsWeights5.4 GB af95d9f47a6e
model-00056-of-00141.safetensorsWeights5.4 GB caa384d78da8
model-00057-of-00141.safetensorsWeights5.4 GB b12ee91dfcc4
model-00058-of-00141.safetensorsWeights5.4 GB 168dc3d43287
model-00059-of-00141.safetensorsWeights5.4 GB 518e10cc0120
model-00060-of-00141.safetensorsWeights5.4 GB feda1de42cc4
model-00061-of-00141.safetensorsWeights5.4 GB bf5f4cac3681
model-00062-of-00141.safetensorsWeights5.4 GB 230563a0ce9a
model-00063-of-00141.safetensorsWeights5.4 GB ad9f7f2d988c
model-00064-of-00141.safetensorsWeights5.4 GB 7adf78e82363
model-00065-of-00141.safetensorsWeights5.4 GB 679c665e9ada
model-00066-of-00141.safetensorsWeights5.4 GB 81e07ea3663a
model-00067-of-00141.safetensorsWeights5.4 GB 4893e7714579
model-00068-of-00141.safetensorsWeights5.4 GB 98fc8956ad13
model-00069-of-00141.safetensorsWeights5.4 GB bee7792fea84
model-00070-of-00141.safetensorsWeights5.4 GB 9f61f7fb310c
model-00071-of-00141.safetensorsWeights5.4 GB 6012dfecf3f6
model-00072-of-00141.safetensorsWeights5.4 GB 260db65aac41
model-00073-of-00141.safetensorsWeights5.4 GB 8f5e929e8206
model-00074-of-00141.safetensorsWeights5.4 GB 27c1ccde7476
model-00075-of-00141.safetensorsWeights5.4 GB fa483d44e064
model-00076-of-00141.safetensorsWeights5.4 GB 77cf94357920
model-00077-of-00141.safetensorsWeights5.4 GB dfc0832bb33d
model-00078-of-00141.safetensorsWeights5.4 GB 4c536713f0d3
model-00079-of-00141.safetensorsWeights5.4 GB 899b86d6f672
model-00080-of-00141.safetensorsWeights5.4 GB bb9cec7efbc4
model-00081-of-00141.safetensorsWeights5.4 GB fa599e07f0d4
model-00082-of-00141.safetensorsWeights5.4 GB 6dd9121cd2fe
model-00083-of-00141.safetensorsWeights5.4 GB 6c515be1a654
model-00084-of-00141.safetensorsWeights5.4 GB 1d59ac473eeb
model-00085-of-00141.safetensorsWeights5.4 GB fdb9c8d5026c
model-00086-of-00141.safetensorsWeights5.4 GB aaca5f2ddd5b
model-00087-of-00141.safetensorsWeights5.4 GB 7958dda4bcc9
model-00088-of-00141.safetensorsWeights5.4 GB 12142bbf24a5
model-00089-of-00141.safetensorsWeights5.4 GB dc4e3c065c71
model-00090-of-00141.safetensorsWeights5.4 GB 5953bca43c6f
model-00091-of-00141.safetensorsWeights5.4 GB f6abe4b934d3
model-00092-of-00141.safetensorsWeights5.4 GB 4a56f759b410
model-00093-of-00141.safetensorsWeights5.4 GB d8da10c49c13
model-00094-of-00141.safetensorsWeights5.4 GB 7a3f836bd20e
model-00095-of-00141.safetensorsWeights5.4 GB 34488a283558
model-00096-of-00141.safetensorsWeights5.4 GB 24e80bb3f655
model-00097-of-00141.safetensorsWeights5.4 GB 6faaa6f276cd
model-00098-of-00141.safetensorsWeights5.4 GB 3f6ba5deab11
model-00099-of-00141.safetensorsWeights5.4 GB 3ddd0f314489
model-00100-of-00141.safetensorsWeights5.4 GB a68b875a8bb2
model-00101-of-00141.safetensorsWeights5.4 GB 3a5d30bac831
model-00102-of-00141.safetensorsWeights5.4 GB 01873cd29e26
model-00103-of-00141.safetensorsWeights5.4 GB 628c1817216f
model-00104-of-00141.safetensorsWeights5.4 GB 7641a66c311d
model-00105-of-00141.safetensorsWeights5.4 GB 5d5b6fed5cb6
model-00106-of-00141.safetensorsWeights5.4 GB b3843515722d
model-00107-of-00141.safetensorsWeights5.4 GB 46f3230ecb54
model-00108-of-00141.safetensorsWeights5.4 GB b7b62322bf48
model-00109-of-00141.safetensorsWeights5.4 GB 0a595c9d83a9
model-00110-of-00141.safetensorsWeights5.4 GB d671cca9545e
model-00111-of-00141.safetensorsWeights5.4 GB c4214b75400e
model-00112-of-00141.safetensorsWeights5.4 GB aa9965212e74
model-00113-of-00141.safetensorsWeights5.4 GB 1c7242a0eb57
model-00114-of-00141.safetensorsWeights5.4 GB 69d98d0b412b
model-00115-of-00141.safetensorsWeights5.4 GB df690e0bd353
model-00116-of-00141.safetensorsWeights5.4 GB 03904100c9a8
model-00117-of-00141.safetensorsWeights5.4 GB bd7d01dd2697
model-00118-of-00141.safetensorsWeights5.4 GB b008875bf308
model-00119-of-00141.safetensorsWeights5.4 GB d954e2505d28
model-00120-of-00141.safetensorsWeights5.4 GB c12905e36603
model-00121-of-00141.safetensorsWeights5.4 GB 0be5012d19d8
model-00122-of-00141.safetensorsWeights5.4 GB 5a73f6134a57
model-00123-of-00141.safetensorsWeights5.4 GB 61f76624bb2f
model-00124-of-00141.safetensorsWeights5.4 GB 39c94c6de9f0
model-00125-of-00141.safetensorsWeights5.4 GB ce597eead0b4
model-00126-of-00141.safetensorsWeights5.4 GB dd05c147ad93
model-00127-of-00141.safetensorsWeights5.4 GB ccd37d509020
model-00128-of-00141.safetensorsWeights5.4 GB d1a5a126e029
model-00129-of-00141.safetensorsWeights5.4 GB 9d8f5dbbbd1e
model-00130-of-00141.safetensorsWeights5.4 GB eeab5728f01b
model-00131-of-00141.safetensorsWeights5.4 GB c1904484c321
model-00132-of-00141.safetensorsWeights5.4 GB 54d258bddf2d
model-00133-of-00141.safetensorsWeights5.4 GB a829a7a92936
model-00134-of-00141.safetensorsWeights5.4 GB e39b80363535
model-00135-of-00141.safetensorsWeights5.4 GB 880037c72bff
model-00136-of-00141.safetensorsWeights5.4 GB 26d141a174b7
model-00137-of-00141.safetensorsWeights5.4 GB 4558ee4c2692
model-00138-of-00141.safetensorsWeights5.4 GB f6c7894ac9e5
model-00139-of-00141.safetensorsWeights5.4 GB 996571ec96e0
model-00140-of-00141.safetensorsWeights5.4 GB 0aed808a4aa4
model-00141-of-00141.safetensorsWeights4.7 GB 83b5ab7fb7f7
.eval_results/deep-swe.yamlConfiguration153 B
.eval_results/hle.yamlConfiguration161 B
.eval_results/terminal-bench-2.1.yamlConfiguration210 B
.eval_results/terminal-bench-3.0.yamlConfiguration204 B
.eval_results/zai-org__GLM-5.3.yamlConfiguration205 B
config.jsonConfiguration29.5 KB
generation_config.jsonConfiguration194 B
model.safetensors.index.jsonConfiguration11.4 MB e0fe7f28c1f8
LICENSEDocumentation4.3 KB
README.mdDocumentation14.2 KB
chat_template.jinjaOther10.7 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer20.2 MB 19e773648cb4
tokenizer_config.jsonTokenizer761 B

License and Download

License
other
Access
Open weights, no gate
Download size
755.6 GB
Download from Z.ai

Released by Z.ai through its official repository on Hugging Face.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
cais/hle Task hleMetric hleSetup With tools.Comparison conditions not established 62.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-29
datacurve/deep-swe Task deep_sweMetric deep_sweComparison conditions not established 66.9 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-28
harborframework/terminal-bench Task terminalbench_3Metric terminalbench_3Setup harness: claude codeComparison conditions not established 28.3 Model Card
Reported by a third party
Evaluated revision not stated 2026-09-04
harborframework/terminal-bench-2.1 Task terminalbench_2_1Metric terminalbench_2_1Setup harness: claude codeComparison conditions not established 88.2 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-31
hkust-nlp/Toolathlon Task toolathlon_verifiedMetric toolathlon_verifiedComparison conditions not established 73 zai-org/GLM-5.3 model card
Reported by a third party
Evaluated revision not stated 2026-08-31

Memory Requirements

PrecisionWeights in memory
As published755.6 GB
16-bit1506.7 GB
8-bit753.3 GB
4-bit376.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
Baseten$1.40 / $4.40input / output, per million tokensSep 18, 2026
DeepInfra$1.20 / $4.00input / output, per million tokensSep 18, 2026
Fireworks$1.40 / $4.40input / output, per million tokensSep 18, 2026
Novita$1.40 / $4.40input / output, per million tokensSep 18, 2026
Together AI$1.40 / $4.40input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Questions About GLM-5.3

How much GPU memory does GLM-5.3 need?

About 1808 GB at 16-bit and 452 GB at 4-bit: the weights (753.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run GLM-5.3 on?

At 16-bit, 8x MI325X from $16.00 an hour; at 4-bit, 2x MI325X from $4.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is GLM-5.3 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is GLM-5.3's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

GLM-5.2-FP8

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.2

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

DeepSeek-V3.2

DeepSeek

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs: 1. DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance, specifically optimized for long-context scenarios. 2. Scalable Reinforcement Learning Framework: By implementing a robust RL protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par…

Open weights mit 685.4B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-V3

DeepSeek

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations…

Open weights 684.5B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-V3-0324

DeepSeek

DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects. - More aesthetically pleasing web pages and game front-ends - Enhanced report analysis requests with more detailed outputs - Increased accuracy in Function Calling, fixing issues from previous V3 versions In the official DeepSeek web/app, we use the same system prompt with a specific date. For example, In our web and application environments, the temperature parameter $T{model}$ is set to 0.3. Because many users use the default temperature 1.0 in API call, we have implemented an API temperature $T{api}$ mapping mechanism that adjusts the input API temperature value of 1.0 to the…

Open weights mit 684.5B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-R1

DeepSeek

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across…

Open weights mit 684.5B parameters 163,840 tokens transformers