SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DeepSeek-R1

by DeepSeek deepseek-ai/DeepSeek-R1

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1.

Parameters684.5B
Context163,840
Weights688.6 GB
Licensemit
AccessOpen weights
Monthly Downloads741k

Runs On

What it takes to serve DeepSeek-R1 (684.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1369.0 GB 1642.8 GB 7x MI325X (256 GB)
Vultr
$14.00 6x MI355X $15.54 · 7x B300 $46.20
8-bit 684.5 GB 821.4 GB 3x MI355X (288 GB)
Vultr
$7.77 4x MI325X $8.00 · 5x MI300X $9.25
4-bit 342.2 GB 410.7 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on DeepSeek-R1

Plan around 410.7 GB for DeepSeek-R1. That is what the 4-bit build needs, and the cheapest cover is two MI325X cards at 256 GB each, $4.00 an hour on demand. The 8-bit build doubles the need to 821.4 GB and three MI355X cards at $7.77 an hour; full 16-bit weights take 1,642.8 GB across seven MI325X cards at $14.00. Only 8 of the 256 routed experts work on each token, so the 684.5 billion parameters are mostly a memory bill, and the 163,840-token context is what that memory buys you.

MIT terms permit commercial use, modification and redistribution with the notice attached, so a fine-tuned internal copy can be run and shipped. Before buying, set the $4.00-an-hour floor against Novita's SAVRN Index rate of $0.70 in and $2.50 out per million tokens, and budget the 688.6 GB, 174-file transfer.

Model Card

By DeepSeek, published under mit, revision 56d4cbbb4d29.

Paper Link

1. Introduction

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across math, code, and reasoning tasks. To support the research community, we have open-sourced DeepSeek-R1-Zero, DeepSeek-R1, and six dense models distilled from DeepSeek-R1 based on Llama and Qwen. DeepSeek-R1-Distill-Qwen-32B outperforms OpenAI-o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.

NOTE: Before running DeepSeek-R1 series models locally, we kindly recommend reviewing the Usage Recommendation section.

2. Model Summary

Post-Training: Large-Scale Reinforcement Learning on the Base Model

Read the full model card (1,629 words)

Configuration

Architecture
DeepseekV3ForCausalLM
Context length (tokens)
163,840
Layers
61
Hidden size
7,168
Feed-forward size
18,432
Attention heads
128
Key/value heads
128
Vocabulary size
129,280
Routed experts
256
Experts active per token
8
RoPE base
10,000
Stored precision
bfloat16
Model type
deepseek_v3
Quantization
fp8

Identity and Version

Repository
deepseek-ai/DeepSeek-R1
Publisher
DeepSeek
Task
Text generation
Modality
Text
Library
transformers
Parameters
684.5B parameters
Languages
Not stated by the source
Revision
56d4cbbb4d29f4355bab4b9a39ccb717a14ad5ad
First published
2025-01-20
Last updated
2025-03-27

Files and Weights

174 files, 688.6 GB in total. The weights are 163 files totalling 688.6 GB in safetensors.

Weights163 files · 688.6 GB
Configuration5 files · 9.0 MB
Tokenizer2 files · 7.9 MB
Documentation2 files · 17.1 KB
Other1 file · 777.3 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-000163.safetensorsWeights5.2 GB c2388e6b127c
model-00002-of-000163.safetensorsWeights4.3 GB 5f450c75da7e
model-00003-of-000163.safetensorsWeights4.3 GB c79b34183573
model-00004-of-000163.safetensorsWeights4.3 GB ad5ca32dcb6c
model-00005-of-000163.safetensorsWeights4.3 GB e06768f65df4
model-00006-of-000163.safetensorsWeights4.4 GB 3eb84d165db0
model-00007-of-000163.safetensorsWeights4.3 GB d6f299f7b410
model-00008-of-000163.safetensorsWeights4.3 GB da197fe67d24
model-00009-of-000163.safetensorsWeights4.3 GB dbd2a790cc3d
model-00010-of-000163.safetensorsWeights4.3 GB 137f0f040729
model-00011-of-000163.safetensorsWeights4.3 GB 47b308ea2f7f
model-00012-of-000163.safetensorsWeights1.3 GB 4cda99d04536
model-00013-of-000163.safetensorsWeights4.3 GB 33adbb52d544
model-00014-of-000163.safetensorsWeights4.3 GB 05b19f1bd146
model-00015-of-000163.safetensorsWeights4.3 GB 0484ca365f84
model-00016-of-000163.safetensorsWeights4.3 GB 1a70800c3373
model-00017-of-000163.safetensorsWeights4.3 GB bd5abae258a5
model-00018-of-000163.safetensorsWeights4.3 GB b5418a2987ad
model-00019-of-000163.safetensorsWeights4.3 GB 69bbd687eafe
model-00020-of-000163.safetensorsWeights4.3 GB 2560e8782214
model-00021-of-000163.safetensorsWeights4.3 GB 09e6825f850b
model-00022-of-000163.safetensorsWeights4.3 GB 083a0025a915
model-00023-of-000163.safetensorsWeights4.3 GB a10ddf9357a7
model-00024-of-000163.safetensorsWeights4.3 GB 5f13210eca8e
model-00025-of-000163.safetensorsWeights4.3 GB d511a1879fd6
model-00026-of-000163.safetensorsWeights4.3 GB 0c50460d2ee1
model-00027-of-000163.safetensorsWeights4.3 GB 260ad43d8099
model-00028-of-000163.safetensorsWeights4.3 GB 1acbd0856de2
model-00029-of-000163.safetensorsWeights4.3 GB 20d3ef472a46
model-00030-of-000163.safetensorsWeights4.3 GB ac1d12d32f3a
model-00031-of-000163.safetensorsWeights4.3 GB eace4ef31c4f
model-00032-of-000163.safetensorsWeights4.3 GB ac7f5250660b
model-00033-of-000163.safetensorsWeights4.3 GB f71549c7a33e
model-00034-of-000163.safetensorsWeights1.7 GB dc52ac9cd64b
model-00035-of-000163.safetensorsWeights4.3 GB 5f4677233a2a
model-00036-of-000163.safetensorsWeights4.3 GB f0b62c567e3c
model-00037-of-000163.safetensorsWeights4.3 GB 9d7c8a1c35d1
model-00038-of-000163.safetensorsWeights4.3 GB ea241680b942
model-00039-of-000163.safetensorsWeights4.3 GB 6dfd35cf8163
model-00040-of-000163.safetensorsWeights4.3 GB fc4c69c0e91e
model-00041-of-000163.safetensorsWeights4.3 GB 333c52631a3d
model-00042-of-000163.safetensorsWeights4.3 GB 450f75fd01ff
model-00043-of-000163.safetensorsWeights4.3 GB aab6deabc711
model-00044-of-000163.safetensorsWeights4.3 GB c9d6582f530a
model-00045-of-000163.safetensorsWeights4.3 GB 938c7306ce56
model-00046-of-000163.safetensorsWeights4.3 GB e13e11c8bacd
model-00047-of-000163.safetensorsWeights4.3 GB 3ae5dc07da78
model-00048-of-000163.safetensorsWeights4.3 GB df85e837c896
model-00049-of-000163.safetensorsWeights4.3 GB 207e8af06dfc
model-00050-of-000163.safetensorsWeights4.3 GB 7dfd5bcb9a53
model-00051-of-000163.safetensorsWeights4.3 GB fe8ba090ec48
model-00052-of-000163.safetensorsWeights4.3 GB 03945249e7af
model-00053-of-000163.safetensorsWeights4.3 GB 1a764331dbf2
model-00054-of-000163.safetensorsWeights4.3 GB a13688518c52
model-00055-of-000163.safetensorsWeights4.3 GB bb473f1ccda3
model-00056-of-000163.safetensorsWeights1.7 GB 210545cb12fb
model-00057-of-000163.safetensorsWeights4.3 GB 3b35f6380a25
model-00058-of-000163.safetensorsWeights4.3 GB 28d7ed16f520
model-00059-of-000163.safetensorsWeights4.3 GB 6edacb001b24
model-00060-of-000163.safetensorsWeights4.3 GB fef1246c153b
model-00061-of-000163.safetensorsWeights4.3 GB f88ba0e9b666
model-00062-of-000163.safetensorsWeights4.3 GB 009804d0502a
model-00063-of-000163.safetensorsWeights4.3 GB 9f762a03e483
model-00064-of-000163.safetensorsWeights4.3 GB cdbe9c409a1a
model-00065-of-000163.safetensorsWeights4.3 GB 9849c2c30fa0
model-00066-of-000163.safetensorsWeights4.3 GB 88a3879e8d89
model-00067-of-000163.safetensorsWeights4.3 GB 7d4c58f31a14
model-00068-of-000163.safetensorsWeights4.3 GB 910927c6e17d
model-00069-of-000163.safetensorsWeights4.3 GB bcfacce03108
model-00070-of-000163.safetensorsWeights4.3 GB 276a349ba4ce
model-00071-of-000163.safetensorsWeights4.3 GB c550537013ea
model-00072-of-000163.safetensorsWeights4.3 GB 39f8920d14e9
model-00073-of-000163.safetensorsWeights4.3 GB 960e55ecc062
model-00074-of-000163.safetensorsWeights4.3 GB 3fcd1a5bf85d
model-00075-of-000163.safetensorsWeights4.3 GB 66e03af200bf
model-00076-of-000163.safetensorsWeights4.3 GB 43cfb5095eb8
model-00077-of-000163.safetensorsWeights4.3 GB 1411532f8eab
model-00078-of-000163.safetensorsWeights1.7 GB f8ad3219d64a
model-00079-of-000163.safetensorsWeights4.3 GB 1d74c9e9383b
model-00080-of-000163.safetensorsWeights4.3 GB edfb4c02a4a6
model-00081-of-000163.safetensorsWeights4.3 GB 9b1e9d08aba3
model-00082-of-000163.safetensorsWeights4.3 GB d8c4620641dd
model-00083-of-000163.safetensorsWeights4.3 GB c273f728b6c9
model-00084-of-000163.safetensorsWeights4.3 GB 424e18ad35df
model-00085-of-000163.safetensorsWeights4.3 GB b1690a7567a6
model-00086-of-000163.safetensorsWeights4.3 GB e65eb5913f33
model-00087-of-000163.safetensorsWeights4.3 GB 2dc4e3d4751f
model-00088-of-000163.safetensorsWeights4.3 GB 54940d661422
model-00089-of-000163.safetensorsWeights4.3 GB 38e2520475fd
model-00090-of-000163.safetensorsWeights4.3 GB ee933f59ddd2
model-00091-of-000163.safetensorsWeights4.3 GB 4bacfdad2f2c
model-00092-of-000163.safetensorsWeights4.3 GB 74fd6b1a280b
model-00093-of-000163.safetensorsWeights4.3 GB 714adb4ffc41
model-00094-of-000163.safetensorsWeights4.3 GB b3bcdaa226ac
model-00095-of-000163.safetensorsWeights4.3 GB 1e0ad4783bf4
model-00096-of-000163.safetensorsWeights4.3 GB dac3287c0ab7
model-00097-of-000163.safetensorsWeights4.3 GB 17bd26dc53c5
model-00098-of-000163.safetensorsWeights4.3 GB 0c1ccb0c95ec
model-00099-of-000163.safetensorsWeights4.3 GB aa3fe3e6fb2e
model-00100-of-000163.safetensorsWeights1.7 GB 300556119da4
model-00101-of-000163.safetensorsWeights4.3 GB d3c4705db628
model-00102-of-000163.safetensorsWeights4.3 GB 945d2f833dfc
model-00103-of-000163.safetensorsWeights4.3 GB db5d341e101a
model-00104-of-000163.safetensorsWeights4.3 GB 0d7766179014
model-00105-of-000163.safetensorsWeights4.3 GB f1466e6a5fd8
model-00106-of-000163.safetensorsWeights4.3 GB 84f318ee430b
model-00107-of-000163.safetensorsWeights4.3 GB 6b37a206419d
model-00108-of-000163.safetensorsWeights4.3 GB 6b6e4d233218
model-00109-of-000163.safetensorsWeights4.3 GB 4d67e9fd9d13
model-00110-of-000163.safetensorsWeights4.3 GB 5a42f528fabf
model-00111-of-000163.safetensorsWeights4.3 GB 1eff3d3853ef
model-00112-of-000163.safetensorsWeights4.3 GB 0fe9b6e8662d
model-00113-of-000163.safetensorsWeights4.3 GB 8c1b3e7e97cb
model-00114-of-000163.safetensorsWeights4.3 GB dddf3e5c7d5d
model-00115-of-000163.safetensorsWeights4.3 GB 13368ae6cba0
model-00116-of-000163.safetensorsWeights4.3 GB 3e933d3d9c42
model-00117-of-000163.safetensorsWeights4.3 GB f74cdd2b2799
model-00118-of-000163.safetensorsWeights4.3 GB 8e27a9cb1d1c
model-00119-of-000163.safetensorsWeights4.3 GB a46e4aa463db
model-00120-of-000163.safetensorsWeights4.3 GB af132a07e3e8
model-00121-of-000163.safetensorsWeights4.3 GB b5cf7a403d84
model-00122-of-000163.safetensorsWeights1.7 GB 7337345a219a
model-00123-of-000163.safetensorsWeights4.3 GB d50d3129ef6f
model-00124-of-000163.safetensorsWeights4.3 GB 6f207cfa394f
model-00125-of-000163.safetensorsWeights4.3 GB 222ef3a2c582
model-00126-of-000163.safetensorsWeights4.3 GB 2f6ba5a5471d
model-00127-of-000163.safetensorsWeights4.3 GB 42cde7a7cc25
model-00128-of-000163.safetensorsWeights4.3 GB 98fba66ba204
model-00129-of-000163.safetensorsWeights4.3 GB 3379b2eec5c2
model-00130-of-000163.safetensorsWeights4.3 GB 4ef03f44980f
model-00131-of-000163.safetensorsWeights4.3 GB 0199456e2b1a
model-00132-of-000163.safetensorsWeights4.3 GB 962b8f44e5f2
model-00133-of-000163.safetensorsWeights4.3 GB 4b6729c4e655
model-00134-of-000163.safetensorsWeights4.3 GB 1b3b36d497b9
model-00135-of-000163.safetensorsWeights4.3 GB d29025efeda5
model-00136-of-000163.safetensorsWeights4.3 GB 2a8117c3aa75
model-00137-of-000163.safetensorsWeights4.3 GB 3e13e0244f29
model-00138-of-000163.safetensorsWeights4.3 GB a18aa50aebf7
model-00139-of-000163.safetensorsWeights4.3 GB 9728520094bf
model-00140-of-000163.safetensorsWeights4.3 GB 11418dfbacad
model-00141-of-000163.safetensorsWeights3.1 GB dd1c282ddba1
model-00142-of-000163.safetensorsWeights4.3 GB 89790f24021d
model-00143-of-000163.safetensorsWeights4.3 GB ff5a31006f8d
model-00144-of-000163.safetensorsWeights4.3 GB b76fd57f1662
model-00145-of-000163.safetensorsWeights4.3 GB e1738a3c47df
model-00146-of-000163.safetensorsWeights4.3 GB 981ed532393e
model-00147-of-000163.safetensorsWeights4.3 GB e5f87a816f13
model-00148-of-000163.safetensorsWeights4.3 GB 6b47933264dd
model-00149-of-000163.safetensorsWeights4.3 GB 22a2f518d7bc
model-00150-of-000163.safetensorsWeights4.3 GB 4b4ca747261d
model-00151-of-000163.safetensorsWeights4.3 GB 8a7c9787942c
model-00152-of-000163.safetensorsWeights4.3 GB e6adb0bb8071
model-00153-of-000163.safetensorsWeights4.3 GB fe3867b4e246
model-00154-of-000163.safetensorsWeights4.3 GB 78fdc257544f
model-00155-of-000163.safetensorsWeights4.3 GB 834632439e88
model-00156-of-000163.safetensorsWeights4.3 GB 4beb34c5173f
model-00157-of-000163.safetensorsWeights4.3 GB c313d28a29f8
model-00158-of-000163.safetensorsWeights4.3 GB f0448fba1ddb
model-00159-of-000163.safetensorsWeights4.3 GB 77a6ec3076e5
model-00160-of-000163.safetensorsWeights5.2 GB f621991fbc9c
model-00161-of-000163.safetensorsWeights4.3 GB e1e75d7f8086
model-00162-of-000163.safetensorsWeights4.3 GB 53e717e22e28
model-00163-of-000163.safetensorsWeights6.6 GB 913177d9e0df
config.jsonConfiguration1.7 KB
configuration_deepseek.pyConfiguration9.9 KB
generation_config.jsonConfiguration171 B
model.safetensors.index.jsonConfiguration8.9 MB
modeling_deepseek.pyConfiguration75.7 KB
LICENSEDocumentation1.1 KB
README.mdDocumentation16.0 KB
figures/benchmark.jpgOther777.3 KB
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer7.8 MB
tokenizer_config.jsonTokenizer3.6 KB

License and Download

License
mit
Access
Open weights, no gate
Download size
688.6 GB
Download from DeepSeek

Released by DeepSeek through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 71.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-01-27
LEXam-Benchmark/LEXam Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established 52.41 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
LEXam-Benchmark/LEXam Task open_questionMetric open_questionComparison conditions not established 55.91 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 84 Model Card
Reported by a third party
Evaluated revision not stated 2026-01-28

Memory Requirements

PrecisionWeights in memory
As published688.6 GB
16-bit1369.0 GB
8-bit684.5 GB
4-bit342.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
Novita$0.70 / $2.50input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Questions About DeepSeek-R1

How much GPU memory does DeepSeek-R1 need?

About 1642.8 GB at 16-bit and 410.7 GB at 4-bit: the weights (684.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-R1 on?

At 16-bit, 7x MI325X from $14.00 an hour; at 4-bit, 2x MI325X from $4.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-R1 commercially?

Yes. DeepSeek-R1 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-R1's context length?

163,840 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

DeepSeek-V3

DeepSeek

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations…

Open weights 684.5B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-V3-0324

DeepSeek

DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects. - More aesthetically pleasing web pages and game front-ends - Enhanced report analysis requests with more detailed outputs - Increased accuracy in Function Calling, fixing issues from previous V3 versions In the official DeepSeek web/app, we use the same system prompt with a specific date. For example, In our web and application environments, the temperature parameter $T{model}$ is set to 0.3. Because many users use the default temperature 1.0 in API call, we have implemented an API temperature $T{api}$ mapping mechanism that adjusts the input API temperature value of 1.0 to the…

Open weights mit 684.5B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-V3.2

DeepSeek

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs: 1. DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance, specifically optimized for long-context scenarios. 2. Scalable Reinforcement Learning Framework: By implementing a robust RL protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par…

Open weights mit 685.4B parameters 163,840 tokens transformers

Model · Text generation

GLM-5.2-FP8

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.2

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.3

Z.ai

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the Ascend NPU platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. - GLM-5.3 supports controlling the thinking budget through the reasoningeffort parameter, which accepts three levels: low, high, and max. It defaults to max…

Open weights other 753.3B parameters 1,048,576 tokens transformers