SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DeepSeek-V3-0324

by DeepSeek deepseek-ai/DeepSeek-V3-0324

DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.

Parameters684.5B
Context163,840
Weights688.6 GB
Licensemit
AccessOpen weights
Monthly Downloads1M

Runs On

What it takes to serve DeepSeek-V3-0324 (684.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1369.1 GB 1642.9 GB 7x MI325X (256 GB)
Vultr
$14.00 6x MI355X $15.54 · 7x B300 $46.20
8-bit 684.5 GB 821.4 GB 3x MI355X (288 GB)
Vultr
$7.77 4x MI325X $8.00 · 5x MI300X $9.25
4-bit 342.3 GB 410.7 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on DeepSeek-V3-0324

How far you are willing to quantize sets the hardware bill. At 16-bit the working memory is 1,642.9 GB, met by seven MI325X at 256 GB each for $14.00 an hour. At 8-bit it is 821.4 GB on three MI355X at $7.77, and at 4-bit 410.7 GB on two MI325X at $4.00. Only 8 of 256 routed experts fire per token, yet every expert has to be resident, so size for the memory figure; 163,840 tokens of context is what you get for it. The files total 688.6 GB; budget the storage and the transfer before the first GPU hour.

MIT is about as short as licenses get: commercial use, modification and redistribution, keep the notices. Before committing, price your own volume against the Index host, DeepInfra at $0.24 in and $0.90 out per million tokens. The one benchmark on file, 81.2 on MMLU-Pro, is the publisher's own number, not independently verified.

Model Card

By DeepSeek, published under mit, revision e9b33add7688.

Features

DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.

Reasoning Capabilities

  • Significant improvements in benchmark performance:
  • MMLU-Pro: 75.9 → 81.2 (+5.3)
  • GPQA: 59.1 → 68.4 (+9.3)
  • AIME: 39.6 → 59.4 (+19.8)
  • LiveCodeBench: 39.2 → 49.2 (+10.0)

Front-End Web Development

  • Improved the executability of the code
  • More aesthetically pleasing web pages and game front-ends

Chinese Writing Proficiency

  • Enhanced style and content quality:
  • Aligned with the R1 writing style
  • Better quality in medium-to-long-form writing

  • Feature Enhancements

  • Improved multi-turn interactive rewriting
  • Optimized translation quality and letter writing

Chinese Search Capabilities

  • Enhanced report analysis requests with more detailed outputs

Function Calling Improvements

  • Increased accuracy in Function Calling, fixing issues from previous V3 versions

Usage Recommendations

System Prompt

In the official DeepSeek web/app, we use the same system prompt with a specific date.

该助手为DeepSeek Chat,由深度求索公司创造。
今天是{current date}。

For example,

该助手为DeepSeek Chat,由深度求索公司创造。
今天是3月24日,星期一。

Temperature

Read the full model card (843 words)

Configuration

Architecture
DeepseekV3ForCausalLM
Context length (tokens)
163,840
Layers
61
Hidden size
7,168
Feed-forward size
18,432
Attention heads
128
Key/value heads
128
Vocabulary size
129,280
Routed experts
256
Experts active per token
8
RoPE base
10,000
Stored precision
bfloat16
Model type
deepseek_v3
Quantization
fp8

Identity and Version

Repository
deepseek-ai/DeepSeek-V3-0324
Publisher
DeepSeek
Task
Text generation
Modality
Text
Library
transformers
Parameters
684.5B parameters
Languages
Not stated by the source
Revision
e9b33add76883f293d6bf61f6bd89b497e80e335
First published
2025-03-24
Last updated
2025-03-27

Files and Weights

173 files, 688.6 GB in total. The weights are 163 files totalling 688.6 GB in safetensors.

Weights163 files · 688.6 GB
Configuration4 files · 9.0 MB
Tokenizer2 files · 7.9 MB
Documentation2 files · 11.7 KB
Other1 file · 292.1 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-000163.safetensorsWeights5.2 GB 134f51f4642d
model-00002-of-000163.safetensorsWeights4.3 GB 400b76ca4599
model-00003-of-000163.safetensorsWeights4.3 GB 1e0192d491a2
model-00004-of-000163.safetensorsWeights4.3 GB 6f8ca069011e
model-00005-of-000163.safetensorsWeights4.3 GB 1973c53f4920
model-00006-of-000163.safetensorsWeights4.4 GB 2250eb7be1d2
model-00007-of-000163.safetensorsWeights4.3 GB 8632d0847cd3
model-00008-of-000163.safetensorsWeights4.3 GB 85daa7fd14a6
model-00009-of-000163.safetensorsWeights4.3 GB b55f6542adae
model-00010-of-000163.safetensorsWeights4.3 GB 0317ba8c8544
model-00011-of-000163.safetensorsWeights4.3 GB c5291900ff83
model-00012-of-000163.safetensorsWeights1.3 GB 1596e384e287
model-00013-of-000163.safetensorsWeights4.3 GB 0b5b8ac70def
model-00014-of-000163.safetensorsWeights4.3 GB df727377a86d
model-00015-of-000163.safetensorsWeights4.3 GB ae96924ba8d4
model-00016-of-000163.safetensorsWeights4.3 GB 506313cbaf4e
model-00017-of-000163.safetensorsWeights4.3 GB d396bde693ed
model-00018-of-000163.safetensorsWeights4.3 GB 4ebc50b2e0e2
model-00019-of-000163.safetensorsWeights4.3 GB 0a9d59229d19
model-00020-of-000163.safetensorsWeights4.3 GB ac27c6cbc271
model-00021-of-000163.safetensorsWeights4.3 GB 8afd4121eeb8
model-00022-of-000163.safetensorsWeights4.3 GB d6e5d4ea01bc
model-00023-of-000163.safetensorsWeights4.3 GB 940662482556
model-00024-of-000163.safetensorsWeights4.3 GB 488969a9e8a3
model-00025-of-000163.safetensorsWeights4.3 GB 450cc3f29910
model-00026-of-000163.safetensorsWeights4.3 GB bdfe115c31c7
model-00027-of-000163.safetensorsWeights4.3 GB 13e981a9bdc7
model-00028-of-000163.safetensorsWeights4.3 GB d767507070d1
model-00029-of-000163.safetensorsWeights4.3 GB 11b1ef947996
model-00030-of-000163.safetensorsWeights4.3 GB 6ff7740ee870
model-00031-of-000163.safetensorsWeights4.3 GB 07afa8b68a43
model-00032-of-000163.safetensorsWeights4.3 GB 361f93b0bcbf
model-00033-of-000163.safetensorsWeights4.3 GB d5cd3a3bd61d
model-00034-of-000163.safetensorsWeights1.7 GB df4c5f0efee2
model-00035-of-000163.safetensorsWeights4.3 GB 59907496ec35
model-00036-of-000163.safetensorsWeights4.3 GB 816a8da95174
model-00037-of-000163.safetensorsWeights4.3 GB 03da5c0e525f
model-00038-of-000163.safetensorsWeights4.3 GB c0e6e082c73b
model-00039-of-000163.safetensorsWeights4.3 GB a8dba03ca01e
model-00040-of-000163.safetensorsWeights4.3 GB 2ebb24adff0b
model-00041-of-000163.safetensorsWeights4.3 GB 5a50db6c7cb7
model-00042-of-000163.safetensorsWeights4.3 GB 9666e834327b
model-00043-of-000163.safetensorsWeights4.3 GB dda5691da3ed
model-00044-of-000163.safetensorsWeights4.3 GB 743f7e5979b3
model-00045-of-000163.safetensorsWeights4.3 GB 929fa255d587
model-00046-of-000163.safetensorsWeights4.3 GB 3ae202c6641b
model-00047-of-000163.safetensorsWeights4.3 GB e3c9d4d06ab0
model-00048-of-000163.safetensorsWeights4.3 GB b5dbbd1df4a0
model-00049-of-000163.safetensorsWeights4.3 GB 308925e3cd15
model-00050-of-000163.safetensorsWeights4.3 GB b97524be8e16
model-00051-of-000163.safetensorsWeights4.3 GB 68fa813be7c3
model-00052-of-000163.safetensorsWeights4.3 GB 806f360b5b46
model-00053-of-000163.safetensorsWeights4.3 GB 9a8867bd41c3
model-00054-of-000163.safetensorsWeights4.3 GB 274d3dd5e855
model-00055-of-000163.safetensorsWeights4.3 GB c4978652ad09
model-00056-of-000163.safetensorsWeights1.7 GB 68130d200659
model-00057-of-000163.safetensorsWeights4.3 GB 3b4826e9817a
model-00058-of-000163.safetensorsWeights4.3 GB 80630144922c
model-00059-of-000163.safetensorsWeights4.3 GB 432a6a2fbb10
model-00060-of-000163.safetensorsWeights4.3 GB 59f81c7cef16
model-00061-of-000163.safetensorsWeights4.3 GB 69cde84d3893
model-00062-of-000163.safetensorsWeights4.3 GB 715c4c8739ff
model-00063-of-000163.safetensorsWeights4.3 GB bf0f5c6e778f
model-00064-of-000163.safetensorsWeights4.3 GB d99efa8b327f
model-00065-of-000163.safetensorsWeights4.3 GB 898c045b5e06
model-00066-of-000163.safetensorsWeights4.3 GB 01e069367ad8
model-00067-of-000163.safetensorsWeights4.3 GB f0cbc915d403
model-00068-of-000163.safetensorsWeights4.3 GB 2594adae97d7
model-00069-of-000163.safetensorsWeights4.3 GB 01daa8eb84a9
model-00070-of-000163.safetensorsWeights4.3 GB 1f0368ad214a
model-00071-of-000163.safetensorsWeights4.3 GB 0cc594400988
model-00072-of-000163.safetensorsWeights4.3 GB 6940c2e5c775
model-00073-of-000163.safetensorsWeights4.3 GB 3fc8ab4be0f4
model-00074-of-000163.safetensorsWeights4.3 GB 9651fe2a556d
model-00075-of-000163.safetensorsWeights4.3 GB f74e66e34acd
model-00076-of-000163.safetensorsWeights4.3 GB 986f6a11367f
model-00077-of-000163.safetensorsWeights4.3 GB 6b0e946a0d1d
model-00078-of-000163.safetensorsWeights1.7 GB 9835f6d79440
model-00079-of-000163.safetensorsWeights4.3 GB 8dfe2c683545
model-00080-of-000163.safetensorsWeights4.3 GB 5afc8aa8b94c
model-00081-of-000163.safetensorsWeights4.3 GB 757fc061d21e
model-00082-of-000163.safetensorsWeights4.3 GB 0eadd45aa5af
model-00083-of-000163.safetensorsWeights4.3 GB 095348fdecd5
model-00084-of-000163.safetensorsWeights4.3 GB ffcf0052d53f
model-00085-of-000163.safetensorsWeights4.3 GB 7b210ddf7172
model-00086-of-000163.safetensorsWeights4.3 GB 99d1090b9423
model-00087-of-000163.safetensorsWeights4.3 GB b89cf8ea275a
model-00088-of-000163.safetensorsWeights4.3 GB 7219a03969a1
model-00089-of-000163.safetensorsWeights4.3 GB e57eef5e6eb8
model-00090-of-000163.safetensorsWeights4.3 GB b98655409b66
model-00091-of-000163.safetensorsWeights4.3 GB a2c5d3f06613
model-00092-of-000163.safetensorsWeights4.3 GB f9e99f4a9745
model-00093-of-000163.safetensorsWeights4.3 GB bc7300167ed6
model-00094-of-000163.safetensorsWeights4.3 GB 6cb0e4055409
model-00095-of-000163.safetensorsWeights4.3 GB f1387857bd9d
model-00096-of-000163.safetensorsWeights4.3 GB b58d7c50e54b
model-00097-of-000163.safetensorsWeights4.3 GB fa678aaed884
model-00098-of-000163.safetensorsWeights4.3 GB a2cd517b5675
model-00099-of-000163.safetensorsWeights4.3 GB 06d5aa90ee2e
model-00100-of-000163.safetensorsWeights1.7 GB 90e85dc99298
model-00101-of-000163.safetensorsWeights4.3 GB 9976dae2a44b
model-00102-of-000163.safetensorsWeights4.3 GB cf36bd6e7a27
model-00103-of-000163.safetensorsWeights4.3 GB f8e08c2159b5
model-00104-of-000163.safetensorsWeights4.3 GB f1bb0364aa51
model-00105-of-000163.safetensorsWeights4.3 GB 5541411c798d
model-00106-of-000163.safetensorsWeights4.3 GB 5a2d163f3ca1
model-00107-of-000163.safetensorsWeights4.3 GB 5a7d97379fce
model-00108-of-000163.safetensorsWeights4.3 GB 2813892afb5b
model-00109-of-000163.safetensorsWeights4.3 GB 70b59cb0fe6e
model-00110-of-000163.safetensorsWeights4.3 GB ce860b0acca5
model-00111-of-000163.safetensorsWeights4.3 GB 912c0c99c138
model-00112-of-000163.safetensorsWeights4.3 GB 8d9989812d88
model-00113-of-000163.safetensorsWeights4.3 GB 02c14d13f644
model-00114-of-000163.safetensorsWeights4.3 GB 881e73a71172
model-00115-of-000163.safetensorsWeights4.3 GB 5a66f0fde0ff
model-00116-of-000163.safetensorsWeights4.3 GB 1102daefb72e
model-00117-of-000163.safetensorsWeights4.3 GB d92157af0d4a
model-00118-of-000163.safetensorsWeights4.3 GB ba30d60916a7
model-00119-of-000163.safetensorsWeights4.3 GB 8938205f9d55
model-00120-of-000163.safetensorsWeights4.3 GB 18fa9b624779
model-00121-of-000163.safetensorsWeights4.3 GB 669570df0b21
model-00122-of-000163.safetensorsWeights1.7 GB 9a9a82578c3a
model-00123-of-000163.safetensorsWeights4.3 GB 84723e0f5c3f
model-00124-of-000163.safetensorsWeights4.3 GB 01e9467b7094
model-00125-of-000163.safetensorsWeights4.3 GB 562c9706e1b1
model-00126-of-000163.safetensorsWeights4.3 GB 949da3c082f9
model-00127-of-000163.safetensorsWeights4.3 GB 8871e7ca5651
model-00128-of-000163.safetensorsWeights4.3 GB cb8a5a4d53a6
model-00129-of-000163.safetensorsWeights4.3 GB 9a49ccda2e63
model-00130-of-000163.safetensorsWeights4.3 GB 00ceb885b7ad
model-00131-of-000163.safetensorsWeights4.3 GB 2af936134d19
model-00132-of-000163.safetensorsWeights4.3 GB 7ceabb5bfd0b
model-00133-of-000163.safetensorsWeights4.3 GB 213c0ae67ab8
model-00134-of-000163.safetensorsWeights4.3 GB d2cd64a8c397
model-00135-of-000163.safetensorsWeights4.3 GB 4c05694e1d64
model-00136-of-000163.safetensorsWeights4.3 GB ca04d3839513
model-00137-of-000163.safetensorsWeights4.3 GB c45464fc9bee
model-00138-of-000163.safetensorsWeights4.3 GB 6a1fe7da4eaa
model-00139-of-000163.safetensorsWeights4.3 GB 3ffc04eb768f
model-00140-of-000163.safetensorsWeights4.3 GB d6ae521438d7
model-00141-of-000163.safetensorsWeights3.1 GB 2e49bea9a512
model-00142-of-000163.safetensorsWeights4.3 GB a84052a92f73
model-00143-of-000163.safetensorsWeights4.3 GB 7682d68689d9
model-00144-of-000163.safetensorsWeights4.3 GB e44a3f368062
model-00145-of-000163.safetensorsWeights4.3 GB c9c978c5afe3
model-00146-of-000163.safetensorsWeights4.3 GB 7f2a681f4931
model-00147-of-000163.safetensorsWeights4.3 GB 14cb65858eaf
model-00148-of-000163.safetensorsWeights4.3 GB a4185359e4e4
model-00149-of-000163.safetensorsWeights4.3 GB d9be3b4e6790
model-00150-of-000163.safetensorsWeights4.3 GB 90d40cf5fb98
model-00151-of-000163.safetensorsWeights4.3 GB 7222b3366b27
model-00152-of-000163.safetensorsWeights4.3 GB 0654913bb563
model-00153-of-000163.safetensorsWeights4.3 GB 4e09a5eb4e1e
model-00154-of-000163.safetensorsWeights4.3 GB ffc11b87992c
model-00155-of-000163.safetensorsWeights4.3 GB d47a88b4d737
model-00156-of-000163.safetensorsWeights4.3 GB eeadc5eb1bd9
model-00157-of-000163.safetensorsWeights4.3 GB 8205af59d738
model-00158-of-000163.safetensorsWeights4.3 GB c122494d6674
model-00159-of-000163.safetensorsWeights4.3 GB 6e01a5c5b28b
model-00160-of-000163.safetensorsWeights5.2 GB 5163fbdd76df
model-00161-of-000163.safetensorsWeights4.3 GB dc4497ab378f
model-00162-of-000163.safetensorsWeights4.3 GB 9c8582309da9
model-00163-of-000163.safetensorsWeights6.6 GB 0a97cb51f6a9
config.jsonConfiguration1.7 KB
configuration_deepseek.pyConfiguration9.9 KB
model.safetensors.index.jsonConfiguration8.9 MB
modeling_deepseek.pyConfiguration75.7 KB
LICENSEDocumentation1.1 KB
README.mdDocumentation10.7 KB
figures/0324_comparison.pngOther292.1 KB
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer7.8 MB
tokenizer_config.jsonTokenizer3.9 KB

License and Download

License
mit
Access
Open weights, no gate
Download size
688.6 GB
Download from DeepSeek

Released by DeepSeek through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 81.2 Model Card
Reported by a third party
Evaluated revision not stated 2026-01-28

Memory Requirements

PrecisionWeights in memory
As published688.6 GB
16-bit1369.1 GB
8-bit684.5 GB
4-bit342.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
DeepInfra$0.24 / $0.90input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Questions About DeepSeek-V3-0324

How much GPU memory does DeepSeek-V3-0324 need?

About 1642.9 GB at 16-bit and 410.7 GB at 4-bit: the weights (684.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-V3-0324 on?

At 16-bit, 7x MI325X from $14.00 an hour; at 4-bit, 2x MI325X from $4.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-V3-0324 commercially?

Yes. DeepSeek-V3-0324 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-V3-0324's context length?

163,840 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

DeepSeek-V3

DeepSeek

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations…

Open weights 684.5B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-R1

DeepSeek

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across…

Open weights mit 684.5B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-V3.2

DeepSeek

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs: 1. DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance, specifically optimized for long-context scenarios. 2. Scalable Reinforcement Learning Framework: By implementing a robust RL protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par…

Open weights mit 685.4B parameters 163,840 tokens transformers

Model · Text generation

GLM-5.2-FP8

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.2

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.3

Z.ai

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the Ascend NPU platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. - GLM-5.3 supports controlling the thinking budget through the reasoningeffort parameter, which accepts three levels: low, high, and max. It defaults to max…

Open weights other 753.3B parameters 1,048,576 tokens transformers