SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DeepSeek-V3

by DeepSeek deepseek-ai/DeepSeek-V3

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.

Parameters684.5B
Context163,840
Weights688.6 GB
License
AccessOpen weights
Monthly Downloads1.1M

Runs On

What it takes to serve DeepSeek-V3 (684.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1369.1 GB 1642.9 GB 7x MI325X (256 GB)
Vultr
$14.00 6x MI355X $15.54 · 7x B300 $46.20
8-bit 684.5 GB 821.4 GB 3x MI355X (288 GB)
Vultr
$7.77 4x MI325X $8.00 · 5x MI300X $9.25
4-bit 342.3 GB 410.7 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on DeepSeek-V3

1,642.9 GB. That is what DeepSeek-V3 needs at 16-bit, so the cheapest Index setup there is seven MI325X cards at $14.00 per hour. At 8-bit it drops to 821.4 GB on three MI355X at $7.77; at 4-bit, 410.7 GB on two MI325X at $4.00. What makes a 684.5B-parameter text model livable is the routing: 256 experts, 8 active per token, so compute per token is a fraction of the memory you provision. The 163,840-token context takes whole document sets in one call.

Our facts carry no license for this model, so get the license text from the publisher before committing hardware. The weights are 185 files and 688.6 GB, so plan transfer and storage ahead of the cards. And run the math against the API: DeepInfra lists $0.32 per million input tokens and $0.89 output, so a $4.00-per-hour 4-bit rig needs a known token volume to pay off.

Model Card

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations…

Excerpt from the card by DeepSeek.

Configuration

Architecture
DeepseekV3ForCausalLM
Context length (tokens)
163,840
Layers
61
Hidden size
7,168
Feed-forward size
18,432
Attention heads
128
Key/value heads
128
Vocabulary size
129,280
Routed experts
256
Experts active per token
8
RoPE base
10,000
Stored precision
bfloat16
Model type
deepseek_v3
Quantization
fp8

Identity and Version

Repository
deepseek-ai/DeepSeek-V3
Publisher
DeepSeek
Task
Text generation
Modality
Text
Library
transformers
Parameters
684.5B parameters
Languages
Not stated by the source
Revision
e815299b0bcbac849fa540c768ef21845365c9eb
First published
2024-12-25
Last updated
2025-03-27

Files and Weights

185 files, 688.6 GB in total. The weights are 163 files totalling 688.6 GB in safetensors.

Weights163 files · 688.6 GB
Configuration12 files · 9.0 MB
Tokenizer2 files · 7.9 MB
Documentation4 files · 38.1 KB
Other3 files · 292.1 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-000163.safetensorsWeights5.2 GB b933b099f335
model-00002-of-000163.safetensorsWeights4.3 GB f619dca0db00
model-00003-of-000163.safetensorsWeights4.3 GB 1f23a8b42d45
model-00004-of-000163.safetensorsWeights4.3 GB 423e7a58cbb9
model-00005-of-000163.safetensorsWeights4.3 GB 8613aed68453
model-00006-of-000163.safetensorsWeights4.4 GB 7938cb303f2c
model-00007-of-000163.safetensorsWeights4.3 GB 3e4d025d5510
model-00008-of-000163.safetensorsWeights4.3 GB a1c86b74e9cb
model-00009-of-000163.safetensorsWeights4.3 GB 0a9a0804e3fb
model-00010-of-000163.safetensorsWeights4.3 GB ba45df8c31e8
model-00011-of-000163.safetensorsWeights4.3 GB 268851646260
model-00012-of-000163.safetensorsWeights1.3 GB 71986c26dcc3
model-00013-of-000163.safetensorsWeights4.3 GB befea8fe1acb
model-00014-of-000163.safetensorsWeights4.3 GB 53a9be88a0aa
model-00015-of-000163.safetensorsWeights4.3 GB a1220514c2f0
model-00016-of-000163.safetensorsWeights4.3 GB 0838ed15d5bb
model-00017-of-000163.safetensorsWeights4.3 GB c56501d42f3e
model-00018-of-000163.safetensorsWeights4.3 GB 15ba8e1a5149
model-00019-of-000163.safetensorsWeights4.3 GB 1e21f56c613a
model-00020-of-000163.safetensorsWeights4.3 GB 56d73a3442aa
model-00021-of-000163.safetensorsWeights4.3 GB 3465da1ad357
model-00022-of-000163.safetensorsWeights4.3 GB 00a47040de71
model-00023-of-000163.safetensorsWeights4.3 GB 2b1972709df3
model-00024-of-000163.safetensorsWeights4.3 GB d86d6a40bc05
model-00025-of-000163.safetensorsWeights4.3 GB 888c83fba1bb
model-00026-of-000163.safetensorsWeights4.3 GB 568555db6a0a
model-00027-of-000163.safetensorsWeights4.3 GB 98656f5ecfe6
model-00028-of-000163.safetensorsWeights4.3 GB 1c45e4a3ba2c
model-00029-of-000163.safetensorsWeights4.3 GB 286a4e7f47c5
model-00030-of-000163.safetensorsWeights4.3 GB c11ead3c5f4a
model-00031-of-000163.safetensorsWeights4.3 GB e94d32e8649e
model-00032-of-000163.safetensorsWeights4.3 GB e7333388eedf
model-00033-of-000163.safetensorsWeights4.3 GB f6b8eded7cb9
model-00034-of-000163.safetensorsWeights1.7 GB bdf369e4ecb2
model-00035-of-000163.safetensorsWeights4.3 GB 5c7afe2627b1
model-00036-of-000163.safetensorsWeights4.3 GB 2eb3e5f66889
model-00037-of-000163.safetensorsWeights4.3 GB f5eedc7c4581
model-00038-of-000163.safetensorsWeights4.3 GB 5be6811ab78a
model-00039-of-000163.safetensorsWeights4.3 GB be4ce145bc0a
model-00040-of-000163.safetensorsWeights4.3 GB b421975c36f7
model-00041-of-000163.safetensorsWeights4.3 GB 98cb15c98687
model-00042-of-000163.safetensorsWeights4.3 GB 69d2383fe605
model-00043-of-000163.safetensorsWeights4.3 GB f5a9e6a958dc
model-00044-of-000163.safetensorsWeights4.3 GB 2d32eff0a6ce
model-00045-of-000163.safetensorsWeights4.3 GB 1b5aeae7ecce
model-00046-of-000163.safetensorsWeights4.3 GB f842c6a79d05
model-00047-of-000163.safetensorsWeights4.3 GB e1ab502eae5b
model-00048-of-000163.safetensorsWeights4.3 GB 7cabe3dc1d4b
model-00049-of-000163.safetensorsWeights4.3 GB 0af950690c49
model-00050-of-000163.safetensorsWeights4.3 GB 805ca3bc43e3
model-00051-of-000163.safetensorsWeights4.3 GB 76e249f3e1e8
model-00052-of-000163.safetensorsWeights4.3 GB ba4543f76284
model-00053-of-000163.safetensorsWeights4.3 GB c726a7c86b4f
model-00054-of-000163.safetensorsWeights4.3 GB 09bf2a5fc1cc
model-00055-of-000163.safetensorsWeights4.3 GB 12ea0f770efa
model-00056-of-000163.safetensorsWeights1.7 GB 1e094059b3f5
model-00057-of-000163.safetensorsWeights4.3 GB 86f30117b460
model-00058-of-000163.safetensorsWeights4.3 GB b131e54bc6bd
model-00059-of-000163.safetensorsWeights4.3 GB c6cf937071dd
model-00060-of-000163.safetensorsWeights4.3 GB d04f96624640
model-00061-of-000163.safetensorsWeights4.3 GB 122af5c0a0ca
model-00062-of-000163.safetensorsWeights4.3 GB 3845f9eba351
model-00063-of-000163.safetensorsWeights4.3 GB 8bc190d0f67b
model-00064-of-000163.safetensorsWeights4.3 GB f8d8d94619a1
model-00065-of-000163.safetensorsWeights4.3 GB f0bda27a01c5
model-00066-of-000163.safetensorsWeights4.3 GB 4f907d245b6d
model-00067-of-000163.safetensorsWeights4.3 GB 5107a41e1771
model-00068-of-000163.safetensorsWeights4.3 GB fa5187cd0f09
model-00069-of-000163.safetensorsWeights4.3 GB 31707005180b
model-00070-of-000163.safetensorsWeights4.3 GB a350ac1e4754
model-00071-of-000163.safetensorsWeights4.3 GB 60424a0f4c25
model-00072-of-000163.safetensorsWeights4.3 GB 6b0b6f9a1fb3
model-00073-of-000163.safetensorsWeights4.3 GB 4d386236d747
model-00074-of-000163.safetensorsWeights4.3 GB f663d4e39baa
model-00075-of-000163.safetensorsWeights4.3 GB 712e24abbb57
model-00076-of-000163.safetensorsWeights4.3 GB 5f922ddc4f32
model-00077-of-000163.safetensorsWeights4.3 GB c299d2121231
model-00078-of-000163.safetensorsWeights1.7 GB 7fa2a67820d2
model-00079-of-000163.safetensorsWeights4.3 GB 38fb787ec3ed
model-00080-of-000163.safetensorsWeights4.3 GB 07556b85b207
model-00081-of-000163.safetensorsWeights4.3 GB 92c6b2c837a4
model-00082-of-000163.safetensorsWeights4.3 GB 28d9cc27f97d
model-00083-of-000163.safetensorsWeights4.3 GB 9390263503c9
model-00084-of-000163.safetensorsWeights4.3 GB c254740ff8e4
model-00085-of-000163.safetensorsWeights4.3 GB b832c3b79a8a
model-00086-of-000163.safetensorsWeights4.3 GB 7cf824bf62e6
model-00087-of-000163.safetensorsWeights4.3 GB af791e4477bd
model-00088-of-000163.safetensorsWeights4.3 GB 6f7db096b2cc
model-00089-of-000163.safetensorsWeights4.3 GB 2129867f49a1
model-00090-of-000163.safetensorsWeights4.3 GB 7fab650d1ed4
model-00091-of-000163.safetensorsWeights4.3 GB 519a413487c5
model-00092-of-000163.safetensorsWeights4.3 GB e830585cb310
model-00093-of-000163.safetensorsWeights4.3 GB ea2b9fc654be
model-00094-of-000163.safetensorsWeights4.3 GB c44b1d39f801
model-00095-of-000163.safetensorsWeights4.3 GB fe8cf8717b94
model-00096-of-000163.safetensorsWeights4.3 GB 18af8ddf19f2
model-00097-of-000163.safetensorsWeights4.3 GB 3b6b7f63017d
model-00098-of-000163.safetensorsWeights4.3 GB 69825814a5ae
model-00099-of-000163.safetensorsWeights4.3 GB 495dd6d48845
model-00100-of-000163.safetensorsWeights1.7 GB 8e14037bf546
model-00101-of-000163.safetensorsWeights4.3 GB d2e75ba894c2
model-00102-of-000163.safetensorsWeights4.3 GB c52f51a3e20e
model-00103-of-000163.safetensorsWeights4.3 GB 86652a9d272c
model-00104-of-000163.safetensorsWeights4.3 GB ecde06a98c84
model-00105-of-000163.safetensorsWeights4.3 GB dd3836e12e35
model-00106-of-000163.safetensorsWeights4.3 GB a4c63a600528
model-00107-of-000163.safetensorsWeights4.3 GB a4e1a19f5340
model-00108-of-000163.safetensorsWeights4.3 GB 3e1efd9300e6
model-00109-of-000163.safetensorsWeights4.3 GB fca398562437
model-00110-of-000163.safetensorsWeights4.3 GB dac3c17b6f31
model-00111-of-000163.safetensorsWeights4.3 GB e8d0b963898c
model-00112-of-000163.safetensorsWeights4.3 GB 0c7030edeb4d
model-00113-of-000163.safetensorsWeights4.3 GB 2ca5c1ad7b1d
model-00114-of-000163.safetensorsWeights4.3 GB f7bc372c8e65
model-00115-of-000163.safetensorsWeights4.3 GB 336a2969eb6b
model-00116-of-000163.safetensorsWeights4.3 GB 1cb0cb2c7437
model-00117-of-000163.safetensorsWeights4.3 GB d160873806f7
model-00118-of-000163.safetensorsWeights4.3 GB 72680742383e
model-00119-of-000163.safetensorsWeights4.3 GB a7f04447b66d
model-00120-of-000163.safetensorsWeights4.3 GB a540ce2f0766
model-00121-of-000163.safetensorsWeights4.3 GB e37eecb8bf6e
model-00122-of-000163.safetensorsWeights1.7 GB 1c5691a0293c
model-00123-of-000163.safetensorsWeights4.3 GB c7f647d98cb4
model-00124-of-000163.safetensorsWeights4.3 GB f3ddfe99d4b4
model-00125-of-000163.safetensorsWeights4.3 GB 0558a282fb72
model-00126-of-000163.safetensorsWeights4.3 GB 79744a5d57eb
model-00127-of-000163.safetensorsWeights4.3 GB bc29d0be4287
model-00128-of-000163.safetensorsWeights4.3 GB fe5b46ca8a8a
model-00129-of-000163.safetensorsWeights4.3 GB 5ccf311da30b
model-00130-of-000163.safetensorsWeights4.3 GB 1ef6f7e7e0fa
model-00131-of-000163.safetensorsWeights4.3 GB 13265becdc7a
model-00132-of-000163.safetensorsWeights4.3 GB 66f55bd2bb59
model-00133-of-000163.safetensorsWeights4.3 GB 2a810f80e3a1
model-00134-of-000163.safetensorsWeights4.3 GB 710bd36d7f3a
model-00135-of-000163.safetensorsWeights4.3 GB dfa8902cbe1f
model-00136-of-000163.safetensorsWeights4.3 GB 95771ee9e65c
model-00137-of-000163.safetensorsWeights4.3 GB e5099a159fa7
model-00138-of-000163.safetensorsWeights4.3 GB ce57c5d6ca9f
model-00139-of-000163.safetensorsWeights4.3 GB e45e57d8b348
model-00140-of-000163.safetensorsWeights4.3 GB 64e8d879017d
model-00141-of-000163.safetensorsWeights3.1 GB 53a063da2194
model-00142-of-000163.safetensorsWeights4.3 GB 86161cab1bba
model-00143-of-000163.safetensorsWeights4.3 GB c154533d5dd9
model-00144-of-000163.safetensorsWeights4.3 GB 9cfe71f25918
model-00145-of-000163.safetensorsWeights4.3 GB 6e75678f19b4
model-00146-of-000163.safetensorsWeights4.3 GB a73fa1faff5d
model-00147-of-000163.safetensorsWeights4.3 GB 7b787dd38a96
model-00148-of-000163.safetensorsWeights4.3 GB 229518f00be3
model-00149-of-000163.safetensorsWeights4.3 GB e87c53c353c3
model-00150-of-000163.safetensorsWeights4.3 GB 24f203609ab4
model-00151-of-000163.safetensorsWeights4.3 GB 68909c3f465d
model-00152-of-000163.safetensorsWeights4.3 GB e1a3478e9d55
model-00153-of-000163.safetensorsWeights4.3 GB 7df9a5d1f1e3
model-00154-of-000163.safetensorsWeights4.3 GB 12a3c02867c9
model-00155-of-000163.safetensorsWeights4.3 GB db7df665d353
model-00156-of-000163.safetensorsWeights4.3 GB 88c75db2a58a
model-00157-of-000163.safetensorsWeights4.3 GB 1da4a1f20375
model-00158-of-000163.safetensorsWeights4.3 GB 9c12b470327e
model-00159-of-000163.safetensorsWeights4.3 GB 5e4be3f94f2d
model-00160-of-000163.safetensorsWeights5.2 GB d99c3c4ebe91
model-00161-of-000163.safetensorsWeights4.3 GB ccaa50981675
model-00162-of-000163.safetensorsWeights4.3 GB b233f58b3441
model-00163-of-000163.safetensorsWeights6.6 GB 1ad67ea80cad
config.jsonConfiguration1.7 KB
configuration_deepseek.pyConfiguration9.9 KB
inference/configs/config_16B.jsonConfiguration417 B
inference/configs/config_236B.jsonConfiguration455 B
inference/configs/config_671B.jsonConfiguration503 B
inference/convert.pyConfiguration3.2 KB
inference/fp8_cast_bf16.pyConfiguration3.2 KB
inference/generate.pyConfiguration5.4 KB
inference/kernel.pyConfiguration4.3 KB
inference/model.pyConfiguration17.6 KB
model.safetensors.index.jsonConfiguration8.9 MB
modeling_deepseek.pyConfiguration75.7 KB
LICENSE-CODEDocumentation1.1 KB
LICENSE-MODELDocumentation13.8 KB
README.mdDocumentation19.6 KB
README_WEIGHTS.mdDocumentation3.6 KB
figures/benchmark.pngOther183.6 KB
figures/niah.pngOther108.5 KB
inference/requirements.txtOther66 B
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer7.8 MB
tokenizer_config.jsonTokenizer3.1 KB

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
688.6 GB
Download from DeepSeek

Released by DeepSeek through its official repository on Hugging Face.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established 58.2071 EvalEval
Reported by a third party
Evaluated revision not stated 2025-03-19
LEXam-Benchmark/LEXam Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established 46.57 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
LEXam-Benchmark/LEXam Task open_questionMetric open_questionComparison conditions not established 52.53 LEXam Leaderboard
Reported by a third party
Evaluated revision not stated 2026-06-02
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 64.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-01-28
openai/gsm8k Task gsm8kMetric gsm8kComparison conditions not established 89.3 Model Card
Reported by a third party
Evaluated revision not stated 2024-12-25
thamilvendhan/signalbench Task access_denyMetric access_denySetup family=access_deny; n=12Comparison conditions not established 0.5833 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task bot_policyMetric bot_policySetup family=bot_policy; n=12Comparison conditions not established 0.25 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task injectionMetric injectionSetup family=injection; n=12Comparison conditions not established 0.5833 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task memory_labelMetric memory_labelSetup family=memory_label; n=12Comparison conditions not established 0.5833 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task srcMetric srcSetup SRC overall; deterministic action-based grader, no LLM judge; seed 0, n=75Comparison conditions not established 0.5667 signalbench raw per-item responses
Reported by a third party
Evaluated revision not stated 2026-07-08
thamilvendhan/signalbench Task timeMetric timeSetup family=time; n=12Comparison conditions not established 0.8333 thamilvendhan
Reported by a third party
Evaluated revision not stated 2026-07-08

Memory Requirements

PrecisionWeights in memory
As published688.6 GB
16-bit1369.1 GB
8-bit684.5 GB
4-bit342.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
DeepInfra$0.32 / $0.89input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Questions About DeepSeek-V3

How much GPU memory does DeepSeek-V3 need?

About 1642.9 GB at 16-bit and 410.7 GB at 4-bit: the weights (684.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-V3 on?

At 16-bit, 7x MI325X from $14.00 an hour; at 4-bit, 2x MI325X from $4.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

What is DeepSeek-V3's context length?

163,840 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

DeepSeek-V3-0324

DeepSeek

DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects. - More aesthetically pleasing web pages and game front-ends - Enhanced report analysis requests with more detailed outputs - Increased accuracy in Function Calling, fixing issues from previous V3 versions In the official DeepSeek web/app, we use the same system prompt with a specific date. For example, In our web and application environments, the temperature parameter $T{model}$ is set to 0.3. Because many users use the default temperature 1.0 in API call, we have implemented an API temperature $T{api}$ mapping mechanism that adjusts the input API temperature value of 1.0 to the…

Open weights mit 684.5B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-R1

DeepSeek

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across…

Open weights mit 684.5B parameters 163,840 tokens transformers

Model · Text generation

DeepSeek-V3.2

DeepSeek

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs: 1. DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance, specifically optimized for long-context scenarios. 2. Scalable Reinforcement Learning Framework: By implementing a robust RL protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par…

Open weights mit 685.4B parameters 163,840 tokens transformers

Model · Text generation

GLM-5.2-FP8

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.2

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.2 blog and GLM-5 Technical report. Use GLM-5.2 API services on Z.ai API Platform. Try GLM-5.2 here. [ Paper ] [ GitHub ] We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: GLM-5.2 supports deployment with the following frameworks. Feel free to try them out: - SGLang (v0.5.13.post1+) — see cookbook - vLLM (v0.23.0+) — see recipes - Transformers (v0.5.12+) — see transformers docs - KTransformers (v0.5.12+)…

Open weights mit 753.3B parameters 1,048,576 tokens transformers

Model · Text generation

GLM-5.3

Z.ai

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the Ascend NPU platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. - GLM-5.3 supports controlling the thinking budget through the reasoningeffort parameter, which accepts three levels: low, high, and max. It defaults to max…

Open weights other 753.3B parameters 1,048,576 tokens transformers