SAVRN
Search Contact SAVRN

Open-weight model · Text generation

GLM-5.2-NVFP4

by NVIDIA nvidia/GLM-5.2-NVFP4

The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture.

Parameters381B
Context1,048,576
Weights464.8 GB
Licensemit
AccessOpen weights
Monthly Downloads859.8k

Runs On

What it takes to serve GLM-5.2-NVFP4 (381B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 762.0 GB 914.4 GB 4x MI325X (256 GB)
Vultr
$8.00 5x MI300X $9.25 · 4x MI355X $10.36
8-bit 381.0 GB 457.2 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55
4-bit 190.5 GB 228.6 GB 1x MI325X (256 GB)
Vultr
$2.00 1x MI355X $2.59 · 2x MI300X $3.70

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on GLM-5.2-NVFP4

We read this build as the way to get GLM-5.2 onto one card. NVIDIA's NVFP4 quantization of ZAI's 381B-parameter mixture-of-experts model, 256 routed experts, built for reasoning and coding, needs 228.6 GB at 4-bit and fits one MI325X with 256 GB at $2.00 an hour on demand. At 16-bit the same model needs 914.4 GB and four of those cards at $8.00 an hour. Plan disk separately: the 56 safetensors files total 464.9 GB.

The MIT terms match the base model and allow commercial use, modification and redistribution with the notices kept. Before you buy hardware for it, check how much of the 1,048,576-token window you can serve on one 256 GB card with 190.5 GB of weights already loaded, confirm your inference stack handles the GlmMoeDsaForCausalLM architecture with its sparse attention, and compare against the unquantized zai-org/GLM-5.2 if 4-bit is not your target.

Model Card

By NVIDIA, published under mit, revision 53e0691e2189.

Model Overview

Description:

The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention (with an IndexShare indexer) to support a long context. For more information, please check here. The NVIDIA GLM-5.2 NVFP4 model is quantized with Model Optimizer.

This model is ready for commercial or non-commercial use.

License/Terms of Use:

GOVERNING TERMS: Use of the model is governed by the MIT License, same as the base model.

Deployment Geography:

Global

Use Case:

Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG systems, and other AI-powered applications.

Release Date:

Hugging Face 06/25/2026 via https://huggingface.co/nvidia/GLM-5.2-NVFP4

References

Nvidia Model Optimizer: https://github.com/NVIDIA/Model-Optimizer

Model Architecture:

Architecture Type: Transformers
Network Architecture: GLM-5.2 (GlmMoeDsaForCausalLM)
Number of Model Parameters: 753B in total and 40B activated

Input:

Read the full model card (1,412 words)

Configuration

Architecture
GlmMoeDsaForCausalLM
Context length (tokens)
1,048,576
Layers
78
Hidden size
6,144
Feed-forward size
12,288
Attention heads
64
Key/value heads
64
Head dimension
192
Vocabulary size
154,880
Routed experts
256
Experts
256
Experts active per token
8
Model type
glm_moe_dsa
Quantization
modelopt

Identity and Version

Repository
nvidia/GLM-5.2-NVFP4
Publisher
NVIDIA
Task
Text generation
Modality
Text
Library
Model Optimizer
Parameters
381B parameters
Languages
Not stated by the source
Revision
53e0691e21895a3863a606dfd12910c69eba94ab
First published
2026-06-22
Last updated
2026-08-31

Files and Weights

56 files, 464.9 GB in total. The weights are 47 files totalling 464.8 GB in safetensors.

Weights47 files · 464.8 GB
Configuration4 files · 22.2 MB
Tokenizer2 files · 20.2 MB
Documentation1 file · 12.4 KB
Other1 file · 5.1 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00047.safetensorsWeights10.0 GB 7355ca8065da
model-00002-of-00047.safetensorsWeights10.0 GB 5a93bf6840c9
model-00003-of-00047.safetensorsWeights10.0 GB 240485288e04
model-00004-of-00047.safetensorsWeights10.0 GB 790fe046fb8b
model-00005-of-00047.safetensorsWeights10.0 GB 9c6a3381b16f
model-00006-of-00047.safetensorsWeights10.0 GB 902e254158f0
model-00007-of-00047.safetensorsWeights10.0 GB 40299690b696
model-00008-of-00047.safetensorsWeights10.0 GB 139f1c107077
model-00009-of-00047.safetensorsWeights10.0 GB e1db5be0dac7
model-00010-of-00047.safetensorsWeights10.0 GB 617a2ca6fd37
model-00011-of-00047.safetensorsWeights10.0 GB 5099883c548f
model-00012-of-00047.safetensorsWeights10.0 GB df6a0c6ab9fa
model-00013-of-00047.safetensorsWeights10.0 GB 2bfd1392eda1
model-00014-of-00047.safetensorsWeights10.0 GB 6320d889ab91
model-00015-of-00047.safetensorsWeights10.0 GB 499d121db3e7
model-00016-of-00047.safetensorsWeights10.0 GB b0bfeb109e4a
model-00017-of-00047.safetensorsWeights10.0 GB ca88817f6f02
model-00018-of-00047.safetensorsWeights10.0 GB 4d32422c750f
model-00019-of-00047.safetensorsWeights10.0 GB 74ecfce9cddd
model-00020-of-00047.safetensorsWeights10.0 GB 23a23304160c
model-00021-of-00047.safetensorsWeights10.0 GB db14cdfbe3e5
model-00022-of-00047.safetensorsWeights10.0 GB bf5f1c293e95
model-00023-of-00047.safetensorsWeights10.0 GB 2fb70768ccd1
model-00024-of-00047.safetensorsWeights10.0 GB ccf3b2e4ca1e
model-00025-of-00047.safetensorsWeights10.0 GB 67f233905c40
model-00026-of-00047.safetensorsWeights10.0 GB daf23e3ccbf4
model-00027-of-00047.safetensorsWeights10.0 GB 1b4c7404684d
model-00028-of-00047.safetensorsWeights10.0 GB a895f3e5a815
model-00029-of-00047.safetensorsWeights10.0 GB 5601bf446f5f
model-00030-of-00047.safetensorsWeights10.0 GB 355b7b4665a8
model-00031-of-00047.safetensorsWeights10.0 GB 7bc1ca71bffd
model-00032-of-00047.safetensorsWeights10.0 GB d11bd3fe6f4c
model-00033-of-00047.safetensorsWeights10.0 GB e39c9ff0d42c
model-00034-of-00047.safetensorsWeights10.0 GB a81b529612c9
model-00035-of-00047.safetensorsWeights10.0 GB 70d19f0b6180
model-00036-of-00047.safetensorsWeights10.0 GB cab4d531d014
model-00037-of-00047.safetensorsWeights10.0 GB dc213d3550fc
model-00038-of-00047.safetensorsWeights10.0 GB 50c4a5522bd9
model-00039-of-00047.safetensorsWeights10.0 GB ca17fd221731
model-00040-of-00047.safetensorsWeights10.0 GB 103a33ce23be
model-00041-of-00047.safetensorsWeights10.0 GB 4d398e69fdd5
model-00042-of-00047.safetensorsWeights9.9 GB 573afa976e2f
model-00043-of-00047.safetensorsWeights10.0 GB a459ccf072fc
model-00044-of-00047.safetensorsWeights10.0 GB 5ef1a45b4d69
model-00045-of-00047.safetensorsWeights10.0 GB 0fa60cd2ffdd
model-00046-of-00047.safetensorsWeights10.0 GB 6edd678c7c89
model-00047-of-00047.safetensorsWeights5.1 GB ca2dfdb8f16f
config.jsonConfiguration15.5 KB
generation_config.jsonConfiguration215 B
hf_quant_config.jsonConfiguration7.4 KB
model.safetensors.index.jsonConfiguration22.2 MB 2aa8397b501d
README.mdDocumentation12.4 KB
chat_template.jinjaOther5.1 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer20.2 MB 19e773648cb4
tokenizer_config.jsonTokenizer790 B

License and Download

License
mit
Access
Open weights, no gate
Download size
464.8 GB
Download from NVIDIA

Released by NVIDIA through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published464.8 GB
16-bit762.0 GB
8-bit381.0 GB
4-bit190.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare GLM-5.2-NVFP4

Questions About GLM-5.2-NVFP4

How much GPU memory does GLM-5.2-NVFP4 need?

About 914.4 GB at 16-bit and 228.6 GB at 4-bit: the weights (381B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run GLM-5.2-NVFP4 on?

At 16-bit, 4x MI325X from $8.00 an hour; at 4-bit, 1x MI325X from $2.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use GLM-5.2-NVFP4 commercially?

Yes. GLM-5.2-NVFP4 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is GLM-5.2-NVFP4's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

DeepSeek-V4-Flash-0731

DeepSeek

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached. DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available. 1. For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort…

Open weights mit 304.2B parameters 1,048,576 tokens transformers

Model · Text generation

DeepSeek-V4-Flash

DeepSeek

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache…

Open weights mit 290.9B parameters 1,048,576 tokens transformers

Model · Text generation

Qwen3-Coder-480B-A35B-Instruct-FP8

Qwen

Today, we're announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we're excited to introduce its most powerful variant first: Qwen3-Coder-480B-A35B-Instruct. featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks, achieving results comparable to Claude Sonnet. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format.…

Open weights apache-2.0 480.2B parameters 262,144 tokens transformers

Model · Text generation

MiniMax-M2.7

MiniMax

Join Our WeChat Discord community. MiniMax Agent API CLI MiniMax Website Hugging Face GitHub ModelScope LICENSE MiniMax-M2.7 is our first model deeply participating in its own evolution. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search. For more details, see our blog post. M2.7 initiates a cycle of model self-evolution: during development, we let the model update its own memory, build dozens of complex skills for RL experiments, and improve its own learning process based on experiment results. An internal version of M2.7 autonomously optimized a programming scaffold over…

Open weights other 228.7B parameters 204,800 tokens transformers

Model · Text generation

DeepSeek-V4-Flash-DSpark

DeepSeek

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid…

Open weights mit 165.3B parameters 1,048,576 tokens transformers

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers