SAVRN
Search Contact SAVRN

Open-weight model · Text generation

GLM-4.7-Flash

by Z.ai zai-org/GLM-4.7-Flash

Join our Discord community. Check out the GLM-4.7 technical blog, technical report(GLM-4.5). Use GLM-4.7-Flash API services on Z.ai API Platform. One click to GLM-4.7. GLM-4.7-Flash is a 30B-A3B MoE model.

Parameters31.2B
Context202,752
Weights62.4 GB
Licensemit
AccessOpen weights
Monthly Downloads1.9M

Runs On

What it takes to serve GLM-4.7-Flash (31.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 62.4 GB 74.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 31.2 GB 37.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 15.6 GB 18.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Z.ai, published under mit, revision 7dd20894a642.

Join ourDiscord community.
Check out the GLM-4.7technical blog, technical report(GLM-4.5).
Use GLM-4.7-Flash API services onZ.ai API Platform.
One click toGLM-4.7.

Introduction

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

Performances on Benchmarks

Benchmark GLM-4.7-Flash Qwen3-30B-A3B-Thinking-2507 GPT-OSS-20B
AIME 25 91.6 85.0 91.7
GPQA 75.2 73.4 71.5
LCB v6 64.0 66.0 61.0
HLE 14.4 9.8 10.9
SWE-bench Verified 59.2 22.0 34.0
τ²-Bench 79.5 49.0 47.7
BrowseComp 42.8 2.29 28.3

Evaluation Parameters

Default Settings (Most Tasks)

  • temperature: 1.0
  • top-p: 0.95
  • max new tokens: 131072

For multi-turn agentic tasks (τ²-Bench and Terminal Bench 2), please turn on Preserved Thinking mode.

Terminal Bench, SWE Bench Verified

  • temperature: 0.7
  • top-p: 1.0
  • max new tokens: 16384

τ^2-Bench

  • Temperature: 0
  • Max new tokens: 16384

Read the full model card (945 words)

Configuration

Architecture
Glm4MoeLiteForCausalLM
Context length (tokens)
202,752
Layers
47
Hidden size
2,048
Feed-forward size
10,240
Attention heads
20
Key/value heads
20
Vocabulary size
154,880
Routed experts
64
Experts active per token
4
RoPE base
1,000,000
Model type
glm4_moe_lite

Identity and Version

Repository
zai-org/GLM-4.7-Flash
Publisher
Z.ai
Task
Text generation
Modality
Text
Library
transformers
Parameters
31.2B parameters
Languages
en, zh
Revision
7dd20894a642a0aa287e9827cb1a1f7f91386b67
First published
2026-01-19
Last updated
2026-01-29

Files and Weights

58 files, 62.5 GB in total. The weights are 48 files totalling 62.4 GB in safetensors.

Weights48 files · 62.4 GB
Configuration5 files · 873.5 KB
Tokenizer2 files · 20.2 MB
Documentation1 file · 8.2 KB
Other1 file · 3.1 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00048.safetensorsWeights1.4 GB 90abe0d07575
model-00002-of-00048.safetensorsWeights1.3 GB 8c51e2434efe
model-00003-of-00048.safetensorsWeights1.3 GB ab6ebdd01af1
model-00004-of-00048.safetensorsWeights1.3 GB 932e4dbb8f75
model-00005-of-00048.safetensorsWeights1.3 GB 7f56a3e52bf5
model-00006-of-00048.safetensorsWeights1.3 GB 83aa0eeca6e5
model-00007-of-00048.safetensorsWeights1.3 GB a9679076d3bd
model-00008-of-00048.safetensorsWeights1.3 GB 7d1231577a05
model-00009-of-00048.safetensorsWeights1.3 GB 095d62662677
model-00010-of-00048.safetensorsWeights1.3 GB c3f8d197b0d2
model-00011-of-00048.safetensorsWeights1.3 GB a8b7eeef625a
model-00012-of-00048.safetensorsWeights1.3 GB 49adb8b9f02d
model-00013-of-00048.safetensorsWeights1.3 GB f269382ae6ce
model-00014-of-00048.safetensorsWeights1.3 GB acd6f36aadce
model-00015-of-00048.safetensorsWeights1.3 GB 606079fe1e00
model-00016-of-00048.safetensorsWeights1.3 GB b53fb45f7689
model-00017-of-00048.safetensorsWeights1.3 GB a51e478badd5
model-00018-of-00048.safetensorsWeights1.3 GB 680295c83f29
model-00019-of-00048.safetensorsWeights1.3 GB 355f9437e291
model-00020-of-00048.safetensorsWeights1.3 GB 8f8d4698fc5c
model-00021-of-00048.safetensorsWeights1.3 GB 111570fc7f55
model-00022-of-00048.safetensorsWeights1.3 GB ced6a73df393
model-00023-of-00048.safetensorsWeights1.3 GB b287bb0ce12a
model-00024-of-00048.safetensorsWeights1.3 GB c6e77476f0f5
model-00025-of-00048.safetensorsWeights1.3 GB 8f9e643d9d3e
model-00026-of-00048.safetensorsWeights1.3 GB 634da561277b
model-00027-of-00048.safetensorsWeights1.3 GB 3286537b5a5d
model-00028-of-00048.safetensorsWeights1.3 GB 56335c906d44
model-00029-of-00048.safetensorsWeights1.3 GB 3e6591dfb36a
model-00030-of-00048.safetensorsWeights1.3 GB 62373b800980
model-00031-of-00048.safetensorsWeights1.3 GB 332d82cefb0b
model-00032-of-00048.safetensorsWeights1.3 GB 0f9897915eef
model-00033-of-00048.safetensorsWeights1.3 GB 29059fb79e6f
model-00034-of-00048.safetensorsWeights1.3 GB f989cc8af199
model-00035-of-00048.safetensorsWeights1.3 GB aaa91e9b1a07
model-00036-of-00048.safetensorsWeights1.3 GB cf807da31533
model-00037-of-00048.safetensorsWeights1.3 GB 887aeb5568ef
model-00038-of-00048.safetensorsWeights1.3 GB 036a38cc1085
model-00039-of-00048.safetensorsWeights1.3 GB 62ae341291fe
model-00040-of-00048.safetensorsWeights1.3 GB cd8af33462a3
model-00041-of-00048.safetensorsWeights1.3 GB 39098c9b6770
model-00042-of-00048.safetensorsWeights1.3 GB ffe0b23aab19
model-00043-of-00048.safetensorsWeights1.3 GB efe1da68f8c3
model-00044-of-00048.safetensorsWeights1.3 GB a731baa13708
model-00045-of-00048.safetensorsWeights1.3 GB a636f872a931
model-00046-of-00048.safetensorsWeights1.3 GB ba5365e31c27
model-00047-of-00048.safetensorsWeights2.5 GB 1bcc5d06065d
model-00048-of-00048.safetensorsWeights1.3 GB 35fff90a30ca
.eval_results/gpqa.yamlConfiguration197 B
.eval_results/hle.yamlConfiguration186 B
config.jsonConfiguration1.1 KB
generation_config.jsonConfiguration181 B
model.safetensors.index.jsonConfiguration871.8 KB
README.mdDocumentation8.2 KB
chat_template.jinjaOther3.1 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer20.2 MB 19e773648cb4
tokenizer_config.jsonTokenizer7.2 KB

License and Download

License
mit
Access
Open weights, no gate
Download size
62.4 GB
Download from Z.ai

Released by Z.ai through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2508.06471

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 75.2 Model Card
Reported by a third party
Evaluated revision not stated 2026-01-27
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established 59.2 Model Card
Reported by a third party
Evaluated revision not stated 2026-03-18
cais/hle Task hleMetric hleComparison conditions not established 14.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-01-28

Memory Requirements

PrecisionWeights in memory
As published62.4 GB
16-bit62.4 GB
8-bit31.2 GB
4-bit15.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
DeepInfra$0.06 / $0.40input / output, per million tokensSep 18, 2026
Novita$0.07 / $0.40input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Compare GLM-4.7-Flash

Questions About GLM-4.7-Flash

How much GPU memory does GLM-4.7-Flash need?

About 74.9 GB at 16-bit and 18.7 GB at 4-bit: the weights (31.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run GLM-4.7-Flash on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use GLM-4.7-Flash commercially?

Yes. GLM-4.7-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is GLM-4.7-Flash's context length?

202,752 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

OTel-2.0-LLM-31B-IT

Farbod Tavakkoli

OTel-2.0-LLM-31B-IT is a telecom-specialized instruction model post-trained from Gemma 4 31B-IT on approximately 440 billion telecom training tokens. It is the first release in the OTel 2.0 family and is designed to support telco-grade AI workflows across network operations, standards interpretation, product development, network configuration assistance, RAG, and telecom-specific question answering. OTel 2.0 extends the original OTel effort from a RAG-oriented telecom fine-tuning release into a larger domain-adapted training program. The model was trained from a much larger standards and telecom corpus, with new data preparation coverage for direct telecom QnA, abstention, RAG…

Open weights apache-2.0 31.3B parameters 262,144 tokens transformers

Model · Text generation

Qwen3-VL-30B-A3B-Instruct-AWQ

QuantTrio

As of 2025-10-08, create a fresh Python environment and run: For more details, refer to vLLM Official Qwen3-VL Guide Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision…

Open weights apache-2.0 31.1B parameters 262,144 tokens transformers

Model · Text generation

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

NVIDIA

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\. Nemotron-3-Nano-30B-A3B-BF16 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in…

Open weights other 31.6B parameters 262,144 tokens transformers

Fastino-Nemotron-3.5-Lightning-Finance is a 30B-parameter, 3B-active mixture-of-experts model specialized for financial reasoning, extraction, and research fine-tuned on LoRA with the Fastino Fine-Tuning Agent. The model targets financial document reasoning, numerical question answering over filings and tables, numeric span extraction, financial entity recognition, conversational analysis, and source-grounded financial research. The evaluation suite includes FinQA, TAT-QA, SEC-Num, FinEntity, BizFinBench, BigFinanceBench, ConvFinQA, and FiQA. The published weights are BF16 and require about 66 GB before runtime overhead. An 80 GB or larger GPU, or tensor parallelism across multiple GPUs, is…

Open weights apache-2.0 31.6B parameters 262,144 tokens transformers

Model · Text generation

Qwen3-Coder-30B-A3B-Instruct-FP8

Qwen

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct-FP8. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format. Qwen3-Coder-30B-A3B-Instruct-FP8 has the following features: NOTE: This model…

Open weights apache-2.0 30.5B parameters 262,144 tokens transformers