SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound

by Aldo Zampatti azampatti/Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound

Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound is an open-weight model for text generation from Aldo Zampatti, released under other. It has 124B parameters and a 262,144-token context. At 16-bit it needs about 297.5 GB of GPU memory, which fits on 2x MI300X from $3.70 an hour; at 4-bit, 74.4 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index. It draws 612 downloads a month.

Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token.

Parameters124B
Context262,144
Weights127.4 GB
Licenseother
AccessOpen weights
Monthly Downloads612

Runs On

What it takes to serve Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound (124B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 247.9 GB 297.5 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18
8-bit 124.0 GB 148.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 62.0 GB 74.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 19, 2026.

Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound on every accelerator the SAVRN Index prices, at every precision

Model Card

Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token. The original activates 6B. Everything needed to serve it is in this one repository, including the 49 GB FP8 n-gram table under ple-table/. Nothing else to download. That builds the serving image, downloads this repository, and starts an OpenAI-compatible server on port 8000. The scripts and the full explanation are in that repo. Serving by hand needs Saren-Arterius/qwen3.8-Flash-DGX-AutoRound, because a stock vLLM cannot serve this checkpoint's int4 +…

Excerpt from the card by Aldo Zampatti, licensed other.

Configuration

Architecture
Qwen4ExpForConditionalGeneration
Context length (tokens)
262,144
Layers
48
Hidden size
2,560
Attention heads
24
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
512
Experts active per token
5
Model type
qwen4_exp
Quantization
gptq

Identity and Version

Repository
azampatti/Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound
Publisher
Aldo Zampatti
Task
Text generation
Modality
Text
Library
vllm
Parameters
124B parameters
Languages
moe, dgx-spark
Revision
c74982dedd31f6c7d59fed9bd699a99ed36ec407
First published
2026-09-08
Last updated
2026-09-19

Files and Weights

120 files, 127.5 GB in total. The weights are 103 files totalling 127.4 GB in safetensors.

Weights103 files · 127.4 GB
Configuration6 files · 23.9 MB
Tokenizer2 files · 20.0 MB
Documentation3 files · 8.7 KB
Other5 files · 30.4 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00067.safetensorsWeights1.1 GB aa5a710b4282
model-00002-of-00067.safetensorsWeights1.1 GB 8f5b87b772c8
model-00003-of-00067.safetensorsWeights1.1 GB 655ee58a6c9c
model-00004-of-00067.safetensorsWeights1.1 GB 0aa27d8f3054
model-00005-of-00067.safetensorsWeights1.1 GB 735e671f2d1c
model-00006-of-00067.safetensorsWeights1.1 GB 31f925cfb499
model-00007-of-00067.safetensorsWeights1.1 GB b08c6630345c
model-00008-of-00067.safetensorsWeights1.1 GB c41abba70a84
model-00009-of-00067.safetensorsWeights1.1 GB fb6ac3df2d9a
model-00010-of-00067.safetensorsWeights1.1 GB e9d959f8c409
model-00011-of-00067.safetensorsWeights1.1 GB 39791a560489
model-00012-of-00067.safetensorsWeights1.1 GB 5e5abe1a2556
model-00013-of-00067.safetensorsWeights1.1 GB 1927a3b45f3a
model-00014-of-00067.safetensorsWeights1.1 GB fcf3034a11f8
model-00015-of-00067.safetensorsWeights1.1 GB faafa98c7db5
model-00016-of-00067.safetensorsWeights1.1 GB d07c2f8c3e45
model-00017-of-00067.safetensorsWeights1.1 GB fb0f9fb50528
model-00018-of-00067.safetensorsWeights1.1 GB eb0bcd685923
model-00019-of-00067.safetensorsWeights1.1 GB bb42f78d49d0
model-00020-of-00067.safetensorsWeights1.1 GB a27533e4577c
model-00021-of-00067.safetensorsWeights1.1 GB 280b30b2e95f
model-00022-of-00067.safetensorsWeights1.1 GB bcf9a73f653a
model-00023-of-00067.safetensorsWeights1.1 GB 9277de6c2ed1
model-00024-of-00067.safetensorsWeights1.1 GB f501028cb869
model-00025-of-00067.safetensorsWeights1.1 GB 72dd947e51ee
model-00026-of-00067.safetensorsWeights1.1 GB 08ac83c53e91
model-00027-of-00067.safetensorsWeights1.1 GB 316f4189fa56
model-00028-of-00067.safetensorsWeights1.1 GB 0f5af310ca4c
model-00029-of-00067.safetensorsWeights1.1 GB 166c23a040bb
model-00030-of-00067.safetensorsWeights1.1 GB ffca539bda0b
model-00031-of-00067.safetensorsWeights1.1 GB 19f3cb3e1eeb
model-00032-of-00067.safetensorsWeights1.1 GB 862b621e0e2a
model-00033-of-00067.safetensorsWeights1.1 GB e84c144585e8
model-00034-of-00067.safetensorsWeights1.1 GB 9ab6fcce95bf
model-00035-of-00067.safetensorsWeights1.1 GB 013a2360c35e
model-00036-of-00067.safetensorsWeights1.1 GB 0f0191281be6
model-00037-of-00067.safetensorsWeights1.1 GB 7dd7061740ea
model-00038-of-00067.safetensorsWeights1.1 GB d14f46d30f8c
model-00039-of-00067.safetensorsWeights1.1 GB 54951f941578
model-00040-of-00067.safetensorsWeights1.1 GB e217cefc7a76
model-00041-of-00067.safetensorsWeights1.1 GB 08537997956e
model-00042-of-00067.safetensorsWeights1.1 GB 763ab4159175
model-00043-of-00067.safetensorsWeights1.1 GB b788e4a790e0
model-00044-of-00067.safetensorsWeights1.1 GB 2b0ad99d28dc
model-00045-of-00067.safetensorsWeights1.1 GB f5c9cf0e3e5d
model-00046-of-00067.safetensorsWeights1.1 GB 314547259fec
model-00047-of-00067.safetensorsWeights1.1 GB 06a5d39d8971
model-00048-of-00067.safetensorsWeights1.1 GB cf6665dc3163
model-00049-of-00067.safetensorsWeights1.1 GB 53bc345ed221
model-00050-of-00067.safetensorsWeights1.1 GB 1c92c979c2cd
model-00051-of-00067.safetensorsWeights1.1 GB 5c19fb787b13
model-00052-of-00067.safetensorsWeights1.1 GB 447805cd9693
model-00053-of-00067.safetensorsWeights1.1 GB f2ae5ba53061
model-00054-of-00067.safetensorsWeights1.1 GB e5243966232c
model-00055-of-00067.safetensorsWeights1.1 GB d3d044f86831
model-00056-of-00067.safetensorsWeights1.1 GB 2e78806c4618
model-00057-of-00067.safetensorsWeights1.1 GB ac40956caea3
model-00058-of-00067.safetensorsWeights1.1 GB cf9b067bb24e
model-00059-of-00067.safetensorsWeights1.1 GB 3d5d60a10532
model-00060-of-00067.safetensorsWeights1.1 GB 7e5275109d74
model-00061-of-00067.safetensorsWeights31.6 MB 746489b656b6
model-00062-of-00067.safetensorsWeights1.3 GB c8af64f29a21
model-00063-of-00067.safetensorsWeights1.1 GB 9a8ab4ebcedf
model-00064-of-00067.safetensorsWeights1.1 GB e2d8e95c0e5b
model-00065-of-00067.safetensorsWeights1.1 GB 27fb5f70f880
model-00066-of-00067.safetensorsWeights460.5 MB 0fc6f0ead0b3
model-00067-of-00067.safetensorsWeights13.1 MB bb554ccb492b
model-healed-shared-expert.safetensorsWeights472.3 MB 8efd0a570b42
model-lmhead-int4.safetensorsWeights330.3 MB 1f7fc618a7ef
model_extra_tensors.safetensorsWeights5.2 GB f9b93d0f967e
ple-table/model-00005-of-00131.safetensorsWeights1.7 GB c50bf465a4a0
ple-table/model-00006-of-00131.safetensorsWeights1.6 GB 57f103367fe3
ple-table/model-00007-of-00131.safetensorsWeights1.6 GB cde2a1285477
ple-table/model-00008-of-00131.safetensorsWeights1.6 GB 98ae323b5a61
ple-table/model-00009-of-00131.safetensorsWeights1.6 GB 70f4b66bca5c
ple-table/model-00010-of-00131.safetensorsWeights1.6 GB 6d5dfa1dfa9a
ple-table/model-00011-of-00131.safetensorsWeights1.6 GB de085adb563e
ple-table/model-00012-of-00131.safetensorsWeights1.6 GB 97dac7c366ba
ple-table/model-00013-of-00131.safetensorsWeights1.6 GB 467be10b473d
ple-table/model-00014-of-00131.safetensorsWeights1.6 GB 38def6e9bbe6
ple-table/model-00015-of-00131.safetensorsWeights1.6 GB 90c6c665f1f8
ple-table/model-00016-of-00131.safetensorsWeights1.6 GB 88185c253064
ple-table/model-00017-of-00131.safetensorsWeights1.6 GB 42e3b9635911
ple-table/model-00018-of-00131.safetensorsWeights1.6 GB 040adc7270ef
ple-table/model-00019-of-00131.safetensorsWeights1.6 GB b8a831992deb
ple-table/model-00020-of-00131.safetensorsWeights1.6 GB 9426112fbd49
ple-table/model-00021-of-00131.safetensorsWeights1.6 GB fce8426665f7
ple-table/model-00022-of-00131.safetensorsWeights1.6 GB a20dd415f280
ple-table/model-00023-of-00131.safetensorsWeights1.6 GB 410059183947
ple-table/model-00024-of-00131.safetensorsWeights1.6 GB bbd6d7a4b54b
ple-table/model-00025-of-00131.safetensorsWeights1.6 GB f9fa7b04de15
ple-table/model-00026-of-00131.safetensorsWeights1.6 GB 193c6b9da2c6
ple-table/model-00027-of-00131.safetensorsWeights1.6 GB d7efa372552c
ple-table/model-00028-of-00131.safetensorsWeights1.6 GB 5fc13d084d62
ple-table/model-00029-of-00131.safetensorsWeights1.6 GB a11f5a6360b3
ple-table/model-00030-of-00131.safetensorsWeights1.6 GB cac780a675f8
ple-table/model-00031-of-00131.safetensorsWeights1.6 GB d390a872eb39
ple-table/model-00032-of-00131.safetensorsWeights1.6 GB 3b36fe4e7215
ple-table/model-00033-of-00131.safetensorsWeights1.6 GB 3fb0504f406c
ple-table/model-00034-of-00131.safetensorsWeights1.6 GB 6d989a812a8d
ple-table/model-00035-of-00131.safetensorsWeights1.6 GB 8e56ba0a7141
ple-table/model-00036-of-00131.safetensorsWeights1.6 GB c23fd6f6cfa9
ple-table/model-00037-of-00131.safetensorsWeights942.3 MB b824a11bf460
config.jsonConfiguration4.6 KB
generation_config.jsonConfiguration214 B
model.safetensors.index.jsonConfiguration23.8 MB b9caad9aaf03
preprocessor_config.jsonConfiguration443 B
processor_config.jsonConfiguration1.2 KB
quantization_config.jsonConfiguration94.0 KB
LICENSEDocumentation3.2 KB
README.mdDocumentation4.6 KB
ple-table/README_upstream_table.mdDocumentation860 B
chat_template.jinjaOther9.0 KB
medium_chat_template.jinjaOther10.7 KB
model-healed-shared-expert.safetensors.sha256Other65 B
model-lmhead-int4.safetensors.sha256Other65 B
xhigh_chat_template.jinjaOther10.7 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.2 KB

License and Download

License
other
Access
Open weights, no gate
Download size
127.4 GB
Download from Aldo Zampatti

Released by Aldo Zampatti through its official repository on Hugging Face. Read the license.

Built From

  • Derived from Intel/Qwen3.8-Flash-Next-W4A16-AutoRound
  • Quantized from Intel/Qwen3.8-Flash-Next-W4A16-AutoRound

Memory Requirements

PrecisionWeights in memory
As published127.4 GB
16-bit247.9 GB
8-bit124.0 GB
4-bit62.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound

How much GPU memory does Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound need?

About 297.5 GB at 16-bit and 74.4 GB at 4-bit: the weights (124B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound on?

At 16-bit, 2x MI300X from $3.70 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers

Model · Text generation

gpt-oss-120b

OpenAI

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 116.8B parameters 131,072 tokens transformers

Model · Text generation

DeepSeek-V4-Flash-DSpark

DeepSeek

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid…

Open weights mit 165.3B parameters 1,048,576 tokens transformers

Model · Text generation

Qwen3-Coder-Next-FP8

Qwen

Today, we're announcing Qwen3-Coder-Next-FP8, an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements: Qwen3-Coder-Next-FP8 has the following features: NOTE: This model supports only non-thinking mode and does not generate blocks in its output. Meanwhile, specifying enablethinking=False is no longer required. For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. We advise you to use the latest version of transformers. The following contains a code snippet illustrating how to use the model generate content based on…

Open weights apache-2.0 79.7B parameters 262,144 tokens transformers

Model · Text generation

Qwen-72B

Qwen

通义千问-72B(Qwen-72B)是阿里云研发的通义千问大模型系列的720亿参数规模的模型。Qwen-72B是基于Transformer的大语言模型, 在超大规模的预训练数据上进行训练得到。预训练数据类型多样,覆盖广泛,包括大量网络文本、专业书籍、代码等。同时,在Qwen-72B的基础上,我们使用对齐机制打造了基于大语言模型的AI助手Qwen-72B-Chat。本仓库为Qwen-72B的仓库。 通义千问-72B(Qwen-72B)主要有以下特点: 1. 大规模高质量训练语料:使用超过3万亿tokens的数据进行预训练,包含高质量中、英、多语言、代码、数学等数据,涵盖通用及专业领域的训练语料。通过大量对比实验对预训练语料分布进行了优化。 2. 强大的性能:Qwen-72B在多个中英文下游评测任务上(涵盖常识推理、代码、数学、翻译等),效果显著超越现有的开源模型。具体评测结果请详见下文。 3. 覆盖更全面的词表:相比目前以中英词表为主的开源模型,Qwen-72B使用了约15万大小的词表。该词表对多语言更加友好,方便用户在不扩展词表的情况下对部分语种进行能力增强和扩展。 4. 较长的上下文支持:Qwen-72B支持32k的上下文长度。 Qwen-72B is the 72B-parameter version of the large language model series, Qwen (abbr. Tongyi Qianwen), proposed by Alibaba Cloud. Qwen-72B is a Transformer-based large…

Open weights other 72.3B parameters 32,768 tokens transformers

Model · Text generation

Llama-3.3-70B-Instruct

Meta Llama

The Meta Llama 3.3 multilingual large language model (LLM) is an instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.3 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and…

Access requested at publisher llama3.3 70.6B parameters transformers