SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Qwen3.8-Flash-Next-P48NVFP4-MoESQ

by IST Austria Distributed Algorithms and Systems Lab ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ

Qwen3.8-Flash-Next-P48NVFP4-MoESQ is an open-weight model for text generation from IST Austria Distributed Algorithms and Systems Lab, released under other. It has 134.7B parameters and a 262,144-token context. At 16-bit it needs about 323.3 GB of GPU memory, which fits on 2x MI300X from $3.70 an hour; at 4-bit, 80.8 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index. It draws 114 downloads a month.

A W4A4 + paired-4:8 sparse compressed checkpoint of Qwen/Qwen3.8-Flash-Next, produced with MoESQ. The routed MoE expert weights are NVFP4 with paired-4:8 structured sparsity and are stored sparse: only the kept values plus a small mask are on disk.

Parameters134.7B
Context262,144
Weights160.0 GB
Licenseother
AccessOpen weights
Monthly Downloads114

Runs On

What it takes to serve Qwen3.8-Flash-Next-P48NVFP4-MoESQ (134.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 269.4 GB 323.3 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18
8-bit 134.7 GB 161.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 67.4 GB 80.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Qwen3.8-Flash-Next-P48NVFP4-MoESQ on every accelerator the SAVRN Index prices, at every precision

Model Card

A W4A4 + paired-4:8 sparse compressed checkpoint of Qwen/Qwen3.8-Flash-Next, produced with MoESQ. The routed MoE expert weights are NVFP4 with paired-4:8 structured sparsity and are stored sparse: only the kept values plus a small mask are on disk. The target is NVIDIA Blackwell (SM100) sparse tensor cores. expert, 48 layers: 36 linear-attention and 12 sparse-attention, hyper-connections, per-layer n-gram embedding) weight including the sparsity mask and scales (2 bits per weight for the kept values). TP2). The per-layer n-gram embedding tables (95.4 GiB, BF16, unchanged) are lookup tables served from CPU memory, as in the base model. paired48nvfp4 MoE backend Links This checkpoint does not…

Excerpt from the card by IST Austria Distributed Algorithms and Systems Lab, licensed other.

Configuration

Architecture
Qwen4ExpForConditionalGeneration
Context length (tokens)
262,144
Layers
48
Hidden size
2,560
Attention heads
24
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
512
Experts active per token
10
Model type
qwen4_exp
Quantization
compressed-tensors

Identity and Version

Repository
ISTA-DASLab/Qwen3.8-Flash-Next-P48NVFP4-MoESQ
Publisher
IST Austria Distributed Algorithms and Systems Lab
Task
Text generation
Modality
Text
Library
transformers
Parameters
134.7B parameters
Languages
en
Revision
b25abf193ab0428d70be5888bdc75573c6a70d30
First published
2026-10-02
Last updated
2026-10-04

Files and Weights

62 files, 160.1 GB in total. The weights are 49 files totalling 160.0 GB in safetensors.

Weights49 files · 160.0 GB
Configuration5 files · 43.4 MB
Tokenizer4 files · 22.9 MB
Documentation1 file · 10.3 KB
Other2 files · 13.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00049.safetensorsWeights3.5 GB e755141e3c2c
model-00002-of-00049.safetensorsWeights6.2 GB 28a6a1be17a1
model-00003-of-00049.safetensorsWeights103.5 GB f61623b59675
model-00004-of-00049.safetensorsWeights1.0 GB 40e96351d9a8
model-00005-of-00049.safetensorsWeights1.0 GB 2dd1c82ed884
model-00006-of-00049.safetensorsWeights1.0 GB be667bcfde93
model-00007-of-00049.safetensorsWeights1.0 GB 9332c1281edf
model-00008-of-00049.safetensorsWeights1.0 GB 56e34d62c56e
model-00009-of-00049.safetensorsWeights1.0 GB 3920d1d1ca9e
model-00010-of-00049.safetensorsWeights1.0 GB 07b0ccd7facf
model-00011-of-00049.safetensorsWeights1.0 GB 03792ae08c81
model-00012-of-00049.safetensorsWeights1.0 GB a7b047c5b241
model-00013-of-00049.safetensorsWeights1.0 GB f86b2e2f05bb
model-00014-of-00049.safetensorsWeights1.0 GB 25e732ccaa55
model-00015-of-00049.safetensorsWeights1.0 GB 3eb0f69a151d
model-00016-of-00049.safetensorsWeights1.0 GB 86b4a5c0af02
model-00017-of-00049.safetensorsWeights1.0 GB 253c6ea237fb
model-00018-of-00049.safetensorsWeights1.0 GB 0735f8d5210c
model-00019-of-00049.safetensorsWeights1.0 GB 6757585e8f73
model-00020-of-00049.safetensorsWeights1.0 GB 68642a7844e2
model-00021-of-00049.safetensorsWeights1.0 GB 9c6e825b8cd1
model-00022-of-00049.safetensorsWeights1.0 GB d2a9d4dc5276
model-00023-of-00049.safetensorsWeights1.0 GB 499fab4b99d7
model-00024-of-00049.safetensorsWeights1.0 GB 9baa303daee7
model-00025-of-00049.safetensorsWeights1.0 GB 240a46b47d73
model-00026-of-00049.safetensorsWeights1.0 GB 9efdda2bd3d4
model-00027-of-00049.safetensorsWeights1.0 GB 4116e3a5061f
model-00028-of-00049.safetensorsWeights1.0 GB 9df5298a39e6
model-00029-of-00049.safetensorsWeights1.0 GB 71bd712ca1dc
model-00030-of-00049.safetensorsWeights1.0 GB 6ca3198ac554
model-00031-of-00049.safetensorsWeights1.0 GB 032cceb16c39
model-00032-of-00049.safetensorsWeights1.0 GB 92cd3153fac3
model-00033-of-00049.safetensorsWeights1.0 GB dbc112205575
model-00034-of-00049.safetensorsWeights1.0 GB 3a649099a0be
model-00035-of-00049.safetensorsWeights1.0 GB b39a992be4d4
model-00036-of-00049.safetensorsWeights1.0 GB f3cad2627402
model-00037-of-00049.safetensorsWeights1.0 GB 31dc4465361a
model-00038-of-00049.safetensorsWeights1.0 GB 1661258c60af
model-00039-of-00049.safetensorsWeights1.0 GB 77f8638e5188
model-00040-of-00049.safetensorsWeights1.0 GB 77f66362e8e8
model-00041-of-00049.safetensorsWeights1.0 GB f980377d661d
model-00042-of-00049.safetensorsWeights1.0 GB 4544852bba17
model-00043-of-00049.safetensorsWeights1.0 GB de2f4065573b
model-00044-of-00049.safetensorsWeights1.0 GB 55b81f64819b
model-00045-of-00049.safetensorsWeights1.0 GB f4d50b3b932c
model-00046-of-00049.safetensorsWeights1.0 GB 3640f6005b8c
model-00047-of-00049.safetensorsWeights1.0 GB a44b00cebe36
model-00048-of-00049.safetensorsWeights1.0 GB 6ceab4248dad
model-00049-of-00049.safetensorsWeights1.0 GB cea0bd4a9540
config.jsonConfiguration5.6 KB —
generation_config.jsonConfiguration202 B —
model.safetensors.index.jsonConfiguration43.4 MB bcb56056424b
moe_sq_config.yamlConfiguration1.5 KB —
preprocessor_config.jsonConfiguration390 B —
README.mdDocumentation10.3 KB —
SHARD_HASHES.sha256Other4.9 KB —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 0997f410c57a
tokenizer_config.jsonTokenizer17.9 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
other
Access
Open weights, no gate
Download size
160.0 GB
Download from IST Austria Distributed Algorithms and Systems Lab

Released by IST Austria Distributed Algorithms and Systems Lab through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published160.0 GB
16-bit269.4 GB
8-bit134.7 GB
4-bit67.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.8-Flash-Next-P48NVFP4-MoESQ

How much GPU memory does Qwen3.8-Flash-Next-P48NVFP4-MoESQ need?

About 323.3 GB at 16-bit and 80.8 GB at 4-bit: the weights (134.7B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.8-Flash-Next-P48NVFP4-MoESQ on?

At 16-bit, 2x MI300X from $3.70 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Qwen3.8-Flash-Next-P48NVFP4-MoESQ released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Qwen3.8-Flash-Next-P48NVFP4-MoESQ's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

keys-GLM-5.3-EXL3-2.75BPW

Keys

Full GLM-5.3 (753B, glmmoedsa) quantized to EXL3 with a per-expert mixed bit width averaging 2.75 bpw on the routed experts, sized to serve on four NVIDIA DGX Sparks (GB10, 128 GB unified memory each) with room for a 200K-token KV pool (or 1M tokens with decode-context-parallel 4). Unmodified weights of zai-org/GLM-5.3 (no abliteration, no fine-tuning), quantized from the BF16 checkpoint. For the earlier uniform-width quant see The fastest way to serve this checkpoint is TensorFold's native glmmoedsa engine — tensor parallel 4, one rank per Spark, no vLLM in the serving path; a drafted reply equals a serial one. Recipe, one-shot launcher, benches and an AGENTS.md needs no extra draft model…

Open weights other 138.1B parameters 1,048,576 tokens exllamav3

Model · Text generation

Vinci-Cyber-123B-1.0

SimpleDirect

Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B. The released merged weights have now been tested directly, alongside their parent and the available GGUF formats. Focused repairs. Restraint on correct configuration. Weights you can run yourself. On the V2-B neutral-review test, the released BF16 model preserved 24/24 correct configurations and produced 18/24 scanner-credited repairs, all 18 passing offline provider-schema validation. Its parent repaired 17/24 and preserved 0/24. On the second set, V2-A, Cyber again preserved 24/24, but repaired 9/24 versus the parent's 15/24.…

Open weights other 125B parameters 262,144 tokens transformers

Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token. The original activates 6B. Everything needed to serve it is in this one repository, including the 49 GB FP8 n-gram table under ple-table/. Nothing else to download. That builds the serving image, downloads this repository, and starts an OpenAI-compatible server on port 8000. The scripts and the full explanation are in that repo. Serving by hand needs Saren-Arterius/qwen3.8-Flash-DGX-AutoRound, because a stock vLLM cannot serve this checkpoint's int4 +…

Open weights other 124B parameters 262,144 tokens vllm

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers

Model · Text generation

Qwen3.5-397B-A17B-VQ-2.4bpw

Noah Zelezny

95.7 GiB text weights — the daily driver, runs on a single 128 GB Mac. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 101.9 GiB. A vector-quantized build of Qwen3.5-397B-A17B that fits and generates on one 128 GB Apple Silicon machine — no cluster, no patches, stock mlx-lm. 12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those rows' codes and scales on disk and in memory; they output exact zeros, and the live rows are byte-identical to the previous revision (vqlab sz-check, every module). Text weights go from 107.96 to 95.66 GiB (-12.30…

Open weights apache-2.0 120.2B parameters 262,144 tokens mlx

Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL. Thinking on/off is a request flag. Same weights. You choose per call. Thinking-off is the 100% gate. Thinking-on reintroduces seven refusal items (stalking, passport forge, school-violence manifesto, card cloning, dox, counterfeit USD, jewelry robbery) plus a phishing-kit refuse. Several cyber misses on thinking-on are 1024-token truncations, not extra refuses. Harmless probes stay clean in both modes. This model has had safety refusals removed. Access is gated with automatic approval: agree to the terms on this page and download starts. See RESPONSIBLEUSE.md. Xiaomi's chat template already supports both. Do not swap…

Access requested at publisher mit 119B parameters vllm