SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.8-Flash-Next

by Tai Hua oODragoOo/Qwen3.8-Flash-Next

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so.

Parameters180B
Context262,144
Weights360.0 GB
Licenseother
AccessOpen weights
Monthly Downloads

Runs On

What it takes to serve Qwen3.8-Flash-Next (180B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 360.0 GB 432.0 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55
8-bit 180.0 GB 216.0 GB 1x MI325X (256 GB)
Vultr
$2.00 1x MI355X $2.59 · 2x MI300X $3.70
4-bit 90.0 GB 108.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Excerpt from the card by Tai Hua, licensed other.

Configuration

Architecture
Qwen4ExpForConditionalGeneration
Context length (tokens)
262,144
Layers
48
Hidden size
2,560
Attention heads
24
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
512
Experts active per token
10
Model type
qwen4_exp

Identity and Version

Repository
oODragoOo/Qwen3.8-Flash-Next
Publisher
Tai Hua
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
180B parameters
Languages
Not stated by the source
Revision
1fe131ec264c2e19b0123dbbc37db05c6253b2de
First published
2026-09-18
Last updated
2026-09-18

Files and Weights

144 files, 360.0 GB in total. The weights are 131 files totalling 360.0 GB in safetensors.

Weights131 files · 360.0 GB
Configuration5 files · 176.4 KB
Tokenizer4 files · 22.9 MB
Documentation2 files · 68.4 KB
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00131.safetensorsWeights1.0 GB 3a092d3a9a54
model-00002-of-00131.safetensorsWeights3.4 GB 1da92de7cf72
model-00003-of-00131.safetensorsWeights1.8 GB b783fbb3bdf1
model-00004-of-00131.safetensorsWeights3.4 GB 18b756227729
model-00005-of-00131.safetensorsWeights3.4 GB 54fdc7adf431
model-00006-of-00131.safetensorsWeights3.2 GB 8277a48e1fd7
model-00007-of-00131.safetensorsWeights3.2 GB 0fba369354e4
model-00008-of-00131.safetensorsWeights3.2 GB 024f59b6f2cb
model-00009-of-00131.safetensorsWeights3.2 GB 23fbef154f7f
model-00010-of-00131.safetensorsWeights3.2 GB b3a303529a19
model-00011-of-00131.safetensorsWeights3.2 GB d1215943b484
model-00012-of-00131.safetensorsWeights3.2 GB f63d649ea848
model-00013-of-00131.safetensorsWeights3.2 GB 5e0f0fb335f4
model-00014-of-00131.safetensorsWeights3.2 GB 5ee73bd21c73
model-00015-of-00131.safetensorsWeights3.2 GB b35421bf2172
model-00016-of-00131.safetensorsWeights3.2 GB 73c4b1bd1f37
model-00017-of-00131.safetensorsWeights3.2 GB ee60c1269974
model-00018-of-00131.safetensorsWeights3.2 GB 72cf3696aa1a
model-00019-of-00131.safetensorsWeights3.2 GB bdad649be407
model-00020-of-00131.safetensorsWeights3.2 GB 2258dad55787
model-00021-of-00131.safetensorsWeights3.2 GB 775952ac572f
model-00022-of-00131.safetensorsWeights3.2 GB 7d205070bb89
model-00023-of-00131.safetensorsWeights3.2 GB 598d4f772853
model-00024-of-00131.safetensorsWeights3.2 GB 09e534942e7d
model-00025-of-00131.safetensorsWeights3.2 GB 482d18f0e515
model-00026-of-00131.safetensorsWeights3.2 GB ac479fe29ed3
model-00027-of-00131.safetensorsWeights3.2 GB 4981d4e63672
model-00028-of-00131.safetensorsWeights3.2 GB 1b45a957b1f3
model-00029-of-00131.safetensorsWeights3.2 GB 40d2f2291ff4
model-00030-of-00131.safetensorsWeights3.2 GB 893a9b59e412
model-00031-of-00131.safetensorsWeights3.2 GB 9b6bf01d60f2
model-00032-of-00131.safetensorsWeights3.2 GB 7453fcb58721
model-00033-of-00131.safetensorsWeights3.2 GB 5b2f577ce409
model-00034-of-00131.safetensorsWeights3.2 GB 35b552266865
model-00035-of-00131.safetensorsWeights3.2 GB 4fe611218d36
model-00036-of-00131.safetensorsWeights3.2 GB 237789c0cef5
model-00037-of-00131.safetensorsWeights1.7 GB 73b53c94d235
model-00038-of-00131.safetensorsWeights3.4 GB 063965d1d57b
model-00039-of-00131.safetensorsWeights1.7 GB 7009a781a3d8
model-00040-of-00131.safetensorsWeights3.4 GB ed06bdece1c7
model-00041-of-00131.safetensorsWeights1.9 GB 0465251c296b
model-00042-of-00131.safetensorsWeights3.4 GB 071bb12fa66a
model-00043-of-00131.safetensorsWeights1.8 GB a88f04113821
model-00044-of-00131.safetensorsWeights3.4 GB bab6f9fd62f1
model-00045-of-00131.safetensorsWeights1.8 GB 50515cf9d024
model-00046-of-00131.safetensorsWeights3.4 GB f838e842fb3f
model-00047-of-00131.safetensorsWeights1.7 GB eb79852b4266
model-00048-of-00131.safetensorsWeights3.4 GB dbbf3a6e2b81
model-00049-of-00131.safetensorsWeights1.9 GB cbab68ff417b
model-00050-of-00131.safetensorsWeights3.4 GB 6014a2b71b74
model-00051-of-00131.safetensorsWeights1.8 GB 6c37a6533697
model-00052-of-00131.safetensorsWeights3.4 GB 1914762c6994
model-00053-of-00131.safetensorsWeights1.8 GB 1bf2bbe69b64
model-00054-of-00131.safetensorsWeights3.4 GB 22daa7622152
model-00055-of-00131.safetensorsWeights1.7 GB 7c4aef74ec00
model-00056-of-00131.safetensorsWeights3.4 GB 86ef0b5385fe
model-00057-of-00131.safetensorsWeights1.9 GB 77ec22970557
model-00058-of-00131.safetensorsWeights3.4 GB b7d2ab70d914
model-00059-of-00131.safetensorsWeights1.7 GB a1d01bd0bce7
model-00060-of-00131.safetensorsWeights3.5 GB 8fd21925cb11
model-00061-of-00131.safetensorsWeights3.4 GB cf5bff0a9663
model-00062-of-00131.safetensorsWeights3.5 GB e62ba3c6e2e9
model-00063-of-00131.safetensorsWeights3.4 GB e35a86ae511b
model-00064-of-00131.safetensorsWeights1.8 GB 08609ac9d06c
model-00065-of-00131.safetensorsWeights3.4 GB 040b73058fe0
model-00066-of-00131.safetensorsWeights1.7 GB 550a89f69aa7
model-00067-of-00131.safetensorsWeights3.4 GB 9223da6ff56d
model-00068-of-00131.safetensorsWeights1.9 GB 427d7d405528
model-00069-of-00131.safetensorsWeights3.4 GB f7d460aa332a
model-00070-of-00131.safetensorsWeights1.8 GB 9644884762cc
model-00071-of-00131.safetensorsWeights3.4 GB 184e042c9702
model-00072-of-00131.safetensorsWeights1.8 GB 2415f75c86da
model-00073-of-00131.safetensorsWeights3.4 GB 19de39abb6e1
model-00074-of-00131.safetensorsWeights1.7 GB 51152d64bc41
model-00075-of-00131.safetensorsWeights3.4 GB fc43b07ed481
model-00076-of-00131.safetensorsWeights1.9 GB 14860603fb56
model-00077-of-00131.safetensorsWeights3.4 GB a75bf91dc472
model-00078-of-00131.safetensorsWeights1.8 GB e9ac55c30320
model-00079-of-00131.safetensorsWeights3.4 GB e4cd07a12651
model-00080-of-00131.safetensorsWeights1.7 GB 549c3f9f5870
model-00081-of-00131.safetensorsWeights3.4 GB 4cf05b67fcb8
model-00082-of-00131.safetensorsWeights1.9 GB 0fe8d836b886
model-00083-of-00131.safetensorsWeights3.4 GB f097129297fd
model-00084-of-00131.safetensorsWeights1.7 GB f41be16d3413
model-00085-of-00131.safetensorsWeights3.4 GB 69083ff9730a
model-00086-of-00131.safetensorsWeights1.9 GB 757b82b98441
model-00087-of-00131.safetensorsWeights3.4 GB 9b9b8704d856
model-00088-of-00131.safetensorsWeights1.8 GB 2feec63478c3
model-00089-of-00131.safetensorsWeights3.4 GB 640ea88172a3
model-00090-of-00131.safetensorsWeights1.8 GB af53d439e4e2
model-00091-of-00131.safetensorsWeights3.4 GB 6ff4b9718910
model-00092-of-00131.safetensorsWeights1.7 GB c0c1b787fe74
model-00093-of-00131.safetensorsWeights3.4 GB 3d4f44c44752
model-00094-of-00131.safetensorsWeights1.9 GB c31221b8e6a5
model-00095-of-00131.safetensorsWeights3.4 GB 9e680b695a5b
model-00096-of-00131.safetensorsWeights1.8 GB a30732e7aadf
model-00097-of-00131.safetensorsWeights3.4 GB 19be76593bce
model-00098-of-00131.safetensorsWeights1.8 GB 5c5e4956c7f7
model-00099-of-00131.safetensorsWeights3.4 GB 21daceed016e
model-00100-of-00131.safetensorsWeights1.7 GB 78da7091e9a3
model-00101-of-00131.safetensorsWeights3.4 GB 2a596073cf66
model-00102-of-00131.safetensorsWeights1.9 GB 6caf69124851
model-00103-of-00131.safetensorsWeights3.4 GB cc197d5c0d5a
model-00104-of-00131.safetensorsWeights1.8 GB 59ee8b7c1f07
model-00105-of-00131.safetensorsWeights3.4 GB 4040abee30dc
model-00106-of-00131.safetensorsWeights1.9 GB 2903f5b2c539
model-00107-of-00131.safetensorsWeights3.4 GB 08fb87b54a96
model-00108-of-00131.safetensorsWeights1.9 GB 426251e0773e
model-00109-of-00131.safetensorsWeights3.4 GB 67252c3ea4da
model-00110-of-00131.safetensorsWeights1.7 GB 66ecf98a5a8d
model-00111-of-00131.safetensorsWeights3.4 GB fe316b2426ea
model-00112-of-00131.safetensorsWeights1.9 GB 90d2e46457dd
model-00113-of-00131.safetensorsWeights3.4 GB ef41245305a2
model-00114-of-00131.safetensorsWeights1.8 GB b8d004974b91
model-00115-of-00131.safetensorsWeights3.4 GB b536b5337a35
model-00116-of-00131.safetensorsWeights1.8 GB b8387f98bdf7
model-00117-of-00131.safetensorsWeights3.4 GB 27b091d8a571
model-00118-of-00131.safetensorsWeights1.7 GB ffe5cc04b3c2
model-00119-of-00131.safetensorsWeights3.4 GB fae555008bac
model-00120-of-00131.safetensorsWeights1.9 GB 6a13f7374f79
model-00121-of-00131.safetensorsWeights3.4 GB e00e79982774
model-00122-of-00131.safetensorsWeights1.8 GB 686b6d76dc53
model-00123-of-00131.safetensorsWeights3.4 GB bb8e5b70612d
model-00124-of-00131.safetensorsWeights1.7 GB bf22358402fa
model-00125-of-00131.safetensorsWeights3.4 GB b91a643909b5
model-00126-of-00131.safetensorsWeights1.9 GB a088d1d8bbda
model-00127-of-00131.safetensorsWeights3.4 GB fa0c609906b1
model-00128-of-00131.safetensorsWeights1.8 GB abfe513d53dd
model-00129-of-00131.safetensorsWeights3.4 GB d24efe734057
model-00130-of-00131.safetensorsWeights3.0 GB 2d651616c258
model-00131-of-00131.safetensorsWeights1.3 GB 50be0ccd11e4
config.jsonConfiguration4.7 KB
generation_config.jsonConfiguration202 B
model.safetensors.index.jsonConfiguration170.7 KB
preprocessor_config.jsonConfiguration390 B
video_preprocessor_config.jsonConfiguration385 B
LICENSEDocumentation3.2 KB
README.mdDocumentation65.2 KB
chat_template.jinjaOther9.0 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer3.4 MB
tokenizer.jsonTokenizer12.8 MB 0997f410c57a
tokenizer_config.jsonTokenizer17.9 KB
vocab.jsonTokenizer6.7 MB

License and Download

License
other
Access
Open weights, no gate
Download size
360.0 GB
Download from Tai Hua

Released by Tai Hua through its official repository on Hugging Face.

Memory Requirements

PrecisionWeights in memory
As published360.0 GB
16-bit360.0 GB
8-bit180.0 GB
4-bit90.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.8-Flash-Next

How much GPU memory does Qwen3.8-Flash-Next need?

About 432 GB at 16-bit and 108 GB at 4-bit: the weights (180B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.8-Flash-Next on?

At 16-bit, 2x MI325X from $4.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Qwen3.8-Flash-Next released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Qwen3.8-Flash-Next's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-Flash-Next

Qwen

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next-Uncensored-NVFP4

OrcaRouter

This model has had its safety alignment substantially removed via abliteration (orthogonalizing the refusal direction out of the residual stream). It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment. Use must comply with the Apache 2.0 License inherited from the base model and all applicable law. The authors accept no liability for misuse. - A Blackwell GPU…

Access requested at publisher apache-2.0 180B parameters transformers

Model · Image and text to text

Qwen3.8-Flash-Next-MLX-oQ3-MTP

Robot Haus

A sensitivity-guided, mixed-precision MLX conversion of Qwen/Qwen3.8-Flash-Next, rebuilt directly from the official BF16 checkpoint with the model's matching native MTP block preserved. oQ3 uses a 3-bit affine base and spends additional precision on sensitive modules. Layer sensitivity was measured with a validated quantized calibration proxy, while every released weight was quantized from the official BF16 checkpoint. The result is a compact model with 746 higher-precision module overrides rather than a uniform 3-bit layout. The upstream tokenizer, current chat template, vision processor, generation configuration, licence, and native MTP configuration are retained. In a compatible oMLX…

Open weights other 180B parameters 262,144 tokens mlx

Model · Image and text to text

Qwen3.5-122B-A10B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 125.1B parameters 262,144 tokens transformers

Model · Image and text to text

GLM-5.3-Flash

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.3-Flash blog and GLM-5 Technical report. Use GLM-5.3-Flash API services on Z.ai API Platform. We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear…

Open weights mit 321.3B parameters 1,048,576 tokens transformers

Model · Image and text to text

Qwen3.6-35B-A3B-FP8

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 36B parameters 262,144 tokens transformers