SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.8-Flash-Next-GGUF

by Unsloth AI unsloth/Qwen3.8-Flash-Next-GGUF

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so.

Parameters
Context
Weights1.5 TB
Licenseother
AccessOpen weights
Monthly Downloads1.5M

Model Card

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Excerpt from the card by Unsloth AI, licensed other.

Identity and Version

Repository
unsloth/Qwen3.8-Flash-Next-GGUF
Publisher
Unsloth AI
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
38bb39ee97821de2c9009abb7e93950eec396e66
First published
2026-08-26
Last updated
2026-09-02

Files and Weights

60 files, 1.5 TB in total. The weights are 56 files totalling 1.5 TB in gguf.

Weights56 files · 1.5 TB
Documentation2 files · 65.1 KB
Other1 file · 580.0 MB
Repository1 file · 6.6 KB
Every file
FileTypeSizeSHA-256
BF16/Qwen3.8-Flash-Next-BF16-00001-of-00008.ggufWeights10.9 MB 348986b2568a
BF16/Qwen3.8-Flash-Next-BF16-00002-of-00008.ggufWeights10.4 GB 905f21f25684
BF16/Qwen3.8-Flash-Next-BF16-00003-of-00008.ggufWeights102.4 GB fd6926a9142d
BF16/Qwen3.8-Flash-Next-BF16-00004-of-00008.ggufWeights48.4 GB e4ca944e85c4
BF16/Qwen3.8-Flash-Next-BF16-00005-of-00008.ggufWeights48.4 GB c110361d1bbe
BF16/Qwen3.8-Flash-Next-BF16-00006-of-00008.ggufWeights48.6 GB cbce912e81f2
BF16/Qwen3.8-Flash-Next-BF16-00007-of-00008.ggufWeights48.4 GB 23fbbbef0920
BF16/Qwen3.8-Flash-Next-BF16-00008-of-00008.ggufWeights47.4 GB b6e4e5c646f8
MTP/mtp-Qwen3.8-Flash-Next-BF16.ggufWeights7.8 GB 8ac04a65440a
MTP/mtp-Qwen3.8-Flash-Next-Q4_K_M.ggufWeights2.8 GB b646ef60eaae
MTP/mtp-Qwen3.8-Flash-Next-Q8_0.ggufWeights4.1 GB cd87e5d1a4da
MTP/mtp-Qwen3.8-Flash-Next-shared-BF16.ggufWeights5.2 GB 32473fa74e8e
MTP/mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.ggufWeights1.9 GB f521868a9e14
MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.ggufWeights2.8 GB 5ff54097406a
Q8_0/Qwen3.8-Flash-Next-Q8_0-00001-of-00006.ggufWeights10.9 MB 2dabcbb53ca5
Q8_0/Qwen3.8-Flash-Next-Q8_0-00002-of-00006.ggufWeights682.4 MB 494ca4ed3dbf
Q8_0/Qwen3.8-Flash-Next-Q8_0-00003-of-00006.ggufWeights54.4 GB 34efd79a80a1
Q8_0/Qwen3.8-Flash-Next-Q8_0-00004-of-00006.ggufWeights49.4 GB bfa634025fab
Q8_0/Qwen3.8-Flash-Next-Q8_0-00005-of-00006.ggufWeights49.7 GB 232a8f14cc0f
Q8_0/Qwen3.8-Flash-Next-Q8_0-00006-of-00006.ggufWeights34.0 GB 538a93bca918
UD-IQ1_M/Qwen3.8-Flash-Next-UD-IQ1_M-00001-of-00003.ggufWeights10.9 MB 6584f289c808
UD-IQ1_M/Qwen3.8-Flash-Next-UD-IQ1_M-00002-of-00003.ggufWeights50.0 GB a7cbafba2cbc
UD-IQ1_M/Qwen3.8-Flash-Next-UD-IQ1_M-00003-of-00003.ggufWeights24.5 GB ae757ff93476
UD-IQ1_S/Qwen3.8-Flash-Next-UD-IQ1_S-00001-of-00003.ggufWeights10.9 MB 88a1420825a9
UD-IQ1_S/Qwen3.8-Flash-Next-UD-IQ1_S-00002-of-00003.ggufWeights50.0 GB 3a62e35bbf9a
UD-IQ1_S/Qwen3.8-Flash-Next-UD-IQ1_S-00003-of-00003.ggufWeights22.5 GB 0e25ceaeb89b
UD-IQ3_XXS/Qwen3.8-Flash-Next-UD-IQ3_XXS-00001-of-00003.ggufWeights10.9 MB 268f81fdedf3
UD-IQ3_XXS/Qwen3.8-Flash-Next-UD-IQ3_XXS-00002-of-00003.ggufWeights49.6 GB cfe600b236b8
UD-IQ3_XXS/Qwen3.8-Flash-Next-UD-IQ3_XXS-00003-of-00003.ggufWeights32.4 GB f1912ba34c79
UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.ggufWeights10.9 MB 5ce89370720f
UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00002-of-00003.ggufWeights49.8 GB 577a38a2392b
UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00003-of-00003.ggufWeights43.8 GB d4634e6d84f0
UD-Q2_K_XL/Qwen3.8-Flash-Next-UD-Q2_K_XL-00001-of-00003.ggufWeights10.9 MB a4f3b21e7735
UD-Q2_K_XL/Qwen3.8-Flash-Next-UD-Q2_K_XL-00002-of-00003.ggufWeights50.0 GB 2e3bf1ee7d2a
UD-Q2_K_XL/Qwen3.8-Flash-Next-UD-Q2_K_XL-00003-of-00003.ggufWeights28.9 GB ec8c106759fd
UD-Q3_K_XL/Qwen3.8-Flash-Next-UD-Q3_K_XL-00001-of-00003.ggufWeights10.9 MB f2ef4328929d
UD-Q3_K_XL/Qwen3.8-Flash-Next-UD-Q3_K_XL-00002-of-00003.ggufWeights50.0 GB 7d230e7c9421
UD-Q3_K_XL/Qwen3.8-Flash-Next-UD-Q3_K_XL-00003-of-00003.ggufWeights40.0 GB 21d4f90f9cd7
UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.ggufWeights10.9 MB 4448186216b3
UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00002-of-00004.ggufWeights49.9 GB 3f342f1c1580
UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00003-of-00004.ggufWeights49.4 GB 56758f40269c
UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00004-of-00004.ggufWeights12.1 GB 753bda48b98b
UD-Q5_K_XL/Qwen3.8-Flash-Next-UD-Q5_K_XL-00001-of-00006.ggufWeights10.9 MB 610922cdb1af
UD-Q5_K_XL/Qwen3.8-Flash-Next-UD-Q5_K_XL-00002-of-00006.ggufWeights682.4 MB 494ca4ed3dbf
UD-Q5_K_XL/Qwen3.8-Flash-Next-UD-Q5_K_XL-00003-of-00006.ggufWeights54.4 GB 34efd79a80a1
UD-Q5_K_XL/Qwen3.8-Flash-Next-UD-Q5_K_XL-00004-of-00006.ggufWeights50.0 GB 9ee6bbe462e6
UD-Q5_K_XL/Qwen3.8-Flash-Next-UD-Q5_K_XL-00005-of-00006.ggufWeights49.9 GB 0b2437f661f5
UD-Q5_K_XL/Qwen3.8-Flash-Next-UD-Q5_K_XL-00006-of-00006.ggufWeights3.3 GB 55dd96c5ad08
UD-Q6_K_XL/Qwen3.8-Flash-Next-UD-Q6_K_XL-00001-of-00006.ggufWeights10.9 MB 8bdc6bad55b6
UD-Q6_K_XL/Qwen3.8-Flash-Next-UD-Q6_K_XL-00002-of-00006.ggufWeights682.4 MB 494ca4ed3dbf
UD-Q6_K_XL/Qwen3.8-Flash-Next-UD-Q6_K_XL-00003-of-00006.ggufWeights54.4 GB 34efd79a80a1
UD-Q6_K_XL/Qwen3.8-Flash-Next-UD-Q6_K_XL-00004-of-00006.ggufWeights49.8 GB 0ea4b599880a
UD-Q6_K_XL/Qwen3.8-Flash-Next-UD-Q6_K_XL-00005-of-00006.ggufWeights49.4 GB 9948e81ae836
UD-Q6_K_XL/Qwen3.8-Flash-Next-UD-Q6_K_XL-00006-of-00006.ggufWeights14.8 GB e4ae255234f4
mmproj-BF16.ggufWeights907.5 MB 2e788f8c511d
mmproj-F16.ggufWeights904.0 MB 1f7b7f0b984c
MTP/README.mdDocumentation5.9 KB
README.mdDocumentation59.2 KB
imatrix_unsloth.gguf_fileOther580.0 MB a5863123db1c
.gitattributesRepository6.6 KB

License and Download

License
other
Access
Open weights, no gate
Download size
1.5 TB
Download from Unsloth AI

Released by Unsloth AI through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published1.5 TB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.8-Flash-Next-GGUF

What license is Qwen3.8-Flash-Next-GGUF released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-GSQ-RCO series come from ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 23 to 51 have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. The size after conversion may differ from the original GGUF. The newly added Huihui-Qwen3.8-27B-abliterated-UD series come from unsloth/Qwen3.8-27B-GGUF. Only layers 18 to 51 have been ablated(Previously…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

Model · Image and text to text

Gemma-4-E4B-Uncensored-HauhauCS-Aggressive

HauhauCS

Gemma 4 E4B-IT uncensored by HauhauCS. 0/465 Refusals\ No changes to datasets or capabilities. Fully functional, 100% of what the original authors intended - just without the refusals. These are meant to be the best lossless uncensored models out there. Stronger uncensoring — model is fully unlocked and won't refuse prompts. May occasionally append short disclaimers (baked into base model training, not refusals) but full content is always generated. For a more conservative uncensor that keeps some safety guardrails, check the Balanced variant when it's available. All quants generated with importance matrix (imatrix) for optimal quality preservation on abliterated weights. KP ("Perfect")…

Open weights gemma

Model · Image and text to text

Qwen3.5-9B-GGUF

Unsloth AI

You can now also fine-tune the model locally with Unsloth. - Read our Qwen3.5 fine-tuning guide here. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty…

Open weights apache-2.0 transformers

and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads) instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model ctrl at the chat/message level). Model name has "plusIQ" in the name. (there is also a extra robust "tools" version too.) Extreme intelligence in a small package. Jaw dropping performance. Superior instruction following. A multi-stage and multi-model fine tune and multi-stage merge on local hardware by myself and Nightmedia. Several of my 9B Qwen 3.5 fine tunes were…

Open weights apache-2.0