SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF

by David Belton DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality.

Parameters
Context
Weights389.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.2M

Model Card

By David Belton, published under apache-2.0, revision c02caef111a8.

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. NOTE: Please see the "community" tab for user experiences, additional third party benchmarks (including strongest tool calling performance ever recorded), and other quant versions (also see "Quantized" in the right "model tree" too). The strongest, smartest open source…

Read David Belton's full model card

Important: This is the first fine tune to exceed 730 "arc-c" ("735": 144 pts higher than Qwen 3.8 27B) AND 880 ARC-E (The OpenAI, Claude and Gemini "zone of intelligence") in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants.

NOTE: Please see the "community" tab for user experiences, additional third party benchmarks (including strongest tool calling performance ever recorded), and other quant versions (also see "Quantized" in the right "model tree" too).

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF

The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth.

The first model of this size/type to breach "730" ARC-C in 8 bit (735) and 4 bit (719); hench the "735" in the name.

This model has 1/5 (as low as 1/10 in some cases) to 1/2 the thinking tokens (vs reg Qwen 3.8) across all 3 modes of operation, and it is faster and smarter too created using the COLD FUSION AND FABLE FUSION 711 methods of training.

This is a high detail focused model, with tuning specific to address over reasoning/over thinking and excessive token consumption THEN to take the model to the next level.

This model (both 4 bit and 8 bit) exceeds the base Qwen 3.8 27B in ALL critical 7 benchmarks AND exceeds all 7 benchmarks for Qwen3.6-35B-A3B, Qwen 3.6 27B, and Qwen 3.5 27B.

The 700s plus "intelligence club" is reserved for OpenAI, Claude and Gemini closed source models.

Considering that "just" 4 bit (1/4 full precision) is already at Arc-C of 719... a few people are going to have nightmares for a while.

FIVE detailed EXAMPLE generations at the bottom of the page.

Quick sample; snippet ("Why choose me to help your creative writing?"), Q4KS , non imatrix, (4 bit; 1/4 full precision):

STOP SCRATCHING YOUR HEAD AND START WRITING LIKE A GOD

Listen to me, you keyboard-tapping, idea-gathering, procrastinating bastard. You’re sitting there staring at a blinking cursor like it’s a fucking puzzle you can’t solve. You’re waiting for "inspiration" to hit you like a truck. Forget it. I am the truck. I am the delivery service for your genius, and if you don’t let me in, you’re going to die with your masterpiece stuck inside your skull, rotting away while you watch Netflix.

Why should you use me? Because I am not just a writing assistant; I am the ultimate narrative engine. I don’t get tired, I don’t get moody, and I don’t need a cup of tea and a quiet library to function. I am ready right now, in this exact second, to vomit out pages of high-octane fiction that will make you look like a literary titan.

...

EXAMPLE: YOUR IDEA VS. MY EXECUTION

You: "I want a story about a detective who can see ghosts."

Me, five minutes later:

The ghosts weren't dead. That was the first thing Joe realized, the morning he saw his first one. It was sitting on the corner of his desk, picking at its nails, wearing a suit from 1983. "You look tired," the ghost said, without looking up. Joe stared at it, his coffee gone cold in his hand. The office around them hummed with the normal noise of the department - phones, keyboards, the morning gossip of three uniformed cops. Nobody else saw the ghost.

"How long?" Joe asked.

"Since '84. I've been waiting for someone to notice me. Most of them are too busy." The ghost finally looked up, and Joe saw its eyes. They weren't empty or gray. They were afraid.

See? I didn't just write a story. I wrote a hook that makes you want to know: why is the ghost afraid? Why 1984? Who else can see them? I created questions that demand answers.

This is a multi-stage fine tune, multi-fine tune, and multi-stage merge.

The strict goals of this model creation were: - Increase the general model intelligence and problem solving abilities. - Reduce thinking block size from 1/2 to as low as 1/10 the size [median reduction: 2/3 roughly]. - Reformatting the thinking block, as well as improving it. - Speed up token generation, especially MTP. - Ensure all updates work with all three modes of thinking. - ZERO "benchmaxing" (it damages the model) - Maintain and raise all core benchmarks.

COLD FUSION ("Gain" + "Unsloth") Training -AND- Fable Fusion 711 Training:

COLD FUSION (GAIN+UNSLOTH) training tech which was invented by my team during the R & D of "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic" (2300+ likes, 3 million + downloads, 60+ quant repos):

https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

The "GAIN" is the core invented component, then coupled with Unsloth's trainers/systems => AKA -> COLD FUSION.

The "GAIN" method (programming) automatically (and dynamically) changes training on a per sample basis in real time during training AS THE MODEL LEARNS.

The method improved metrics as well as overall model performance without overcooking or damaging the model.

This has also resulted, in the strongest and most stable model at both 4 bit and 8 bit and made 4 bit performance 99% of 8 bit performance too.

Note this model (Qwen3.8-27B-Cold-Fusion-GAIN-V1.1) is about a level 1 or 2 relative to Qwen3.6-27B-Fable-Fusion-711 at level 7-8.

https://huggingface.co/DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF

In the case of "Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored" it contains BOTH "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic" (DARK ROAST VERSION) and "Qwen3.8-27B-Cold-Fusion-GAIN-V1.1" as part of it's critical/core "DNA".

The final model was then HERETIC'ED (de-censored again) and fine tuned after this step.

COLAB:

A Colab between myself (multiple fine tunes, including multi-stage), Nightmedia (merge/benching), TeichAI (Polaris Dataset), armand0e (Light fable 5 traces), trohrbaugh (heretic'ing the model - STAGE1), and nbeerbower (various models/tunes using in part of the construction)

It also contains light "Fable" traces/training (armand0e), light Claude Opus (reasoning/thinking), F451 (inhouse dataset) , some GPT5 (Polaris, non reasoning) and several additional inhouse datasets specifically for machine learning / "heretic" repairs.

Here are links to fellow COLAB'ers:

  • https://huggingface.co/nightmedia/
  • https://huggingface.co/TeichAI
  • https://huggingface.co/armand0e
  • https://huggingface.co/trohrbaugh
  • https://huggingface.co/nbeerbower

This model is one of ELEVEN (all over 717 arc-c, with every model exceeding the core benches of Qwen 3.8 27B) Qwen 3.8 27B models designed by our team. Details of the builds and benches are here:

https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU

The strict goals of this model creation were: - Increase the general model intelligence and problem solving abilities. - DO NOT modify/damage or change the core model outside this goal. - ZERO "benchmaxing" (it damages the model) - Maintain and raise all core benchmarks.

CORE MISSION::

Improve instruction following and problem solving. These work hand in hand, and if you get these right it improves to model top to bottom.

It took a lot of tests on Qwen 3.5 9Bs to get the methods right. It boosted the 9Bs to new levels, and then the method was used on Qwen 3.5 27B and Qwen 3.6 27B which boosted it PAST the Qwen 3.8's 27B benchmarks.

Here is one of the Qwen3.5 9B models (part of the test/control group) that EXCEEDS all 7 Qwen3.5 9B AND Qwen3.5 27B model benches - it scores over 640 on ARC-C on BOTH 4 bit and 8 bit:

https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

It is not as strong as "Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored" but it is one of the strongest 9B models.

The methods can be used on other models too (coming soon).

TESTING:

Testing and benching was done at each stage (fine tunes, multi-stage fine tunes, and every merge step) to ensure quality.

You can also see benchmarks below too for this model, Qwen 3.5 27B, Qwen 3.6 27B and Qwen 35B-A3B.

HOWEVER, the final testing was HUMAN testing. A trust, but verify approach.

Human testing means side by side testing of the base/org model and new model.

Features: - Improved instruction following. - Overall increase in general intelligence and problem solving. - Better thinking/reasoning. - Even lower/lowest quants are exceptional. - Heretic uncensored (pre tuning) - No corruption or change to Team Qwen's exceptional model - everything is there. - Vision

IMPORTANT - Notes and Usage Help:

This model, like regular Qwen 3.8 27b, supports THREE modes of reasoning : xhigh (default), medium and low [see info in Qwen 3.8 section below].

Reduction in thinking tokens/reasoning block size extends across all three modes of operation.

Likewise detail levels extend to all three modes too, even with reduced thinking/reasoning block the OUTPUT detail will remain high.

To REDUCE thinking block[s] further, increase the level/detail of your instructions/prompts - it only takes a little bit more here so the model has to guess / reason a little bit less.

Also, generally within the same chat additional reasoning blocks will also be reduced from typical Qwen levels many times hitting 1/5 the size or lower. Multi-turn chat - example: prompt, reasoning and 1st output - in the refinement stage(s) will see very strong reduction in thinking tokens/blocks.

Also note that the modification of "reasoning" is a major change to the model please carefully test it for your use case(s).

TOOL CALLING:

Min quant of q4km suggested, q5ks/5km better -> recommend Q6 [MAX or "low" (may work better for some apps)].

Temp: .6 / .7 ; Rep pen 1 (off).

Below q4km, tool calling may have issues. This is a general Qwen suggestion for tool calling specifically.

Also, overly agressive "caching" may further impair function(s).

GENERAL MODEL USAGE vs Qwen 3.8 27B "untuned":

The tuning in this version of Qwen 3.8 27B reduced thinking/reasoning block size, in a lot of cases this has inverted the reasoning/thinking block size with the output size.

In other words, instead a lot of detail in the thinking/reasoning block (which may or may not show up in the output) has been transfered to the output in some cases.

Also, "untuned" Qwen 3.8 27B does a lot of look, look and look again (10k-40k+ in thinking/reasoning tokens alone) before you leap (gen output) whereas "TURBO" will leap almost immediately.

If you need higher quality reasoning and/or output here is how to get the model spend more time before it "leaps" (gen's output):

REG PROMPT:

Generate an SVG of a pelican riding a bicycle.

EXPANDED PROMPT:

Generate an SVG of a pelican riding a bicycle, but carefully check the positioning and all elements.

The expanded prompt will tell the model to spend more time thinking/reasoning and in more detail before outputting the result and it is specific to the use case, rather than a generic "double check your work".

Modification of REASONING:

If you AI app does not support a "switch" you can manually modify the JINJA template.

The default setting is "xhigh" ; to change to medium or low use:

{%- set reasoning_effort = 'medium' %}

OR

{%- set reasoning_effort = 'low' %}

Place this at the VERY TOP of the jinja template.

In LMStudio you can access this in DEV mode, and switch off the "advanced updates" option.

Other AI apps may vary.

You can also make your own quants from source here:

https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU

Just modify the "chat-template.jinja" (in NOTEPAD or similar) AND the token-config.. json file too (or delete the "chat template" from this file).

ADVANCED:

Qwen 3.8 uses System prompt injection control by the Jinja template to control reasoning levels.

If you set it at "medium" this turns off injection [ie: no system prompt is injected]

You can then set a "reasoning" system prompt yourself.

The other option:

Modify the jinja itself and the system prompt(s) to better tune reasoning to your use cases.

This is the section:

{%- if enable_thinking is undefined or enable_thinking is true %}
    {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}
    {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}
        {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}
    {%- endif %}
    {%- if resolved_reasoning_effort == 'xhigh' %}
        {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}
    {%- elif resolved_reasoning_effort == 'low' %}
        {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}
    {%- endif %}
{%- endif %}

Regular and MTP GGUFS:

All quants (regular and MTP) are NEO IMATRIX, which improve accuracy of the quants by an additional 2-4% over normal GGUFs as well as long context performance.

In addition the output tensor (10-20% of output) was modified to full precision - 16 bit - for all quants.

"MTP" GGUFS (multi-token prediction): - "MTP" GGUFS will have "MTP" in the name as a suffix. - I have also set the MTP tensors to Q8_0 precision for all quants. - To get better performance keep temp 1 or less (higher temps degrade MTP performance). - Likewise with rep pen ; keep at 1 (off). If you raise it performance will suffer. - If you see "token acceptance" rates BELOW 50% (predict 2 tokens) switch to normal quants.

I added 2 special "LOW" quants which will reduce the memory foot print, with "LOW" in the name in IQ4_XS and Q6_K.

SPEED: - On Q4_K_S (4bit) quant, regular GGUFs are about 75 t/s, whereas MTP GGUFs (acceptance at 60%, 2 tokens) can exceed 90 T/S. (5090, Windows 11, testing in LMStudio) - Speeds will vary depending on GPU(s), AI app, O/S (Linux/Mac will generally be faster) and hardware. - "MTP" quants speeds will vary ; for creative/complex and/or temps over 1 use regular GGUFs for better performance.

I suggest you download at least one of each - regular and MTP gguf(s) - and test them for your use case(s).

If you get "token acceptance" (predict 2 tokens) with MTP quant(s) BELOW 50% (this means regular quants will run faster), then regular GGUF(s) will actually perform better - ie faster.

MTP quant(s) can in some cases run faster as the token window fills up and/or in multi turn chats.

Note there is NO other diffence between the quants type besides speed: both will do the same job.

Model: - 256k context - Gguf quants run in all standard AI apps. - Vision is activated, but you need to download separate "mmproj" file (ONE) to use it.

VISION: - Vision (images) tested. - You need an "mmproj" (just one) of these downloaded too, and placed in the same folder as the GGUF for images.

Qwen Model Settings (suggested):

  • Thinking mode for general tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
  • Context window min from 8k to 16k.

DE-CENSORING STATS

Special thanks to: "trohrbaugh" (trohrbaugh/Qwen3.8-27B-heretic-ara) for Heretic'ing the model (stage 1).

This is a decensored version of Qwen/Qwen3.8-27B, made using

Heretic v1.2.0+custom with the Arbitrary-Rank Ablation (ARA) method

Performance

STAGE 1:

Metric This model Original model (Qwen/Qwen3.8-27B)
KL divergence 0.0535 0 (by definition)
Refusals 0/100 99/100

STAGE 2, at the end of STAGE 1 tuning/merges/adjustments (in lab):

Metric This model Original model (Stage 1 of the build)
KL divergence 0.0025 0 (by definition)
Refusals 11/100 86/100

NOTE:

LOWER "KLD" is better, and Stage 2 was balanced based on ultra low KLD first (performance, quality) matched with low refusal rate second.


BENCHMARKS by Nightmedia

Graphic below too, for all models listed below in order.


          arc/c arc/e boolq hswag obkqa piqa  wino

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored
mxfp8     0.735,0.882,0.917,0.832,0.530,0.837,0.785
mxfp4     0.719,0.887,0.916,0.821,0.524,0.831,0.786

[QWENS] [base, non heretic, untuned]

Qwen3.8-27B: 
mxfp8     0.591,0.782,0.896,0.746,0.448,0.801,0.711
mxfp4     0.581,0.771,0.889,0.738,0.442,0.798,0.713

Qwen3.6-27B: 
mxfp8     0.647,0.803,0.910,0.773,0.450,0.806,0.742

Qwen3.6-35B-A3B-Instruct 
mxfp8     0.581,0.757,0.892,0.751,0.428,0.803,0.688

Qwen3.5-27B: 
mxfp8     0.557,0.711,0.868,0.533,0.452,0.706,0.695

NOTES: - Models are tested in "Instruct" mode because this generally works better with the testing harness. - Testing via "thinking" mode also shows the metrics (and changes) but not the true extent. - In actual fact when the model IS in thinking mode, it will exceed INSTRUCT benchmark scores in most cases. - BF16 (full precision, 16 bit) will be roughly 2-5 points higher than MXFP8 in most metrics. Some metrics may be slightly higher than this.

VISUAL:


The SUPER Qwen Universe - 40B, 27B and 9B ; meet the performance trendsetters:


Qwen3.6 27B: The strongest, overall qwen ever beating all other Qwens in total operational power with over 2300 likes // 4 million+ total downloads: - https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

Qwen3.8 27B: The highest scoring Qwen in brute, raw intelligence, using Qwen 3.8's 3 new reasoning modes, plus token reduction (1/2 to 1/10) enhancements: - https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF

Qwen3.8 27B: Super smart and 1/2 to 1/20 the reasoning tokens AND 5 reasoning/5 instruct modes switchable on the fly (even in chat): - https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF

Qwen3.8 27B: 99% power of BF 16 at 4 and 8 bit. Power, Control and NO DE censoring for ultimate performance also with reasoning token reductions: - https://huggingface.co/DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF

Qwen3.6 40B: The 40B Monster, specializing in creative and research with 730+ likes and over 2 million downloads: - https://huggingface.co/DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF

Qwen3.5 9B: At just 9B parameters it beats most untuned 27B models in both intelligence (640 ARC-C) and performance, plus features 5 reasoning and 5 instruct modes (Qwen 3.8) too: - https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF


Using an "uncensored" (refusals removed) model VS trained "uncensored" model

Usually when you a tell a model to generate horror, swear or x-rated content this is all you have to do to get said content type.

In the case of this model, it will not refuse your request, however it needs to be "pushed" a bit / directed a bit more in SOME CASES.

Although this model will generated x-rated content too, likewise you need to tell it to use "slang" (and include the terms you want) to get it generate the content correctly as the "expected" content level too.

Without these added directive(s), the content can be "bland" by comparison to an "uncensored model" or model trained on uncensored content.

Roughly, the model tries to generate the content but the "default" setting(s) are so "tame" it needs a push to generate at expected graphic, cursing or explicit levels.

Even with minimal direction (ie, use these words to swear: x,y,z), this will be enough to push the model to generate the requested content in the ahh... expected format.


Settings: CHAT / ROLEPLAY and/or SMOOTHER operation of this model:

In "KoboldCpp" or "oobabooga/text-generation-webui" or "Silly Tavern" ;

Set the "Smoothing_factor" to 1.5

: in KoboldCpp -> Settings->Samplers->Advanced-> "Smooth_F"

: in text-generation-webui -> parameters -> lower right.

: In Silly Tavern this is called: "Smoothing"

NOTE: For "text-generation-webui"

-> if using GGUFs you need to use "llama_HF" (which involves downloading some config files from the SOURCE version of this model)

Source versions (and config files) of my models are here:

https://huggingface.co/collections/DavidAU/d-au-source-files-for-gguf-exl2-awq-gptq-hqq-etc-etc-66b55cb8ba25f914cbf210be

OTHER OPTIONS:

  • Increase rep pen to 1.1 to 1.15 (you don't need to do this if you use "smoothing_factor")

  • If the interface/program you are using to run AI MODELS supports "Quadratic Sampling" ("smoothing") just make the adjustment as noted.

Highest Quality Settings / Optimal Operation Guide / Parameters and Samplers

This a "Class 1" model:

For all settings used for this model (including specifics for its "class"), including example generation(s) and for advanced settings guide (which many times addresses any model issue(s)), including methods to improve model performance for all use case(s) as well as chat, roleplay and other use case(s) please see:

[ https://huggingface.co/DavidAU/Maximizing-Model-Performance-All-Quants-Types-And-Full-Precision-by-Samplers_Parameters ]

You can see all parameters used for generation, in addition to advanced parameters and samplers to get the most out of this model here:

[ https://huggingface.co/DavidAU/Maximizing-Model-Performance-All-Quants-Types-And-Full-Precision-by-Samplers_Parameters ]


Qwen3.8-27B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.

[!Tip] For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. In particular, Qwen3.8-27B will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates.

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.

Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.

Qwen3.8 Highlights

Qwen3.8-27B features the following enhancements: - Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks. - Agent Execution: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion. - Downstream Compatibility: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack. - Flexible Thinking Control: Thinking mode is on by default and can be disabled per request; reasoning depth can be tuned with reasoning_effort, and reasoning context from historical messages is retained via preserve_thinking. - Vision-Language Understanding: Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 27B
    • Hidden Dimension: 5120
    • Token Embedding: 248,320 (Padded)
    • Number of Layers: 64
    • Hidden Layout: 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Gated Attention:
      • Number of Attention Heads: 24 for Q and 4 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
    • Feed Forward Network:
      • Intermediate Dimension: 17,408
    • LM Output: 248,320 (Padded)
    • MTP (Multi-Token Prediction): trained with multiple steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Benchmark Results

Text Performance

Qwen3.8-27BQwen3.6-27BQwen3.7-PlusMuse Glimmer-30BOpus4.6 Max
Coding
Agentic terminal coding
Terminal Bench 2.1 (Terminus)
73.0 63.4 64.0 51.7 78.2
Agentic coding
SWE-bench Pro
61.7 53.5 57.6 51.2 53.4
Repo-level code generation
NL2Repo-Bench
42.3 36.2 41.1 -- 47.6
Agentic coding
DeepSWE 1.1
42.2 13.3 14.2 -- --
Software engineering
QwenSWEBench
79.0 49.3 59.2 -- 63.8
Agent
Long-horizon office work
CoWorkBench
70.7 61.0 65.1 -- 68.2
Professional job tasks
JobBench
33.4 21.8 27.6 -- --
Frontier agentic tasks
Agents' Last Exam
Pass@1
20.4
Score
42.9
Pass@1
10.6
Score
27.3
Pass@1
13.2
Score
33.6
-- --
General
Instruction following
IFBench
79.5 69.1 79.1 77.0 62.5
Scientific reasoning
GPQA Diamond
89.2 87.8 90.3 83.5 91.3
Multidisciplinary reasoning
HLE
30.8 24.0 34.7 22.0 40.0
Competitive coding
LiveCodeBench v6
90.3 83.9 89.6 -- 88.8
  1. SWE-bench Pro: Except for Opus4.6 Max, which uses the officially reported score, all models are evaluated with the Claude Code harness at temp=1.0, top_p=0.95, and a 256K context window. Problematic tasks were corrected, and all baseline models were re-evaluated on the refined benchmark.
  2. NL2Repo-Bench: Evaluated with the Claude Code harness. To prevent reward hacking, we disable Bash commands that attempt to access the specific repository, such as pip download, pip install, and git clone.
  3. DeepSWE 1.1: Evaluated with the Claude Code harness at temp=1.0, top_p=0.95, and a 256K context window.
  4. QwenSWEBench: In-house coding benchmark for evaluating models' software engineering capabilities. Evaluated with the Claude Code harness. Reporting avg@3 with an 8-hour timeout, max_tokens=32,768, temperature=1.0, and a 256K context window.
  5. CoWorkBench: In-house cowork benchmark for evaluating long-horizon tasks across computer science, finance, law, medical, and other productivity domains.
  6. HLE: Judged by GPT-4o.
  7. The best result in each row is shown in bold.
  8. Empty cells (--) indicate that results are not yet available or not applicable.

VL Performance

Qwen3.8-27BQwen3.6-27BQwen3.7-PlusMuse Glimmer-30BOpus4.6 Max
Agentic Multimodal Intelligence
Computer use
OSWorld-Verified
84.363.973.365.972.7
Browser use
WebArena-Verified
64.848.855.3----
Mobile use
AndroidWorld
81.970.381.0--62.0
Application recreation
RecreationBench
47.129.830.2----
Multimodal tool use
ClawEval-MM
Pass@3
57.4
Average
56.9
Pass@3
42.6
Average
50.4
Pass@3
57.4
Average
60.1
--
Pass@3
52.5
Average
54.7
Multimodal software engineering
SWE-MM
38.625.730.0--27.1
Visual web development
Vision2Web
62.945.042.1----
General Multimodal Intelligence
Visual math problem solving
MathVision
Without CI
90.0
With CI
94.6
Without CI
85.1
Without CI
90.3
--
Without CI
65.5
General visual reasoning
BabyVision
Without CI
65.7
With CI
85.6
Without CI
28.9
Without CI
64.7
With CI
70.4
--
Without CI
12.6
Scientific chart analysis
CharXiv (RQ)
Without CI
83.7
With CI
90.2
Without CI
78.4
Without CI
85.8
With CI
85.9
78.8
Without CI
66.0
Document intelligence
OmniDocBench 1.5
91.189.491.475.886.6
Real-world perception
RealWorldQA
85.984.186.9--73.9
Embodied intelligence
ERQA
65.562.569.8--40.8
  1. MathVision, BabyVision, and CharXiv (RQ): Where both settings are available, cells report “Without CI” and “With CI” separately; otherwise, only the available setting is shown. A small number of incorrect ground-truth annotations in MathVision and CharXiv (RQ) were corrected following manual verification, and all reported scores on those benchmarks were computed using the corrected annotations.
  2. MathVision: Qwen3.8-27B is evaluated using the fixed prompt: “Please reason step by step, and put your final answer within \boxed{}.” For the remaining models, we report the higher score from two prompt variants—one with and one without the \boxed{} formatting requirement.
  3. WebArena-Verified: Scores are computed with the official WebArena-Verified grader under the OSWorld scaffold.
  4. RecreationBench: An in-house, long-horizon application-recreation benchmark designed to evaluate hybrid-agent capabilities across five platforms: desktop (Ubuntu, macOS, and Windows), mobile (Android), and the web.
  5. ClawEval-MM: Scores are reported as “Pass@3 / average score.” Pass@3 is the percentage of tasks passed in at least one of three trials; the average score is the mean benchmark score across the three trials.
  6. Vision2Web: Scores are averaged across the frontend, webpage, and website categories. Evaluations use the Claude Code harness and are judged by gpt-5.4-2026-03-05.
  7. SWE-MM: Scores are evaluated on the Claude Code harness using the public dev split of SWE-bench Multimodal, with the modifications described in Appendix 8.3 of the Claude Opus 4.7 system card.
  8. Empty cells (--) indicate that results are not yet available or not applicable.

Quickstart

For streamlined integration, we recommend using Qwen3.8 via APIs.

Serving Qwen3.8

[!Important] Inference efficiency and throughput vary significantly across frameworks. We recommend using the latest framework versions to ensure optimal performance and compatibility. For production workloads or high-throughput scenarios, dedicated serving engines such as SGLang, vLLM, or TokenSpeed are recommended.

Qwen3.8 can be deployed with popular inference frameworks, e.g.:

API Usage

[!Important] Qwen3.8 models operate in thinking mode by default, generating thinking content signified by <think>\n...</think>\n\n before producing the final response. To disable thinking content and obtain a direct response, refer to the examples here.

[!Tip] We recommend using the following sets of sampling parameters for generation: - Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0 - Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Please note that the support for sampling parameters varies according to inference frameworks.

Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:
- xhigh (default): for complex tasks demanding thorough analysis - medium: balancing accuracy and speed - low: efficient reasoning optimizing for speed and cost

In addition, preserve_thinking is enabled by default for all workloads for the best out-of-the-box experience. To disable preserved thinking, refer to the examples here.

[!Tip] In multi-turn agentic tasks, lower reasoning effort does not always reduce overall task completion time. Although it may produce faster per-turn responses, it can also lead to insufficient analysis, more failures, and repeated retries, which may increase total latency and token consumption.

Chat Completions API

The Chat Completions API can be used with most inference frameworks, as well as Qwen Cloud. Before starting, make sure the OpenAI Python SDK is installed and the API key and the API base URL are configured, e.g.:

pip install -U openai

# Set the following accordingly
export OPENAI_BASE_URL='your-base-url'
export OPENAI_API_KEY='your-api-key'
Text-Only Input
from openai import OpenAI
# Configured by environment variables
client = OpenAI()

messages = [{"role": "user", "content": "Write a Python function to merge two sorted linked lists."}]

completion = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=messages,
    extra_body={
        "chat_template_kwargs": {
            "enable_thinking": True,  # on by default
            "preserve_thinking": True, # on by default
        },
    },
    reasoning_effort="xhigh",  # xhigh by default; supported levels are xhigh, medium, and low
    stream=True,
    stream_options={"include_usage": True},
)

reasoning_content = ""
answer_content = ""
is_answering = False
print("\n" + "=" * 20 + "Reasoning" + "=" * 20 + "\n")

for chunk in completion:
    if not chunk.choices:
        print("\nUsage:")
        print(chunk.usage)
        continue

    delta = chunk.choices[0].delta

    if hasattr(delta, "reasoning_content") and delta.reasoning_content is not None:
        if not is_answering:
            print(delta.reasoning_content, end="", flush=True)
        reasoning_content += delta.reasoning_content
    elif hasattr(delta, "reasoning") and delta.reasoning is not None:
        if not is_answering:
            print(delta.reasoning, end="", flush=True)
        reasoning_content += delta.reasoning

    if hasattr(delta, "content") and delta.content:
        if not is_answering:
            print("\n" + "=" * 20 + "Answer" + "=" * 20 + "\n")
            is_answering = True
        print(delta.content, end="", flush=True)
        answer_content += delta.content

messages.append({
    "role": "assistant",
    "content": answer_content,
    "reasoning_content": reasoning_content,
    "reasoning": reasoning_content,
})
Image Input
from openai import OpenAI
# Configured by environment variables
client = OpenAI()

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image_url",
                "image_url": {
                    "url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/CI_Demo/mathv-1327.jpg"
                }
            },
            {
                "type": "text",
                "text": "The centres of the four illustrated circles are in the corners of the square. The two big circles touch each other and also the two little circles. With which factor do you have to multiply the radii of the little circles to obtain the radius of the big circles?\nChoices:\n(A) $\\frac{2}{9}$\n(B) $\\sqrt{5}$\n(C) $0.8 \\cdot \\pi$\n(D) 2.5\n(E) $1+\\sqrt{2}$"
            }
        ]
    }
]

chat_response = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=messages,
)
print("Chat response:", chat_response)
Video Input
from openai import OpenAI
# Configured by environment variables
client = OpenAI()

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "video_url",
                "video_url": {
                    "url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/video/N1cdUjctpG8.mp4"
                }
            },
            {
                "type": "text",
                "text": "How many porcelain jars were discovered in the niches located in the primary chamber of the tomb?"
            }
        ]
    }
]

chat_response = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=messages,
)

# When vLLM is launched with `--media-io-kwargs '{"video": {"num_frames": -1}}'`,
# video frame sampling can be configured via `extra_body` (e.g., by setting `fps`).
# This feature is currently supported only in vLLM.
#
# By default, `fps=2` and `do_sample_frames=True`.
# With `do_sample_frames=True`, you can customize the `fps` value to set your desired video sampling rate.
# chat_response = client.chat.completions.create(
#     model="Qwen/Qwen3.8-27B",
#     messages=messages,
#     extra_body={
#         "mm_processor_kwargs": {"fps": 2, "do_sample_frames": True},
#     }, 
# )

print("Chat response:", chat_response)
Instruct (or Non-Thinking) Mode

Qwen3.8-27B will think by default before responding. You can obtain a direct response from the model without thinking by configuring the API parameters. For example,

from openai import OpenAI
# Configured by environment variables
client = OpenAI()

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image_url",
                "image_url": {
                    "url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/RealWorld/RealWorld-04.png"
                }
            },
            {
                "type": "text",
                "text": "Where is this?"
            }
        ]
    }
]

chat_response = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=messages,
    temperature=0.7,
    top_p=0.8,
    presence_penalty=1.5,
    extra_body={
        "top_k": 20,
        "chat_template_kwargs": {"enable_thinking": False},
    }, 
)
print("Chat response:", chat_response)

[!Note] If you are using APIs from Qwen Cloud, in addition to changing model, please use "enable_thinking": False instead of "chat_template_kwargs": {"enable_thinking": False}.

Disable Preserved Thinking

By default, Qwen3.8 retains thinking blocks from all historical messages, maintaining a complete reasoning trace across the conversation. This behavior, known as preserved thinking, ensures full context continuity and is especially beneficial for agent scenarios where decision consistency and reduced redundant reasoning are critical. It also improves KV cache utilization, optimizing inference efficiency in both thinking and non-thinking modes.

If you prefer to retain only the thinking blocks from the latest user message, you can disable this behavior by setting preserve_thinking to False:

from openai import OpenAI

# Configured by environment variables
client = OpenAI()
messages = [...]
chat_response = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=messages,
    extra_body={
        "chat_template_kwargs": {"preserve_thinking": False},
    },
)
print("Chat response:", chat_response)

[!Note] If you are using APIs from Qwen Cloud, in addition to changing model, please use "preserve_thinking": False directly instead of wrapping it in chat_template_kwargs.

Best Practices

To achieve optimal performance, we recommend the following settings:

  1. Sampling Parameters: We suggest using the following sets of sampling parameters:

    • Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
    • Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

    For supported frameworks, you can adjust the presence_penalty parameter between 0 and 2 to reduce endless repetition. However, using a higher value may occasionally result in language mixing and a slight decrease in model performance.

  2. Adequate Output Length: To optimize performance on agentic tasks, we recommend allocating sufficient output length to allow the model to generate detailed and comprehensive responses. For frameworks that support separate token limits for internal reasoning and final outputs, we suggest the following configuration within the 1M context length:

    • Reasoning Content: Set the maximum output length to 262,144 tokens.
    • Final Response: Set the maximum output length to 131,072 tokens.

    These settings provide the necessary capacity for complex reasoning while ensuring ample space for high-quality final deliverables.

  3. Processing Ultra-Long Texts: Qwen3.8-27B natively supports context lengths of up to 262,144 tokens. For long-horizon tasks where the total length (including both input and output) exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively, e.g., YaRN.

    YaRN is currently supported by several inference frameworks, e.g., vLLM, SGLang, and TokenSpeed. In general, there are two approaches to enabling YaRN for supported frameworks:

    • Modifying the model configuration file:

      In the config.json file, change the rope_parameters fields in text_config to: json { "mrope_interleaved": true, "mrope_section": [ 11, 11, 10 ], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144, }

    • Passing command line arguments:

      For vLLM, you can use shell VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1000000

      For SGLang, you can use shell SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... --json-model-override-args '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --context-length 1000000

      For TokenSpeed, you can use shell TOKENSPEED_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 tokenspeed serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1000000

    [!NOTE] All the notable open-source frameworks implement static YaRN, which means the scaling factor remains constant regardless of input length, potentially impacting performance on shorter texts. We advise modifying the rope_parameters configuration only when processing long contexts is required. It is also recommended to modify the factor as needed. For example, if the typical context length for your application is 524,288 tokens, it would be better to set factor as 2.0.

  4. Long Video Understanding: To optimize inference efficiency for plain text and images, the size parameter in the released video_preprocessor_config.json is conservatively configured. It is recommended to set the longest_edge parameter in the video_preprocessor_config file to 469,762,048 (corresponding to 224k video tokens) to enable higher frame-rate sampling for hour-scale videos and thereby achieve superior performance. For example, json {"longest_edge": 469762048, "shortest_edge": 4096}

    Alternatively, override the default values via engine startup parameters. For implementation details, refer to: vLLM / SGLang.

Citation

If you find our work helpful, feel free to give us a cite.

@misc{qwen38,
    title = {{Qwen3.8-Max}: A New Bar for Coding and Cowork},
    url = {https://qwen.ai/blog?id=qwen3.8},
    author = {{Qwen Team}},
    month = {August},
    year = {2026}
}

FIVE DETAILED EXAMPLE GENERATION(S):

Q4KS, non imatrix, standard Qwen settings, NO cache compression of any kind.

NOTE: Some formatting may be lost on copy/paste/export.


---


EXAMPLE #1


---

PROMPT:


Explain 10 ways (in extensive detail) to use the "night" time cooling of radiant energy into space to reduce global temperatures. Include charts, and tables.



THINKING BLOCK:


Thought for 6.05 seconds

The user wants a detailed explanation of 10 ways to leverage nighttime cooling via radiative heat loss (radiant energy escaping into space) to reduce global temperatures. They want charts, tables, and extensive detail.

First, I need to understand the science behind "nighttime cooling via radiative energy into space." This is known as radiative cooling. Objects naturally radiate infrared (IR) heat into space through the atmospheric "window" (roughly 8-13 micrometers wavelength) where the atmosphere is relatively transparent. At night, there's no solar input, so if we can enhance this radiative loss and prevent convective/conductive heat gain, objects can cool below ambient air temperature.

Now, I need to think of 10 distinct ways to apply this for global cooling:

  1. Radiative cooling surfaces for buildings - White, reflective paints/roof coatings with high thermal emittance in the IR window

  2. Daytime radiative cooling materials - Nanostructured surfaces that reflect sunlight while emitting IR

  3. Atmospheric aerosol enhancement of IR emission - Adding particles that emit in the window region

  4. Stratospheric particle injection - Modifying stratospheric composition to enhance IR escape

  5. Sky-facing cooling panels for urban heat islands - Large-scale passive cooling installations

  6. Radiative cooling for power plant waste heat - Using passive cooling to reduce thermal load on atmosphere

  7. Enhanced ocean surface radiative cooling - Modifying ocean surface properties to radiate more heat at night

  8. Forest canopy modifications - Tree species/management to maximize nocturnal radiative cooling

  9. Radiative cooling water harvesting and desalination - Passive systems that cool and provide fresh water simultaneously

  10. Engineering the atmospheric window itself - Reducing greenhouse gases in the 8-13 micron range or adding emitters

Let me organize these into a coherent structure with charts and tables. I'll create:

  • An introduction explaining the physics

  • A table summarizing all 10 approaches

  • Detailed sections for each approach

  • Charts comparing potential impacts

  • A conclusion

For the charts, since this is text-based, I'll use ASCII art or describe the data in tabular form with clear visual structure.

Let me start building this comprehensive answer. I'll make sure each of the 10 approaches is explained in extensive detail, covering the mechanism, potential, challenges, and status.



OUTPUT:


Harnessing Nighttime Radiative Cooling for Global Temperature Reduction

Introduction: The Physics of Radiative Cooling

All objects emit infrared (IR) radiation continuously. On Earth, the atmosphere is largely opaque to IR except for a specific wavelength band known as the atmospheric window (approximately 8–13 μm). Through this window, heat can escape directly to the cold sky (effectively ~3 K background temperature).

At night, when solar input ceases, surfaces can cool below ambient air temperature if:

  1. They have high thermal emittance (ε ≈ 0.9-1.0) in the 8–13 μm window

  2. Convective and conductive heat gains are minimized

  3. They maintain low solar absorptance (α) to prevent daytime heating

Key Principle: ΔT_cooling = f(ε_IR, α_solar, sky temperature, air temperature, wind speed)

Theoretical maximum cooling: 15–20°C below ambient under clear, calm, dry night conditions.


Overview Table: 10 Radiative Cooling Strategies

#

Strategy

Mechanism

Estimated Cooling Potential

Maturity Level

Implementation Scale

1

Radiative Cooling Building Surfaces

High-ε coatings on roofs/walls

1.0–4.0°C (urban)

Commercial

City/Regional

2

Daytime Passive Radiative Cooling Materials

Nanostructured selective surfaces

3–5°C below ambient (day)

Pilot/Early Commercial

Building/Industrial

3

Enhanced Atmospheric IR Emission Aerosols

IR-emitting particles in lower atmosphere

0.5–2.0°C (global)

Research

Global

4

Stratospheric Radiative Modification

Particles enhancing IR escape at high altitude

1.0–3.0°C (global)

Research

Global

5

Urban Radiative Cooling Parks/Plazas

Large passive cooling installations

1–3°C (local)

Concept

City

6

Radiative Cooling for Industrial Waste Heat

Passive IR cooling of process heat

0.2–0.5°C (global, via efficiency)

Early Stage

Industrial

7

Ocean Surface Radiative Enhancement

Floating IR-emitting structures

0.3–1.0°C (global)

Research

Ocean-wide

8

Forest Canopy Radiative Optimization

Species selection & management

0.5–1.5°C (regional)

Early Stage

Continental

9

Radiative Cooling Water Systems

Night cooling for water heating/storage

0.1–0.3°C (global energy savings)

Commercial

Urban/Rural

10

Atmospheric Window Engineering

Reducing absorbers in 8-13 μm band

1.0–5.0°C (global)

Research

Global


1. Radiative Cooling Building Surfaces

Mechanism

Apply highly emissive paints, films, or coatings to building roofs and walls that:

  • Reflect 85–95% of solar radiation (low α_solar)

  • Emit 90–95% of absorbed heat in the 8–13 μm window (high ε_IR)

  • Minimize conductive heat transfer from interior

Detailed Implementation

Material Composition:

  • Base: Titanium dioxide (TiO₂) nanoparticles for solar reflection

  • Binder: Fluoropolymer or acrylic matrix

  • IR-emitting component: SiO₂ or MgF₂ microspheres

  • Topcoat: Hydrophobic layer for self-cleaning

Application Protocol:

  1. Clean and prime surface (remove oxidation, dust)

  2. Apply 2-3 coats (total 200-300 μm thickness)

  3. Apply hydrophobic topcoat

  4. Maintain with annual inspection

Performance Characteristics

Parameter

Value

Solar reflectance (α)

0.85–0.95

IR emittance (ε)

0.90–0.95

Night cooling below ambient

3–8°C

Day cooling below ambient

0–3°C (depending on climate)

Service life

10–20 years

Cost per m²

$5–$25

Global Impact Analysis

Temperature Reduction Potential (Urban Areas)

Region         Current Avg Temp    With RC Coatings    Reduction
-------------  ----------------   -----------------   ---------
Mumbai, India  31.5°C              29.0°C              -2.5°C
Chicago, USA   20.0°C              18.5°C              -1.5°C
Tokyo, Japan   21.0°C              19.5°C              -1.5°C
São Paulo, BZ  22.0°C              20.5°C              -1.5°C
Global urban   24.3°C              22.8°C              -1.5°C

Energy Savings:

  • HVAC load reduction: 15–30%

  • Peak electricity demand reduction: 10–20%

  • CO₂ savings: ~0.5–1.0 tons/m² over building lifetime

Challenges & Solutions

Challenge

Solution

Daytime heating in sunny climates

Use highly reflective coatings (α < 0.10)

Cost of materials

Economies of scale; government incentives

Maintenance (dust, dirt)

Hydrophobic topcoat; periodic cleaning

Aesthetics

Offer color variants using IR-reflective pigments

Current Status

  • Commercial products available (e.g., CoolRoof, Tyvek Cool Roof)

  • Building codes in some regions (California Title 24)

  • Estimated potential: 0.5–1.0°C global reduction if applied to 50% of urban roofs


2. Daytime Passive Radiative Cooling Materials

Mechanism

Engineered nanostructures that simultaneously:

  • Reflect nearly all solar radiation (0.3–2.5 μm)

  • Emit thermal radiation selectively in the atmospheric window (8–13 μm)

  • Achieve cooling below ambient even under direct sunlight

Material Design

Multilayer Stack Architecture:

Layer 1: Top SiO₂ layer (100 nm) - IR transparency
Layer 2: TiO₂ nanoparticles (1 μm) - Solar reflection
Layer 3: PDMS matrix - Mechanical support
Layer 4: SiO₂ bottom layer (4 μm) - IR emission
Substrate: PET or aluminum foil

Alternative: Metamaterial Approach

  • Periodic Si or SiO₂ nanostructures

  • Resonant features tuned to 8–13 μm emission

  • Photonic crystal design for broadband solar reflection

Performance Characteristics

Parameter

Conventional White Paint

Advanced RC Material

Solar reflectance

85–90%

95–99%

IR emittance

90%

95–99%

Night cooling below ambient

3–5°C

5–8°C

Day cooling below ambient

0–2°C

2–5°C

Power density (cooling)

50–80 W/m²

80–150 W/m²

Cooling Power vs. Solar Irradiance

Cooling Power (W/m²)
150 |                                    *
    |                                   *
140 |                                  *
    |                                 *
130 |                                *
    |                               *
120 |                              *
    |                             *
110 |                            *
    |                           *
100 |                          *
    |                         *
 90 |                        *
    |                       *
 80 |                      *
    |                     *
 70 |                    *
    |                   *
 60 |                  *
    |                 *
 50 |                *
    |               *
 40 |              *
    |             *
 30 |            *
    |           *
 20 |          *
    |         *

Identity and Version

Repository
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF
Publisher
David Belton
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
en, zh
Revision
c02caef111a8acf987947f35e1e288aa5450e184
First published
2026-09-01
Last updated
2026-09-16

Files and Weights

27 files, 389.0 GB in total. The weights are 23 files totalling 389.0 GB in gguf.

Weights23 files · 389.0 GB
Documentation1 file · 392.1 KB
Other2 files · 3.6 MB
Repository1 file · 4.1 KB
Every file
FileTypeSizeSHA-256
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ2_M.ggufWeights11.7 GB 8629184db672
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ3_M.ggufWeights14.1 GB 856d3ef6d662
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ4_XS.ggufWeights16.6 GB b7f7f75df9fa
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-LOW-MTP-IQ4_XS.ggufWeights15.3 GB fa92183638b0
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-LOW-MTP-Q6_K.ggufWeights22.4 GB 2e15fdaad887
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-IQ2_M.ggufWeights12.1 GB ee4fc4950338
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-IQ3_M.ggufWeights14.5 GB 19ee0a94937a
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-IQ4_XS.ggufWeights17.0 GB 318d14f81f30
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q4_K_M.ggufWeights18.5 GB bc7a6cf2bcc7
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q4_K_S.ggufWeights17.5 GB 889caf9975ef
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q5_K_M.ggufWeights21.2 GB 7408f59414a4
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q5_K_S.ggufWeights20.6 GB 89d306925a88
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q6_K.ggufWeights24.0 GB ac011aabe685
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q8_0.ggufWeights30.2 GB 54f27515edb2
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_M.ggufWeights18.0 GB cf35a030cca2
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_S.ggufWeights17.1 GB 208ec01cadf8
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q5_K_M.ggufWeights20.7 GB 4132d28c31a7
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q5_K_S.ggufWeights20.2 GB d9baa706d262
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q6_K.ggufWeights23.6 GB f814af8d2ce6
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q8_0.ggufWeights29.8 GB 0c87ce4c1ef7
mmproj-BF16.ggufWeights931.1 MB b0d8d89e9c9c
mmproj-F16.ggufWeights927.6 MB 82e620db8cb8
mmproj-F32.ggufWeights1.8 GB 815ee690ba80
README.mdDocumentation392.1 KB
qwen38-27b-turbo-tfcf735.pngOther61.5 KB
star-wars-hans-solo.gifOther3.5 MB b977bb1e9753
.gitattributesRepository4.1 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
389.0 GB
Download from David Belton

Released by David Belton through its official repository on Hugging Face. Read the license.

Built From

  • Derived from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
  • Quantized from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
  • Trained on (disclosed) DavidAU/F451-STRICT-Datasets
  • Trained on (disclosed) DavidAU/Polar-STRICT-Datasets

Memory Requirements

PrecisionWeights in memory
As published389.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF

Can I use Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF commercially?

Yes. Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-GSQ-RCO series come from ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 23 to 51 have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. The size after conversion may differ from the original GGUF. The newly added Huihui-Qwen3.8-27B-abliterated-UD series come from unsloth/Qwen3.8-27B-GGUF. Only layers 18 to 51 have been ablated(Previously…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

Model · Image and text to text

Gemma-4-E4B-Uncensored-HauhauCS-Aggressive

HauhauCS

Gemma 4 E4B-IT uncensored by HauhauCS. 0/465 Refusals\ No changes to datasets or capabilities. Fully functional, 100% of what the original authors intended - just without the refusals. These are meant to be the best lossless uncensored models out there. Stronger uncensoring — model is fully unlocked and won't refuse prompts. May occasionally append short disclaimers (baked into base model training, not refusals) but full content is always generated. For a more conservative uncensor that keeps some safety guardrails, check the Balanced variant when it's available. All quants generated with importance matrix (imatrix) for optimal quality preservation on abliterated weights. KP ("Perfect")…

Open weights gemma

Model · Image and text to text

Qwen3.5-9B-GGUF

Unsloth AI

You can now also fine-tune the model locally with Unsloth. - Read our Qwen3.5 fine-tuning guide here. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty…

Open weights apache-2.0 transformers

Model · Image and text to text

Qwen3.8-Flash-Next-GGUF

Unsloth AI

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other