SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.8-27B-nvfp4-NInfer

by Neroued neroued/Qwen3.8-27B-nvfp4-NInfer

Qwen3.8-27B-nvfp4-NInfer is an open-weight model for image and text to text from Neroued, released under Apache License 2.0. Its published files total 23.7 GB. It draws 366.3k downloads a month.

This model card is the version-controlled source for The repository contains a mixed NVFP4/FP8 representation of Qwen3.8-27B.

Parameters—
Context—
Weights23.7 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads366.3k

Model Card

By Neroued, published under apache-2.0, revision a107ba1b5b06.

Qwen3.8-27B NVFP4 for NInfer

This model card is the version-controlled source for neroued/Qwen3.8-27B-nvfp4-NInfer.

The repository contains a mixed NVFP4/FP8 representation of Qwen3.8-27B. It combines the official BF16 checkpoint with the fixed packed Text weights from unsloth/Qwen3.8-27B-NVFP4 in the native NInfer .ninfer artifact format. The artifact is intended only for NInfer; it is not a Transformers checkpoint, Safetensors distribution, or GGUF file.

The artifact uses the Qwen3.5 Dense architecture. Text layers 0–55 use NVFP4 MLP weights, while the token embedding, attention input/output projections, GDN Q/K/V/Z and output projections, full output head, and Text layers 56–63 MLP weights use row-scaled FP8. Control weights use BF16, with separate MTP, Vision and DFlash2 weights.

Artifact

Read the full model card (1,437 words)

Identity and Version

Repository
neroued/Qwen3.8-27B-nvfp4-NInfer
Publisher
Neroued
Task
Image and text to text
Modality
Image and text
Library
ninfer
Parameters
Not stated by the source
Languages
rtx-5090
Revision
a107ba1b5b0609d9ca90d5ed61439f5f5cc7d64d
First published
2026-08-06
Last updated
2026-09-29

Files and Weights

6 files, 23.7 GB in total.

Configuration1 file · 2.4 KB
Documentation2 files · 26.5 KB
Other2 files · 23.7 GB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
artifact-manifest.jsonConfiguration2.4 KB —
LICENSEDocumentation11.4 KB —
README.mdDocumentation15.1 KB —
SHA256SUMSOther91 B —
qwen3_8_27b_nvfp4.ninferOther23.7 GB 74d2c57145e6
.gitattributesRepository1.6 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download from Neroued

Released by Neroued through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
AIME 2025 Task Text GenerationMetric Accuracy (0-shot, rule)Comparison conditions not established 96.67 neroued
Publisher reported
Evaluated revision not stated —
AIME 2026 Task Text GenerationMetric Accuracy (0-shot, rule)Comparison conditions not established 96.67 neroued
Publisher reported
Evaluated revision not stated —
ERQA Task Image Text to TextMetric Accuracy (0-shot, rule)Comparison conditions not established 66.25 neroued
Publisher reported
Evaluated revision not stated —
GPQA-Diamond Task Text GenerationMetric Accuracy (0-shot, rule)Comparison conditions not established 90.4 neroued
Publisher reported
Evaluated revision not stated —
IFBench Task Text GenerationMetric Prompt-level strict (0-shot, rule)Comparison conditions not established 77 neroued
Publisher reported
Evaluated revision not stated —
RealWorldQA Task Image Text to TextMetric Accuracy (0-shot, rule)Comparison conditions not established 83.53 neroued
Publisher reported
Evaluated revision not stated —

Questions About Qwen3.8-27B-nvfp4-NInfer

Can I use Qwen3.8-27B-nvfp4-NInfer commercially?

Yes. Qwen3.8-27B-nvfp4-NInfer is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Non-uniform GGUF quantizations of a 512-expert MoE, produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-Flash-Next at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. A…

Open weights apache-2.0 gguf

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads) instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model ctrl at the chat/message level). Model name has "plusIQ" in the name. A 12+12 (12 reasoning and 12 instruct) model with interactive optimization/help system will be releasing shortly too. BF16/16-bit MTP GGUF also avail. (there is also a extra robust "tools" version too - Q6 and Q8.) Extreme intelligence in a small package. Jaw dropping performance. Superior…

Open weights apache-2.0