SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3-VL-4B-Instruct-LoRA-AX650

by AXERA AXERA-TECH/Qwen3-VL-4B-Instruct-LoRA-AX650

Qwen3-VL-4B-Instruct-LoRA-AX650 is an open-weight model for image and text to text from AXERA, released under Apache License 2.0. Its published files total 6.9 GB. It draws 9 downloads a month.

Ready-to-run package for Qwen/Qwen3-VL-4B-Instruct on AX650 / NPU3. It includes an AX650 aarch64 axllm server, 36 compiled text layers, a fixed-shape image encoder, and two runtime-selectable LoRA adapters.

Parameters—
Context—
Weights910.0 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads9

Model Card

By AXERA, published under apache-2.0, revision e68e2e360d6d.

Ready-to-run package for Qwen/Qwen3-VL-4B-Instruct on AX650 / NPU3. It includes an AX650 aarch64 axllm server, 36 compiled text layers, a fixed-shape image encoder, and two runtime-selectable LoRA adapters. This release supports text chat and single-image requests through the OpenAI-compatible chat API. - AX650 / AX650N aarch64; the results below were measured on an AX650 / NPU3 board. The image encoder contributes 144 visual soft tokens per image: (384 / 16 / 2)² = 144, using the packaged 16-pixel patch size and spatial merge size of 2. The total input-token count also includes the text prompt and chat-format tokens. Keep the full request, including visual tokens, within the prefill and…

Read AXERA's full model card

Qwen3-VL-4B-Instruct LoRA on AXERA AX650

Ready-to-run package for Qwen/Qwen3-VL-4B-Instruct on AX650 / NPU3. It includes an AX650 aarch64 axllm server, 36 compiled text layers, a fixed-shape image encoder, and two runtime-selectable LoRA adapters. This release supports text chat and single-image requests through the OpenAI-compatible chat API.

Supported Platform and Configuration

  • AX650 / AX650N aarch64; the results below were measured on an AX650 / NPU3 board.
  • Text prefill limit: 1,536 input tokens, in 128-token chunks. The compiled context is 2,048 positions; the runtime reports max_token_len: 2047. Leave room for the generated response.
  • Packaged image profile: 384 × 384. The runtime resizes input images to this fixed profile.
  • Maximum server concurrency: one request. Send requests serially, especially when changing adapters.

The image encoder contributes 144 visual soft tokens per image: (384 / 16 / 2)² = 144, using the packaged 16-pixel patch size and spatial merge size of 2. The total input-token count also includes the text prompt and chat-format tokens. Keep the full request, including visual tokens, within the prefill and context limits.

Performance

These are single-request measurements on AX650 / NPU3 with requests sent serially. TTFT means time to first token; the image row includes image preparation and encoding.

Request Input tokens Prefill chunks TTFT
Short text, ChartQA adapter 34 1 1,031 ms
Medium text, ChartQA adapter 575 5 6,918 ms
Long text, ChartQA adapter 995 8 14,246 ms
First text request after switching to Design 32 1 2,189 ms
Packaged chart image, ChartQA adapter 173 2 2,828 ms

The 575- and 995-token prompts used a repeated word list and asked for a one-word response; both returned the requested word. The first request after an adapter change includes adapter loading and rebinding. The image request crosses the 128-token first prefill chunk because its 144 visual soft tokens are added to the prompt.

Startup Runtime Footprint

Measured on AX650 / NPU3 with this package.

Item Value
Package size, excluding Git history 6.44 GiB
CMM used by the server (incremental) approximately 5.93 GiB

The package includes 39 AXModels, embedding weights, both LoRA adapters, and runtime files. CMM use may vary with the runtime environment.

Image Profile and Accuracy

Qwen3-VL-4B-Instruct_vision.axmodel is the image encoder selected by config.json. The package also contains Qwen3-VL-4B-Instruct_vision_u8.axmodel; the startup script does not select that file. The fixed 384 × 384 profile limits fine-detail chart reading.

At this fixed resolution, chart reading is unreliable: this build answered 2 of 24 questions correctly in a ChartQA check. Check image-based answers before relying on them.

LoRA Adapters

Set task_id on each chat request:

task_id Intended task
qwen3-vl-lora-chartqa Chart questions
qwen3-vl-lora-design Design assistance

The package stores both adapters as BF16 matrix-input payloads. Adapter selection is process-global, so send requests serially when switching tasks.

These examples demonstrate dynamic LoRA loading and switching; their task accuracy may be insufficient for practical use.

Package Layout

.
├── README.md
├── bin/axllm
├── start_axllm.sh
├── axllm.version.json
├── config.json
├── post_config.json
├── qwen3_tokenizer.txt
├── model.embed_tokens.weight.bfloat16.bin
├── qwen3_vl_text_p128_l0_together.axmodel ... qwen3_vl_text_p128_l35_together.axmodel
├── qwen3_vl_text_post.axmodel
├── Qwen3-VL-4B-Instruct_vision.axmodel
├── Qwen3-VL-4B-Instruct_vision_u8.axmodel
├── lora/
│   ├── qwen3-vl-lora-chartqa/   (layer_00.bf16.bin ... layer_35.bf16.bin, manifest)
│   └── qwen3-vl-lora-design/    (layer_00.bf16.bin ... layer_35.bf16.bin, manifest)
├── assets/chartqa_00.png
└── runtime/lib/libax_engine.so

The original Hugging Face weights are not included. The startup script resolves all runtime files relative to this package.

Download and Run

Download the package with the Hugging Face CLI on a network-connected machine:

mkdir -p Qwen3-VL-4B-Instruct-LoRA-AX650
cd Qwen3-VL-4B-Instruct-LoRA-AX650
hf download AXERA-TECH/Qwen3-VL-4B-Instruct-LoRA-AX650 --local-dir .

Transfer this directory to an AX650 board if it was downloaded elsewhere. From the package directory on the board, start the bundled server:

chmod +x ./bin/axllm ./start_axllm.sh
./start_axllm.sh 8000

start_axllm.sh sets the bundled runtime library path and runs bin/axllm serve. In another terminal on the board, check the server:

curl -fsS http://127.0.0.1:8000/health
curl -fsS http://127.0.0.1:8000/v1/models

A healthy server reports "status": "healthy" and "max_concurrency": 1. The model list reported this ID:

AXERA-TECH/Qwen3-VL-4B-Instruct-LoRA-VLM-AX650-P1536-C2048-Chunk128

Use that exact ID in the model field of every request.

Text Request

This request selects the Design adapter:

curl -sS http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"AXERA-TECH/Qwen3-VL-4B-Instruct-LoRA-VLM-AX650-P1536-C2048-Chunk128","task_id":"qwen3-vl-lora-design","messages":[{"role":"user","content":"Name three colors that work well with navy blue. Be brief."}],"max_tokens":64,"temperature":0}'

Example response content (choices[0].message.content): "White, beige, gold".

Image Request

The packaged sample is a bar chart at assets/chartqa_00.png:

From the package directory on the board, send it as a base64 data URL. The base64 -w 0 option is available on the board's GNU coreutils.

IMAGE_DATA=$(base64 -w 0 assets/chartqa_00.png)
curl -sS http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  --data-binary @- <<JSON
{"model":"AXERA-TECH/Qwen3-VL-4B-Instruct-LoRA-VLM-AX650-P1536-C2048-Chunk128","task_id":"qwen3-vl-lora-chartqa","messages":[{"role":"user","content":[{"type":"text","text":"What is the index value of Coffee?"},{"type":"image_url","image_url":{"url":"data:image/png;base64,$IMAGE_DATA"}}]}],"max_tokens":32,"temperature":0}
JSON

For this sample, the package returned "1" in choices[0].message.content, although the chart shows 82.2 for Coffee. This illustrates the image accuracy limitation above.

Conversion References

If you need the original model files or want to rebuild the deployment artifacts, start with:

Discussion

Identity and Version

Repository
AXERA-TECH/Qwen3-VL-4B-Instruct-LoRA-AX650
Publisher
AXERA
Task
Image and text to text
Modality
Image and text
Library
axllm
Parameters
Not stated by the source
Languages
en, zh
Revision
e68e2e360d6ddcd3de7c8c9dde6040f6d1885679
First published
2026-09-18
Last updated
2026-09-20

Files and Weights

126 files, 6.9 GB in total. The weights are 73 files totalling 910.0 MB in bin.

Weights73 files · 910.0 MB
Configuration7 files · 17.8 KB
Tokenizer1 file · 1.6 MB
Documentation1 file · 7.4 KB
Other43 files · 6.0 GB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
lora/qwen3-vl-lora-chartqa/layer_00.bf16.binWeights1.8 MB 22ce21105681
lora/qwen3-vl-lora-chartqa/layer_01.bf16.binWeights1.8 MB 01970e9efd06
lora/qwen3-vl-lora-chartqa/layer_02.bf16.binWeights1.8 MB dc404af7b0b7
lora/qwen3-vl-lora-chartqa/layer_03.bf16.binWeights1.8 MB 363345d89cae
lora/qwen3-vl-lora-chartqa/layer_04.bf16.binWeights1.8 MB 293f36a24333
lora/qwen3-vl-lora-chartqa/layer_05.bf16.binWeights1.8 MB 7ef006f36178
lora/qwen3-vl-lora-chartqa/layer_06.bf16.binWeights1.8 MB e48a5c9e727c
lora/qwen3-vl-lora-chartqa/layer_07.bf16.binWeights1.8 MB 17f4f51bd17f
lora/qwen3-vl-lora-chartqa/layer_08.bf16.binWeights1.8 MB 00bdc66babf6
lora/qwen3-vl-lora-chartqa/layer_09.bf16.binWeights1.8 MB e7a7a6f71cc9
lora/qwen3-vl-lora-chartqa/layer_10.bf16.binWeights1.8 MB c54ffb674246
lora/qwen3-vl-lora-chartqa/layer_11.bf16.binWeights1.8 MB 60aa1f6bb9b4
lora/qwen3-vl-lora-chartqa/layer_12.bf16.binWeights1.8 MB edaba6ae1df4
lora/qwen3-vl-lora-chartqa/layer_13.bf16.binWeights1.8 MB cbcd35e788a7
lora/qwen3-vl-lora-chartqa/layer_14.bf16.binWeights1.8 MB 1df2ccbebb15
lora/qwen3-vl-lora-chartqa/layer_15.bf16.binWeights1.8 MB a57fd0c8de33
lora/qwen3-vl-lora-chartqa/layer_16.bf16.binWeights1.8 MB bea8c2a72280
lora/qwen3-vl-lora-chartqa/layer_17.bf16.binWeights1.8 MB da390b2be99d
lora/qwen3-vl-lora-chartqa/layer_18.bf16.binWeights1.8 MB 2ea6ba625d2f
lora/qwen3-vl-lora-chartqa/layer_19.bf16.binWeights1.8 MB df3cf57f8fba
lora/qwen3-vl-lora-chartqa/layer_20.bf16.binWeights1.8 MB a37adfd9ddf7
lora/qwen3-vl-lora-chartqa/layer_21.bf16.binWeights1.8 MB acfde32244aa
lora/qwen3-vl-lora-chartqa/layer_22.bf16.binWeights1.8 MB 9b3d59e4eed4
lora/qwen3-vl-lora-chartqa/layer_23.bf16.binWeights1.8 MB 180d655742a4
lora/qwen3-vl-lora-chartqa/layer_24.bf16.binWeights1.8 MB 4951c0152bef
lora/qwen3-vl-lora-chartqa/layer_25.bf16.binWeights1.8 MB 0a2c8f202d3e
lora/qwen3-vl-lora-chartqa/layer_26.bf16.binWeights1.8 MB f2c2346c4eb1
lora/qwen3-vl-lora-chartqa/layer_27.bf16.binWeights1.8 MB 28591883b439
lora/qwen3-vl-lora-chartqa/layer_28.bf16.binWeights1.8 MB aced9996a382
lora/qwen3-vl-lora-chartqa/layer_29.bf16.binWeights1.8 MB e841adf39348
lora/qwen3-vl-lora-chartqa/layer_30.bf16.binWeights1.8 MB bd6a59c34e14
lora/qwen3-vl-lora-chartqa/layer_31.bf16.binWeights1.8 MB 235876f7f7da
lora/qwen3-vl-lora-chartqa/layer_32.bf16.binWeights1.8 MB ae5bcf48ebd5
lora/qwen3-vl-lora-chartqa/layer_33.bf16.binWeights1.8 MB d4a48ba6392f
lora/qwen3-vl-lora-chartqa/layer_34.bf16.binWeights1.8 MB 8b9fd715e1fb
lora/qwen3-vl-lora-chartqa/layer_35.bf16.binWeights1.8 MB 33bea406b8b8
lora/qwen3-vl-lora-design/layer_00.bf16.binWeights1.8 MB b88faed4be10
lora/qwen3-vl-lora-design/layer_01.bf16.binWeights1.8 MB 483048010338
lora/qwen3-vl-lora-design/layer_02.bf16.binWeights1.8 MB efd9be2bf013
lora/qwen3-vl-lora-design/layer_03.bf16.binWeights1.8 MB c9c3a560a799
lora/qwen3-vl-lora-design/layer_04.bf16.binWeights1.8 MB ceb48b5427ca
lora/qwen3-vl-lora-design/layer_05.bf16.binWeights1.8 MB bf806276050e
lora/qwen3-vl-lora-design/layer_06.bf16.binWeights1.8 MB e78ebe7add2b
lora/qwen3-vl-lora-design/layer_07.bf16.binWeights1.8 MB cc5a4cdb1376
lora/qwen3-vl-lora-design/layer_08.bf16.binWeights1.8 MB 7f3bd65fe49e
lora/qwen3-vl-lora-design/layer_09.bf16.binWeights1.8 MB 413aad1c0c93
lora/qwen3-vl-lora-design/layer_10.bf16.binWeights1.8 MB fda9cc9b0df1
lora/qwen3-vl-lora-design/layer_11.bf16.binWeights1.8 MB 4a6886aa0e54
lora/qwen3-vl-lora-design/layer_12.bf16.binWeights1.8 MB b90384ee3e8e
lora/qwen3-vl-lora-design/layer_13.bf16.binWeights1.8 MB fb5774bd10cf
lora/qwen3-vl-lora-design/layer_14.bf16.binWeights1.8 MB 5730493eee8c
lora/qwen3-vl-lora-design/layer_15.bf16.binWeights1.8 MB 10280b71c0f6
lora/qwen3-vl-lora-design/layer_16.bf16.binWeights1.8 MB 53df9a94fd01
lora/qwen3-vl-lora-design/layer_17.bf16.binWeights1.8 MB 4eb4b0c9f5eb
lora/qwen3-vl-lora-design/layer_18.bf16.binWeights1.8 MB a837cff06dd4
lora/qwen3-vl-lora-design/layer_19.bf16.binWeights1.8 MB 3ed07cb2e467
lora/qwen3-vl-lora-design/layer_20.bf16.binWeights1.8 MB 8d196fa43dd0
lora/qwen3-vl-lora-design/layer_21.bf16.binWeights1.8 MB 7c9db8a8de8f
lora/qwen3-vl-lora-design/layer_22.bf16.binWeights1.8 MB 8e6a8a2582b7
lora/qwen3-vl-lora-design/layer_23.bf16.binWeights1.8 MB f72f7d277529
lora/qwen3-vl-lora-design/layer_24.bf16.binWeights1.8 MB 18e7d8e95706
lora/qwen3-vl-lora-design/layer_25.bf16.binWeights1.8 MB 53975b6d6f1b
lora/qwen3-vl-lora-design/layer_26.bf16.binWeights1.8 MB 5f3edce35d9c
lora/qwen3-vl-lora-design/layer_27.bf16.binWeights1.8 MB 836944ee39e4
lora/qwen3-vl-lora-design/layer_28.bf16.binWeights1.8 MB 22d2bef9b1c4
lora/qwen3-vl-lora-design/layer_29.bf16.binWeights1.8 MB c88ad3212fb2
lora/qwen3-vl-lora-design/layer_30.bf16.binWeights1.8 MB cc055b817208
lora/qwen3-vl-lora-design/layer_31.bf16.binWeights1.8 MB bc92e6f27c0e
lora/qwen3-vl-lora-design/layer_32.bf16.binWeights1.8 MB cac84162543b
lora/qwen3-vl-lora-design/layer_33.bf16.binWeights1.8 MB 914b25a3e551
lora/qwen3-vl-lora-design/layer_34.bf16.binWeights1.8 MB 64cc337bfaed
lora/qwen3-vl-lora-design/layer_35.bf16.binWeights1.8 MB 78c1ed08a163
model.embed_tokens.weight.bfloat16.binWeights777.9 MB 2942e869a0df
axllm.version.jsonConfiguration352 B —
config.jsonConfiguration1.5 KB —
lora/qwen3-vl-lora-chartqa/manifest.jsonConfiguration6.7 KB —
lora/qwen3-vl-lora-chartqa/source_adapter_config.jsonConfiguration1.1 KB —
lora/qwen3-vl-lora-design/manifest.jsonConfiguration6.7 KB —
lora/qwen3-vl-lora-design/source_adapter_config.jsonConfiguration1.2 KB —
post_config.jsonConfiguration279 B —
README.mdDocumentation7.4 KB —
Qwen3-VL-4B-Instruct_vision.axmodelOther461.4 MB ec5e5ce70571
Qwen3-VL-4B-Instruct_vision_u8.axmodelOther441.9 MB 0318af7d86d0
assets/chartqa_00.pngOther44.9 KB —
bin/axllmOther2.5 MB 7b1876e0308d
qwen3_vl_text_p128_l0_together.axmodelOther129.8 MB 5011e7fc91e3
qwen3_vl_text_p128_l10_together.axmodelOther129.8 MB 37caf08b2c6f
qwen3_vl_text_p128_l11_together.axmodelOther129.8 MB a5a98bd38666
qwen3_vl_text_p128_l12_together.axmodelOther129.8 MB 02353d0e645b
qwen3_vl_text_p128_l13_together.axmodelOther129.8 MB c9a271486d85
qwen3_vl_text_p128_l14_together.axmodelOther129.8 MB ade15f8bf840
qwen3_vl_text_p128_l15_together.axmodelOther129.8 MB 36f362e14751
qwen3_vl_text_p128_l16_together.axmodelOther129.8 MB fdaf1d31d7c7
qwen3_vl_text_p128_l17_together.axmodelOther129.8 MB 4f154db24fb0
qwen3_vl_text_p128_l18_together.axmodelOther129.8 MB 244a384d60cd
qwen3_vl_text_p128_l19_together.axmodelOther129.8 MB 6d5d9c3a530d
qwen3_vl_text_p128_l1_together.axmodelOther129.8 MB fa183ce6e1f1
qwen3_vl_text_p128_l20_together.axmodelOther129.8 MB 7cdbe19f91f5
qwen3_vl_text_p128_l21_together.axmodelOther129.8 MB 7d508cb727bc
qwen3_vl_text_p128_l22_together.axmodelOther129.8 MB 7d0ad3121fff
qwen3_vl_text_p128_l23_together.axmodelOther129.8 MB 00c48fdf1671
qwen3_vl_text_p128_l24_together.axmodelOther129.8 MB 1c3065ed393c
qwen3_vl_text_p128_l25_together.axmodelOther129.8 MB 431cab734959
qwen3_vl_text_p128_l26_together.axmodelOther129.8 MB 893eba59174f
qwen3_vl_text_p128_l27_together.axmodelOther129.8 MB 390de80cc83c
qwen3_vl_text_p128_l28_together.axmodelOther129.8 MB 92bbb67e5eee
qwen3_vl_text_p128_l29_together.axmodelOther129.8 MB 1c0b2d2bf84b
qwen3_vl_text_p128_l2_together.axmodelOther129.8 MB 9581896e64ba
qwen3_vl_text_p128_l30_together.axmodelOther129.8 MB 8f9531401334
qwen3_vl_text_p128_l31_together.axmodelOther129.8 MB 5753d99ec0df
qwen3_vl_text_p128_l32_together.axmodelOther129.8 MB 8a1eba161478
qwen3_vl_text_p128_l33_together.axmodelOther129.8 MB d00fc79a98a5
qwen3_vl_text_p128_l34_together.axmodelOther129.8 MB f42635e4e0b5
qwen3_vl_text_p128_l35_together.axmodelOther129.8 MB 82f868c9146d
qwen3_vl_text_p128_l3_together.axmodelOther129.8 MB d4d21974476f
qwen3_vl_text_p128_l4_together.axmodelOther129.8 MB 6e22550fc5dc
qwen3_vl_text_p128_l5_together.axmodelOther129.8 MB 440fa899cf0d
qwen3_vl_text_p128_l6_together.axmodelOther129.8 MB 6ade2e87ae24
qwen3_vl_text_p128_l7_together.axmodelOther129.8 MB 57e1024983fb
qwen3_vl_text_p128_l8_together.axmodelOther129.8 MB 7fd22671077b
qwen3_vl_text_p128_l9_together.axmodelOther129.8 MB 2eb4d41398e3
qwen3_vl_text_post.axmodelOther424.0 MB 8dde0ac533b9
runtime/lib/libax_engine.soOther1.9 MB 1f519027ef5e
start_axllm.shOther270 B —
.gitattributesRepository1.7 KB —
qwen3_tokenizer.txtTokenizer1.6 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
910.0 MB
Download from AXERA

Released by AXERA through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published910.0 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3-VL-4B-Instruct-LoRA-AX650

Can I use Qwen3-VL-4B-Instruct-LoRA-AX650 commercially?

Yes. Qwen3-VL-4B-Instruct-LoRA-AX650 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. The size after conversion may differ from the original GGUF (Some of the weights are converted from PTQ1 to Q2K or Q3K.). This is just a test/validation. The ternary hybrid-attention kernels live in the PrismML-Eng/llama.cpp fork.…

Open weights apache-2.0 transformers

and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads) instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model ctrl at the chat/message level). Model name has "plusIQ" in the name. (there is also a extra robust "tools" version too.) A 12+12 (12 reasoning and 12 instruct) model with interactive optimization/help system will be releasing shortly too. Extreme intelligence in a small package. Jaw dropping performance. Superior instruction following. A multi-stage and multi-model…

Open weights apache-2.0

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…

Open weights apache-2.0

Non-uniform GGUF quantizations produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-27B at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. Both methods were developed at the…

Open weights apache-2.0 gguf