SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

RK182X-VLM-Qwen3-VL-2B-Instruct

by RKNNAI RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct

RK182X-VLM-Qwen3-VL-2B-Instruct is an open-weight model for image and text to text from RKNNAI, released under Apache License 2.0. Its published files total 4.0 GB.

本仓库提供由 Qwen/Qwen3-VL-2B-Instruct 转换得到的 RKNN VLM 模型。 - Model ID:RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct - 模型显示名称:RK182X-VLM-Qwen3-VL-2B-Instruct - 源模型:Qwen/Qwen3-VL-2B-Instruct - 模型类型:VLM ModelScope 完整下载: Hugging Face 完整下载: ModelScope 指定配置下载: Hugging Face…

Parameters—
Context—
Weights1.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By RKNNAI, published under apache-2.0, revision e5e82e8ea6e9.

本仓库提供由 Qwen/Qwen3-VL-2B-Instruct 转换得到的 RKNN VLM 模型。 - Model ID:RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct - 模型显示名称:RK182X-VLM-Qwen3-VL-2B-Instruct - 源模型:Qwen/Qwen3-VL-2B-Instruct - 模型类型:VLM ModelScope 完整下载: Hugging Face 完整下载: ModelScope 指定配置下载: Hugging Face 指定配置下载: - 使用配套 RKNN Runtime 和驱动;运行前用 rknn-smi -v 检查设备端版本。 源模型许可证见根目录 LICENSE;同时遵守 RKNN Toolkit 和 RKNN Runtime 许可条款。

Read RKNNAI's full model card

1. 模型介绍

本仓库提供由 Qwen/Qwen3-VL-2B-Instruct 转换得到的 RKNN VLM 模型。

  • Model ID:RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct
  • 模型显示名称:RK182X-VLM-Qwen3-VL-2B-Instruct
  • 发布版本:v1.1.0
  • 源模型:Qwen/Qwen3-VL-2B-Instruct
  • 模型类型:VLM
  • 芯片字段:RK182X
  • 具体支持芯片:RK1828

可用模型

发布版本 配置目录 支持芯片 分辨率 量化方式 NPU 核数 上下文长度(tokens) KVCache
v1.1.0 Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672 RK1828 Vision: 384x384 Vision: w4a16;LLM: w4a16 Vision: 8;LLM: 8 LLM: 28672(28k) LLM: fp16
v1.1.0 Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4 RK1828 Vision: 384x384 Vision: w4a16;LLM: w4a16 Vision: 8;LLM: 8 LLM: 66560(65k) LLM: int4

2. 文件说明

文件或目录 说明
LICENSE 源模型许可证
<配置目录>/ 配套模型文件
子目录 README.md 当前配置说明
子目录 config.json 模型配置与文件清单
子目录 SHA256SUMS 当前目录交付文件的 SHA-256 校验值(不包含自身)

3. 模型下载

ModelScope 完整下载:

modelscope download --model RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct --revision v1.1.0 --local_dir ./RK182X-VLM-Qwen3-VL-2B-Instruct

Hugging Face 完整下载:

hf download RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct --revision v1.1.0 --local-dir ./RK182X-VLM-Qwen3-VL-2B-Instruct

ModelScope 指定配置下载:

from modelscope import snapshot_download

snapshot_download(
    "RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct",
    revision="v1.1.0",
    allow_patterns=["README.md", "LICENSE", "Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/**"],
    local_dir="./RK182X-VLM-Qwen3-VL-2B-Instruct",
)

Hugging Face 指定配置下载:

hf download RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct --revision v1.1.0 --include "README.md" "LICENSE" "Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/**" --local-dir ./RK182X-VLM-Qwen3-VL-2B-Instruct

4. SHA-256 校验

在配置目录执行:

cd ./RK182X-VLM-Qwen3-VL-2B-Instruct/Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672
sha256sum -c SHA256SUMS

所有条目显示 OK 后再部署。

5. 兼容性与限制

  • 支持芯片:RK1828。
  • 使用配套 RKNN Runtime 和驱动;运行前用 rknn-smi -v 检查设备端版本。
  • 所选配置目录内的文件须配套使用。

6. 版权与许可证

源模型许可证见根目录 LICENSE;同时遵守 RKNN Toolkit 和 RKNN Runtime 许可条款。

Identity and Version

Repository
RKNNAI/RK182X-VLM-Qwen3-VL-2B-Instruct
Publisher
RKNNAI
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
vlm
Revision
e5e82e8ea6e9d72266f527d1851a0bca2acb5c04
First published
2026-09-24
Last updated
2026-09-28

Files and Weights

25 files, 4.0 GB in total. The weights are 4 files totalling 1.3 GB in bin, gguf.

Weights4 files · 1.3 GB
Configuration2 files · 2.6 KB
Documentation4 files · 19.3 KB
Other14 files · 2.8 GB
Repository1 file · 2.6 KB
Every file
FileTypeSizeSHA-256
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/Qwen3-VL-2B-llm.embed.binWeights622.3 MB 80a4c412dd89
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/Qwen3-VL-2B-llm.tokenizer.ggufWeights5.9 MB d51077b1e11f
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/Qwen3-VL-2B-llm.embed.binWeights622.3 MB 80a4c412dd89
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/Qwen3-VL-2B-llm.tokenizer.ggufWeights5.9 MB d51077b1e11f
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/config.jsonConfiguration1.3 KB —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/config.jsonConfiguration1.3 KB —
LICENSEDocumentation11.3 KB —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/README.mdDocumentation2.4 KB —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/README.mdDocumentation2.4 KB —
README.mdDocumentation3.1 KB —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/Qwen3-VL-2B-llm.rknnOther33.2 MB 9e921aa1a9d1
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/Qwen3-VL-2B-llm.weightOther1.1 GB f0893007007e
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/Qwen3-VL-2B-vision.rknnOther5.4 MB 92dbe53caab8
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/Qwen3-VL-2B-vision.weightOther240.0 MB c4d48eb33a65
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/SHA256SUMSOther880 B —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/llm_model_report.htmlOther827.4 KB —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-28672/vision_model_report.htmlOther87.1 KB —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/Qwen3-VL-2B-llm.rknnOther65.3 MB 6c4d2dc56ac3
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/Qwen3-VL-2B-llm.weightOther1.1 GB f0893007007e
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/Qwen3-VL-2B-vision.rknnOther5.4 MB 92dbe53caab8
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/Qwen3-VL-2B-vision.weightOther240.0 MB c4d48eb33a65
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/SHA256SUMSOther880 B —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/llm_model_report.htmlOther1.0 MB —
Qwen3-VL-2B-Instruct-384x384-w4a16-8-66560-kv-int4/vision_model_report.htmlOther87.1 KB —
.gitattributesRepository2.6 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.3 GB
Download from RKNNAI

Released by RKNNAI through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published1.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About RK182X-VLM-Qwen3-VL-2B-Instruct

Can I use RK182X-VLM-Qwen3-VL-2B-Instruct commercially?

Yes. RK182X-VLM-Qwen3-VL-2B-Instruct is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Non-uniform GGUF quantizations of a 512-expert MoE, produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-Flash-Next at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. A…

Open weights apache-2.0 gguf

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads) instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model ctrl at the chat/message level). Model name has "plusIQ" in the name. A 12+12 (12 reasoning and 12 instruct) model with interactive optimization/help system will be releasing shortly too. BF16/16-bit MTP GGUF also avail. (there is also a extra robust "tools" version too - Q6 and Q8.) Extreme intelligence in a small package. Jaw dropping performance. Superior…

Open weights apache-2.0