SAVRN
Search Contact SAVRN

Open-weight model · Speech recognition

RK182X-ASR-Qwen3-ASR-0.6B

by RKNNAI RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B

RK182X-ASR-Qwen3-ASR-0.6B is an open-weight model for speech recognition from RKNNAI, released under Apache License 2.0. Its published files total 2.3 GB. It draws 1 downloads a month.

本仓库提供由 Qwen/Qwen3-ASR-0.6B 转换得到的 RKNN ASR 模型。 - Model ID:RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B - 模型显示名称:RK182X-ASR-Qwen3-ASR-0.6B - 源模型:Qwen/Qwen3-ASR-0.6B - 模型类型:ASR ModelScope 完整下载: Hugging Face 完整下载: ModelScope 指定配置下载: Hugging Face 指定配置下载: - 使用配套 RKNN Runtime…

Parameters—
Context—
Weights634.2 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1

Model Card

By RKNNAI, published under apache-2.0, revision 898daacaa983.

本仓库提供由 Qwen/Qwen3-ASR-0.6B 转换得到的 RKNN ASR 模型。 - Model ID:RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B - 模型显示名称:RK182X-ASR-Qwen3-ASR-0.6B - 源模型:Qwen/Qwen3-ASR-0.6B - 模型类型:ASR ModelScope 完整下载: Hugging Face 完整下载: ModelScope 指定配置下载: Hugging Face 指定配置下载: - 使用配套 RKNN Runtime 和驱动;运行前用 rknn-smi -v 检查设备端版本。 源模型许可证见根目录 LICENSE;同时遵守 RKNN Toolkit 和 RKNN Runtime 许可条款。

Read RKNNAI's full model card

1. 模型介绍

本仓库提供由 Qwen/Qwen3-ASR-0.6B 转换得到的 RKNN ASR 模型。

  • Model ID:RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B
  • 模型显示名称:RK182X-ASR-Qwen3-ASR-0.6B
  • 发布版本:v1.1.0
  • 源模型:Qwen/Qwen3-ASR-0.6B
  • 模型类型:ASR
  • 芯片字段:RK182X
  • 具体支持芯片:RK1820、RK1828

可用模型

发布版本 配置目录 支持芯片 量化方式 NPU 核数
v1.1.0 Qwen3-ASR-0.6B-w4a16-8-offline RK1820、RK1828 Encoder: w4a16;LLM: w4a16 Encoder: 8;LLM: 8
v1.1.0 Qwen3-ASR-0.6B-w4a16-8-online RK1820、RK1828 Encoder: w4a16;LLM: w4a16 Encoder: 8;LLM: 8

2. 文件说明

文件或目录 说明
LICENSE 源模型许可证
<配置目录>/ 配套模型文件
子目录 README.md 当前配置说明
子目录 config.json 模型配置与文件清单
子目录 SHA256SUMS 当前目录交付文件的 SHA-256 校验值(不包含自身)

3. 模型下载

ModelScope 完整下载:

modelscope download --model RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B --revision v1.1.0 --local_dir ./RK182X-ASR-Qwen3-ASR-0.6B

Hugging Face 完整下载:

hf download RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B --revision v1.1.0 --local-dir ./RK182X-ASR-Qwen3-ASR-0.6B

ModelScope 指定配置下载:

from modelscope import snapshot_download

snapshot_download(
    "RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B",
    revision="v1.1.0",
    allow_patterns=["README.md", "LICENSE", "Qwen3-ASR-0.6B-w4a16-8-offline/**"],
    local_dir="./RK182X-ASR-Qwen3-ASR-0.6B",
)

Hugging Face 指定配置下载:

hf download RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B --revision v1.1.0 --include "README.md" "LICENSE" "Qwen3-ASR-0.6B-w4a16-8-offline/**" --local-dir ./RK182X-ASR-Qwen3-ASR-0.6B

4. SHA-256 校验

在配置目录执行:

cd ./RK182X-ASR-Qwen3-ASR-0.6B/Qwen3-ASR-0.6B-w4a16-8-offline
sha256sum -c SHA256SUMS

所有条目显示 OK 后再部署。

5. 兼容性与限制

  • 支持芯片:RK1820、RK1828。
  • 使用配套 RKNN Runtime 和驱动;运行前用 rknn-smi -v 检查设备端版本。
  • 所选配置目录内的文件须配套使用。

6. 版权与许可证

源模型许可证见根目录 LICENSE;同时遵守 RKNN Toolkit 和 RKNN Runtime 许可条款。

Identity and Version

Repository
RKNNAI/RK182X-ASR-Qwen3-ASR-0.6B
Publisher
RKNNAI
Task
Speech recognition
Modality
Audio
Library
Not stated by the source
Parameters
Not stated by the source
Languages
asr
Revision
898daacaa9832d6b247e76819e91c4a9e0a33fdc
First published
2026-09-24
Last updated
2026-09-28

Files and Weights

27 files, 2.3 GB in total. The weights are 4 files totalling 634.2 MB in bin, gguf.

Weights4 files · 634.2 MB
Configuration2 files · 2.8 KB
Documentation4 files · 18.7 KB
Other16 files · 1.6 GB
Repository1 file · 2.5 KB
Every file
FileTypeSizeSHA-256
Qwen3-ASR-0.6B-w4a16-8-offline/Qwen3-ASR-0.6B-llm.embed.binWeights311.2 MB e80150119fa5
Qwen3-ASR-0.6B-w4a16-8-offline/Qwen3-ASR-0.6B-llm.tokenizer.ggufWeights5.9 MB 2b89ff9f19cc
Qwen3-ASR-0.6B-w4a16-8-online/Qwen3-ASR-0.6B-llm.embed.binWeights311.2 MB e80150119fa5
Qwen3-ASR-0.6B-w4a16-8-online/Qwen3-ASR-0.6B-llm.tokenizer.ggufWeights5.9 MB 2b89ff9f19cc
Qwen3-ASR-0.6B-w4a16-8-offline/config.jsonConfiguration1.4 KB —
Qwen3-ASR-0.6B-w4a16-8-online/config.jsonConfiguration1.4 KB —
LICENSEDocumentation11.3 KB —
Qwen3-ASR-0.6B-w4a16-8-offline/README.mdDocumentation2.3 KB —
Qwen3-ASR-0.6B-w4a16-8-online/README.mdDocumentation2.3 KB —
README.mdDocumentation2.7 KB —
Qwen3-ASR-0.6B-w4a16-8-offline/Qwen3-ASR-0.6B-encoder.rknnOther3.5 MB bfcb57d2371a
Qwen3-ASR-0.6B-w4a16-8-offline/Qwen3-ASR-0.6B-encoder.weightOther375.0 MB 71aa65ceac7d
Qwen3-ASR-0.6B-w4a16-8-offline/Qwen3-ASR-0.6B-encoder_model_report.htmlOther85.0 KB —
Qwen3-ASR-0.6B-w4a16-8-offline/Qwen3-ASR-0.6B-llm.rknnOther47.4 MB 1088b5963b0b
Qwen3-ASR-0.6B-w4a16-8-offline/Qwen3-ASR-0.6B-llm.weightOther387.2 MB a70fa91ce5c2
Qwen3-ASR-0.6B-w4a16-8-offline/Qwen3-ASR-0.6B-llm_model_report.htmlOther1.0 MB —
Qwen3-ASR-0.6B-w4a16-8-offline/SHA256SUMSOther1.0 KB —
Qwen3-ASR-0.6B-w4a16-8-offline/model_report.htmlOther1.0 MB —
Qwen3-ASR-0.6B-w4a16-8-online/Qwen3-ASR-0.6B-encoder.rknnOther3.1 MB 1096c55bb45e
Qwen3-ASR-0.6B-w4a16-8-online/Qwen3-ASR-0.6B-encoder.weightOther374.8 MB 1b23225f8555
Qwen3-ASR-0.6B-w4a16-8-online/Qwen3-ASR-0.6B-encoder_model_report.htmlOther85.0 KB —
Qwen3-ASR-0.6B-w4a16-8-online/Qwen3-ASR-0.6B-llm.rknnOther47.4 MB 1088b5963b0b
Qwen3-ASR-0.6B-w4a16-8-online/Qwen3-ASR-0.6B-llm.weightOther387.2 MB a70fa91ce5c2
Qwen3-ASR-0.6B-w4a16-8-online/Qwen3-ASR-0.6B-llm_model_report.htmlOther1.0 MB —
Qwen3-ASR-0.6B-w4a16-8-online/SHA256SUMSOther1.0 KB —
Qwen3-ASR-0.6B-w4a16-8-online/model_report.htmlOther1.0 MB —
.gitattributesRepository2.5 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
634.2 MB
Download from RKNNAI

Released by RKNNAI through its official repository on Hugging Face. Read the license.

Built From

  • Derived from Qwen/Qwen3-ASR-0.6B
  • Quantized from Qwen/Qwen3-ASR-0.6B

Memory Requirements

PrecisionWeights in memory
As published634.2 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About RK182X-ASR-Qwen3-ASR-0.6B

Can I use RK182X-ASR-Qwen3-ASR-0.6B commercially?

Yes. RK182X-ASR-Qwen3-ASR-0.6B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Speech recognition

wav2vec2-large-xlsr-53-japanese

Jonatas Grosman

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Japanese using the train and validation splits of Common Voice 6.1, CSS10 and JSUT. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: The model can be evaluated as follows on the Japanese test data of Common Voice. In the table below I report the Word Error Rate (WER) and the Character Error Rate (CER) of the model. I ran the…

Open weights apache-2.0 transformers

Model · Speech recognition

whisperkit-coreml

Argmax

WhisperKit is part of Argmax OSS, an On-device Speech AI SDK for Apple Silicon: https://github.com/argmaxinc/argmax-oss-swift Check out the WhisperKit paper and presentation from ICML 2025: https://icml.cc/virtual/2025/47854 For real-time transcription with speakers and custom vocabulary, check out Argmax Pro SDK: https://www.argmaxinc.com/blog/argmax-sdk-2

Open weights mit whisperkit

Model · Speech recognition

speaker-diarization-3.1

Pyannote

Using this open-source model in production? Consider switching to pyannoteAI for better and faster options. This pipeline is the same as pyannote/speaker-diarization-3.0 except it removes the problematic use of onnxruntime. Both speaker segmentation and embedding now run in pure PyTorch. This should ease deployment and possibly speed up inference. It requires pyannote.audio version 3.1 or higher. It ingests mono audio sampled at 16kHz and outputs speaker diarization as an Annotation instance: - stereo or multi-channel audio files are automatically downmixed to mono by averaging the channels. - audio files sampled at a different rate are resampled to 16kHz automatically upon loading. 1.…

Access requested at publisher mit pyannote-audio

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Portuguese using the train and validation splits of Common Voice 6.1. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: 1. To evaluate on mozilla-foundation/commonvoice60 with split test 2. To evaluate on speech-recognition-community-v2/devdata If you want to cite this model you can use this

Open weights apache-2.0 transformers

Model · Speech recognition

speaker-diarization-community-1

Pyannote

This pipeline ingests mono audio sampled at 16kHz and outputs speaker diarization. - stereo or multi-channel audio files are automatically downmixed to mono by averaging the channels. - audio files sampled at a different rate are resampled to 16kHz automatically upon loading. The main improvements brought by Community-1 are: - improved speaker assignment and counting - simpler reconciliation with transcription timestamps with exclusive speaker diarization - easy offline use (i.e. without internet connection) - (optionally) hosted on pyannoteAI cloud 1. pip install pyannote.audio 3. Create access token at hf.co/settings/tokens. Out of the box, Community-1 is much better than…

Access requested at publisher cc-by-4.0 pyannote-audio

Model · Speech recognition

wav2vec2-large-xlsr-53-russian

Jonatas Grosman

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Russian using the train and validation splits of Common Voice 6.1 and CSS10. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: 1. To evaluate on mozilla-foundation/commonvoice60 with split test 2. To evaluate on speech-recognition-community-v2/devdata If you want to cite this model you can use this

Open weights apache-2.0 transformers