SAVRN
Search Contact SAVRN

Open-weight model · Text to speech

RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base

by RKNNAI RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base

RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base is an open-weight model for text to speech from RKNNAI, released under Apache License 2.0. Its published files total 2.0 GB.

本仓库提供由 Qwen/Qwen3-TTS-12Hz-1.7B-Base 转换得到的 RKNN TTS 模型。 - Model ID:RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base - 模型显示名称:RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base - 源模型:Qwen/Qwen3-TTS-12Hz-1.7B-Base - 模型类型:TTS ModelScope 完整下载: Hugging Face 完整下载: ModelScope 指定配置下载:…

Parameters—
Context—
Weights760.7 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By RKNNAI, published under apache-2.0, revision e4ea77ecb669.

本仓库提供由 Qwen/Qwen3-TTS-12Hz-1.7B-Base 转换得到的 RKNN TTS 模型。 - Model ID:RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base - 模型显示名称:RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base - 源模型:Qwen/Qwen3-TTS-12Hz-1.7B-Base - 模型类型:TTS ModelScope 完整下载: Hugging Face 完整下载: ModelScope 指定配置下载: Hugging Face 指定配置下载: - 使用配套 RKNN Runtime 和驱动;运行前用 rknn-smi -v 检查设备端版本。 源模型许可证见根目录 LICENSE;同时遵守 RKNN Toolkit 和 RKNN Runtime 许可条款。

Read RKNNAI's full model card

1. 模型介绍

本仓库提供由 Qwen/Qwen3-TTS-12Hz-1.7B-Base 转换得到的 RKNN TTS 模型。

  • Model ID:RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base
  • 模型显示名称:RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base
  • 发布版本:v1.1.0
  • 源模型:Qwen/Qwen3-TTS-12Hz-1.7B-Base
  • 模型类型:TTS
  • 芯片字段:RK182X
  • 具体支持芯片:RK1820、RK1828

可用模型

发布版本 配置目录 支持芯片 量化方式 NPU 核数
v1.1.0 Qwen3-TTS-12Hz-1.7B-Base-w4a16-8 RK1820、RK1828 Talker: w4a16;Code Predictor: w4a16;Speech Decoder: w4a16;Text Projector: w4a16 Talker: 8;Code Predictor: 8;Speech Decoder: 8;Text Projector: 1

2. 文件说明

文件或目录 说明
LICENSE 源模型许可证
<配置目录>/ 配套模型文件
子目录 README.md 当前配置说明
子目录 config.json 模型配置与文件清单
子目录 SHA256SUMS 当前目录交付文件的 SHA-256 校验值(不包含自身)

3. 模型下载

ModelScope 完整下载:

modelscope download --model RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base --revision v1.1.0 --local_dir ./RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base

Hugging Face 完整下载:

hf download RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base --revision v1.1.0 --local-dir ./RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base

ModelScope 指定配置下载:

from modelscope import snapshot_download

snapshot_download(
    "RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base",
    revision="v1.1.0",
    allow_patterns=["README.md", "LICENSE", "Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/**"],
    local_dir="./RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base",
)

Hugging Face 指定配置下载:

hf download RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base --revision v1.1.0 --include "README.md" "LICENSE" "Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/**" --local-dir ./RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base

4. SHA-256 校验

在配置目录执行:

cd ./RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base/Qwen3-TTS-12Hz-1.7B-Base-w4a16-8
sha256sum -c SHA256SUMS

所有条目显示 OK 后再部署。

5. 兼容性与限制

  • 支持芯片:RK1820、RK1828。
  • 使用配套 RKNN Runtime 和驱动;运行前用 rknn-smi -v 检查设备端版本。
  • 所选配置目录内的文件须配套使用。

6. 版权与许可证

源模型许可证见根目录 LICENSE;同时遵守 RKNN Toolkit 和 RKNN Runtime 许可条款。

Identity and Version

Repository
RKNNAI/RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base
Publisher
RKNNAI
Task
Text to speech
Modality
Audio
Library
Not stated by the source
Parameters
Not stated by the source
Languages
tts
Revision
e4ea77ecb6695d20656d4e3ed29af57facd6c360
First published
2026-09-24
Last updated
2026-09-28

Files and Weights

23 files, 2.0 GB in total. The weights are 3 files totalling 760.7 MB in bin.

Weights3 files · 760.7 MB
Configuration1 file · 2.5 KB
Tokenizer1 file · 11.4 MB
Documentation3 files · 17.6 KB
Other14 files · 1.3 GB
Repository1 file · 2.6 KB
Every file
FileTypeSizeSHA-256
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/embeds/Qwen3-TTS-12Hz-1.7B-Base-codec_embed.fp16.binWeights125.8 MB 934471109029
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/embeds/Qwen3-TTS-12Hz-1.7B-Base-talker_input_embed.fp16.binWeights12.6 MB b13f93e9ae5b
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/embeds/Qwen3-TTS-12Hz-1.7B-Base-talker_text_embed.fp16.binWeights622.3 MB ad513a33a330
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/config.jsonConfiguration2.5 KB —
LICENSEDocumentation11.3 KB —
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/README.mdDocumentation3.4 KB —
README.mdDocumentation2.8 KB —
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/SHA256SUMSOther2.2 KB —
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/code_predictor/Qwen3-TTS-12Hz-1.7B-Base-code_predictor.rknnOther9.7 MB ebc5957183b0
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/code_predictor/Qwen3-TTS-12Hz-1.7B-Base-code_predictor.weightOther151.3 MB 33110dcb293c
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/code_predictor/Qwen3-TTS-12Hz-1.7B-Base-code_predictor_model_report.htmlOther242.4 KB —
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/model_report.htmlOther816.5 KB —
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/speech_decoder/Qwen3-TTS-12Hz-1.7B-Base-speech_decoder.rknnOther3.8 MB c2e45f8a7929
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/speech_decoder/Qwen3-TTS-12Hz-1.7B-Base-speech_decoder.weightOther228.6 MB cc929a6f957f
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/speech_decoder/Qwen3-TTS-12Hz-1.7B-Base-speech_decoder_model_report.htmlOther85.0 KB —
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/talker/Qwen3-TTS-12Hz-1.7B-Base-talker.rknnOther26.8 MB 917b14dd3285
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/talker/Qwen3-TTS-12Hz-1.7B-Base-talker.weightOther822.4 MB c42285f85063
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/talker/Qwen3-TTS-12Hz-1.7B-Base-talker_model_report.htmlOther816.5 KB —
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/text_projector/Qwen3-TTS-12Hz-1.7B-Base-text_projection.rknnOther154.6 KB 997856223a21
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/text_projector/Qwen3-TTS-12Hz-1.7B-Base-text_projection.weightOther16.8 MB 116316c3fedf
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/text_projector/Qwen3-TTS-12Hz-1.7B-Base-text_projection_model_report.htmlOther74.3 KB —
.gitattributesRepository2.6 KB —
Qwen3-TTS-12Hz-1.7B-Base-w4a16-8/embeds/Qwen3-TTS-12Hz-1.7B-Base-tokenizer.jsonTokenizer11.4 MB 09267689b836

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
760.7 MB
Download from RKNNAI

Released by RKNNAI through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published760.7 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base

Can I use RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base commercially?

Yes. RK182X-TTS-Qwen3-TTS-12Hz-1.7B-Base is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to speech

Kokoro-82M

Hexgrad

Kokoro is an open-weight TTS model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects. You can run this basic cell on Google Colab. Listen to samples. For more languages and details, see Advanced Usage. Under the hood, kokoro uses misaki, a G2P library at https://github.com/hexgrad/misaki Model SHA256 Hash: 496dba118d1a58f5f3db2efc88dbdc216e0483fc89fe6e47ee1f2c53f18ad1e4 Data: Kokoro was trained exclusively on permissive/non-copyrighted audio data and IPA…

Open weights apache-2.0

Model · Text to speech

XTTS-v2

Coqui.ai

ⓍTTS is a Voice generation model that lets you clone voices into different languages by using just a quick 6-second audio clip. There is no need for an excessive amount of training data that spans countless hours. This is the same or similar model to what powers Coqui Studio and Coqui API. - Supports 17 languages. - Voice cloning with just a 6-second audio clip. - Emotion and style transfer by cloning. - Cross-language voice cloning. - Multi-lingual speech generation. - 24khz sampling rate. - 2 new languages; Hungarian and Korean - Architectural improvements for speaker conditioning. - Enables the use of multiple speaker references and interpolation between speakers. - Stability…

Open weights other coqui

Model · Text to speech

audio.cpp-gguf

Audio.cpp

This directory contains audio.cpp-native GGUF conversions of multiple speech models. These files are intended for use with audio.cpp. If you enjoy the project, please star audio.cpp on GitHub and this Hugging Face repository. For conversion details, supported layouts, direct-file loading, sidecar embedding, and the latest compatibility notes, see the audio.cpp GGUF guide: - https://github.com/0xShug0/audio.cpp/blob/main/docs/gguf.md!!! Converted and quantized packages are checked with automated metrics, but perceived quality can still differ for human listeners. Please validate the exact package, backend, and route to confirm the output is acceptable for your use case. The table lists the…

Open weights other audio.cpp

Model · Text to speech

chatterbox

Resemble AI

Chatterbox Multilingual V3 is the latest general-purpose multilingual TTS model in the Chatterbox family. It keeps the same 0.5B model size while improving speaker similarity, reducing hallucinations, and producing more natural, conversational speech across languages. V3 is designed for broad language coverage like V2, but with stronger stability and more expressive generation. It is the recommended multilingual model for users who want one voice cloning model that works across many languages. Try it in the Chatterbox Multilingual TTS V3 Space. Alongside V3, we are releasing the Single Language Pack: dedicated finetunes for priority languages where tighter quality control, stronger…

Open weights mit chatterbox

Model · Text to speech

F5-TTS

Yushen CHEN

Download F5-TTS or E2 TTS and place under ckpts/ Paper: F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Open weights cc-by-nc-4.0 f5-tts

Model · Text to speech

Kokoro-82M-v1.0-ONNX

ONNX Community

Kokoro is a frontier TTS model for its size of 82 million parameters (text in/audio out). First, install the kokoro-js library from NPM using: You can then generate speech as follows: Optionally, save the audio to a file: The model is resilient to quantization, enabling efficient high-quality speech synthesis at a fraction of the original model size.

Open weights apache-2.0 transformers.js