SAVRN
Search Contact SAVRN

Open-weight model · Text to speech

Qwen3-TTS-12Hz-1.7B-VoiceDesign

by Qwen Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign

We release Qwen3-TTS, a series of powerful speech generation models developed by Qwen, offering comprehensive support for voice cloning, voice design, ultra-high-quality human-like speech generation, and natural language-based voice control.

Parameters1.9B
Context
Weights4.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads283.7k

Runs On

What it takes to serve Qwen3-TTS-12Hz-1.7B-VoiceDesign (1.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 3.8 GB 4.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.9 GB 2.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.0 GB 1.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Qwen, published under apache-2.0, revision 5ecdb67327fd.

Qwen3-TTS


Read the full model card (301 words)

Configuration

Architecture
Qwen3TTSForConditionalGeneration
Model type
qwen3_tts

Identity and Version

Repository
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
Publisher
Qwen
Task
Text to speech
Modality
Audio
Library
qwen-tts
Parameters
1.9B parameters
Languages
tts
Revision
5ecdb67327fd37bb2e042aab12ff7391903235d3
First published
2026-01-21
Last updated
2026-01-29

Files and Weights

13 files, 4.5 GB in total. The weights are 2 files totalling 4.5 GB in safetensors.

Weights2 files · 4.5 GB
Configuration6 files · 7.4 KB
Tokenizer3 files · 4.5 MB
Documentation1 file · 3.2 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights3.8 GB 391e8db219f2
speech_tokenizer/model.safetensorsWeights682.3 MB 836b7b357f5e
config.jsonConfiguration4.4 KB
generation_config.jsonConfiguration245 B
preprocessor_config.jsonConfiguration127 B
speech_tokenizer/config.jsonConfiguration2.3 KB
speech_tokenizer/configuration.jsonConfiguration76 B
speech_tokenizer/preprocessor_config.jsonConfiguration234 B
README.mdDocumentation3.2 KB
.gitattributesRepository1.5 KB
merges.txtTokenizer1.7 MB
tokenizer_config.jsonTokenizer7.3 KB
vocab.jsonTokenizer2.8 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
4.5 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published4.5 GB
16-bit3.8 GB
8-bit1.9 GB
4-bit1.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare Qwen3-TTS-12Hz-1.7B-VoiceDesign

Questions About Qwen3-TTS-12Hz-1.7B-VoiceDesign

How much GPU memory does Qwen3-TTS-12Hz-1.7B-VoiceDesign need?

About 4.6 GB at 16-bit and 1.2 GB at 4-bit: the weights (1.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3-TTS-12Hz-1.7B-VoiceDesign on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3-TTS-12Hz-1.7B-VoiceDesign commercially?

Yes. Qwen3-TTS-12Hz-1.7B-VoiceDesign is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to speech

Qwen3-TTS-12Hz-1.7B-CustomVoice

Qwen

Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs. In addition, the models feature strong contextual understanding, enabling adaptive control of tone, speaking rate, and emotional expression based on instructions and text semantics, and they show markedly improved robustness to noisy input text. Key features: Intelligent Text Understanding and Voice Control: Supports speech generation driven by natural language instructions, allowing for flexible control over multi-dimensional acoustic attributes such as timbre, emotion, and prosody.…

Open weights apache-2.0 1.9B parameters

Model · Text to speech

MOSS-VoiceGenerator

OpenMOSS

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS. When a single piece of audio needs to sound like a real person, pronounce every word accurately, switch speaking styles across content, remain stable over tens of minutes, and support dialogue, role‑play, and real‑time interaction, a single TTS model is often not enough. The MOSS‑TTS Family breaks the workflow into five production‑ready models that can be…

Open weights apache-2.0 2.1B parameters 40,960 tokens

Model · Text to speech

Zonos-v0.1-transformer

Zyphra

alt="Title card" style="width: 500px; Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers. Our model enables highly natural speech generation from text prompts when given a speaker embedding or audio prefix, and can accurately perform speech cloning when given a reference clip spanning just a few seconds. The conditioning setup also allows for fine control over speaking rate, pitch variation, audio quality, and emotions such as happiness, fear, sadness, and anger. The model outputs speech natively at 44kHz. Zonos follows a straightforward…

Open weights apache-2.0 1.6B parameters zonos

Model · Text to speech

csm-1b

Sesame

2025/05/20 - CSM is availabile natively in Hugging Face Transformers as of version 4.52.1 2025/03/13 - We are releasing the 1B CSM variant. The checkpoint is hosted on Hugging Face. CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs. The model architecture employs a Llama backbone and a smaller audio decoder that produces Mimi audio codes. A fine-tuned variant of CSM powers the interactive voice demo shown in our blog post. A hosted HuggingFace space is also available for testing audio generation. CSM supports full-graph compilation with CUDA graphs! CSM can be fine-tuned using Transformers' Trainer. Does this…

Access requested at publisher apache-2.0 1.6B parameters transformers

Model · Text to speech

VoxCPM2

OpenBMB

VoxCPM2 is a tokenizer-free, diffusion autoregressive Text-to-Speech model — 2B parameters, 30 languages, 48kHz audio output, trained on over 2 million hours of multilingual speech data. - 30-Language Multilingual — No language tag needed; input text in any supported language directly - Voice Design — Generate a novel voice from a natural-language description alone (gender, age, tone, emotion, pace…); no reference audio required - Controllable Cloning — Clone any voice from a short clip, with optional style guidance to steer emotion, pace, and expression while preserving timbre - Ultimate Cloning — Provide reference audio + its transcript for audio-continuation cloning; every vocal nuance…

Open weights apache-2.0 2.3B parameters voxcpm

Model · Text to speech

VibeVoice-1.5B

Microsoft

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a Large Language Model (LLM) to understand textual…

Open weights mit 2.7B parameters transformers