SAVRN
Search Contact SAVRN

Open-weight model · Text to speech

Qwen3-TTS-12Hz-1.7B-CustomVoice

by Qwen Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs.

Parameters1.9B
Context
Weights4.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads2.6M

Runs On

What it takes to serve Qwen3-TTS-12Hz-1.7B-CustomVoice (1.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 3.8 GB 4.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.9 GB 2.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.0 GB 1.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Qwen, published under apache-2.0, revision 0c0e3051f131.

Qwen3-TTS

Overview

Introduction

Read the full model card (3,390 words)

Configuration

Architecture
Qwen3TTSForConditionalGeneration
Model type
qwen3_tts

Identity and Version

Repository
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Publisher
Qwen
Task
Text to speech
Modality
Audio
Library
Not stated by the source
Parameters
1.9B parameters
Languages
Not stated by the source
Revision
0c0e3051f131929182e2c023b9537f8b1c68adfe
First published
2026-01-21
Last updated
2026-01-29

Files and Weights

13 files, 4.5 GB in total. The weights are 2 files totalling 4.5 GB in safetensors.

Weights2 files · 4.5 GB
Configuration6 files · 7.9 KB
Tokenizer3 files · 4.5 MB
Documentation1 file · 57.8 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights3.8 GB 38b1d5971bdb
speech_tokenizer/model.safetensorsWeights682.3 MB 836b7b357f5e
config.jsonConfiguration4.9 KB
generation_config.jsonConfiguration245 B
preprocessor_config.jsonConfiguration127 B
speech_tokenizer/config.jsonConfiguration2.3 KB
speech_tokenizer/configuration.jsonConfiguration76 B
speech_tokenizer/preprocessor_config.jsonConfiguration234 B
README.mdDocumentation57.8 KB
.gitattributesRepository1.5 KB
merges.txtTokenizer1.7 MB
tokenizer_config.jsonTokenizer7.3 KB
vocab.jsonTokenizer2.8 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
4.5 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published4.5 GB
16-bit3.8 GB
8-bit1.9 GB
4-bit1.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare Qwen3-TTS-12Hz-1.7B-CustomVoice

Questions About Qwen3-TTS-12Hz-1.7B-CustomVoice

How much GPU memory does Qwen3-TTS-12Hz-1.7B-CustomVoice need?

About 4.6 GB at 16-bit and 1.2 GB at 4-bit: the weights (1.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3-TTS-12Hz-1.7B-CustomVoice on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3-TTS-12Hz-1.7B-CustomVoice commercially?

Yes. Qwen3-TTS-12Hz-1.7B-CustomVoice is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to speech

Qwen3-TTS-12Hz-1.7B-VoiceDesign

Qwen

We release Qwen3-TTS, a series of powerful speech generation models developed by Qwen, offering comprehensive support for voice cloning, voice design, ultra-high-quality human-like speech generation, and natural language-based voice control. Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles. Key features: Install the qwen-tts Python package from PyPI: Zero-shot speech generation on the Seed-TTS test set (Word Error Rate (WER, ↓)): If you find our paper and code useful in your research, please consider giving a star and citation

Open weights apache-2.0 1.9B parameters qwen-tts

Model · Text to speech

MOSS-VoiceGenerator

OpenMOSS

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS. When a single piece of audio needs to sound like a real person, pronounce every word accurately, switch speaking styles across content, remain stable over tens of minutes, and support dialogue, role‑play, and real‑time interaction, a single TTS model is often not enough. The MOSS‑TTS Family breaks the workflow into five production‑ready models that can be…

Open weights apache-2.0 2.1B parameters 40,960 tokens

Model · Text to speech

Zonos-v0.1-transformer

Zyphra

alt="Title card" style="width: 500px; Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers. Our model enables highly natural speech generation from text prompts when given a speaker embedding or audio prefix, and can accurately perform speech cloning when given a reference clip spanning just a few seconds. The conditioning setup also allows for fine control over speaking rate, pitch variation, audio quality, and emotions such as happiness, fear, sadness, and anger. The model outputs speech natively at 44kHz. Zonos follows a straightforward…

Open weights apache-2.0 1.6B parameters zonos

Model · Text to speech

csm-1b

Sesame

2025/05/20 - CSM is availabile natively in Hugging Face Transformers as of version 4.52.1 2025/03/13 - We are releasing the 1B CSM variant. The checkpoint is hosted on Hugging Face. CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs. The model architecture employs a Llama backbone and a smaller audio decoder that produces Mimi audio codes. A fine-tuned variant of CSM powers the interactive voice demo shown in our blog post. A hosted HuggingFace space is also available for testing audio generation. CSM supports full-graph compilation with CUDA graphs! CSM can be fine-tuned using Transformers' Trainer. Does this…

Access requested at publisher apache-2.0 1.6B parameters transformers

Model · Text to speech

VoxCPM2

OpenBMB

VoxCPM2 is a tokenizer-free, diffusion autoregressive Text-to-Speech model — 2B parameters, 30 languages, 48kHz audio output, trained on over 2 million hours of multilingual speech data. - 30-Language Multilingual — No language tag needed; input text in any supported language directly - Voice Design — Generate a novel voice from a natural-language description alone (gender, age, tone, emotion, pace…); no reference audio required - Controllable Cloning — Clone any voice from a short clip, with optional style guidance to steer emotion, pace, and expression while preserving timbre - Ultimate Cloning — Provide reference audio + its transcript for audio-continuation cloning; every vocal nuance…

Open weights apache-2.0 2.3B parameters voxcpm

Model · Text to speech

VibeVoice-1.5B

Microsoft

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a Large Language Model (LLM) to understand textual…

Open weights mit 2.7B parameters transformers