SAVRN
Search Contact SAVRN

SAVRN Model Hub · Comparisons

MOSS-VoiceGenerator vs Qwen3-TTS-12Hz-1.7B-CustomVoice

MOSS-VoiceGenerator has 2.1B parameters and Qwen3-TTS-12Hz-1.7B-CustomVoice has 1.9B parameters; both are released under Apache License 2.0; at 16-bit, MOSS-VoiceGenerator needs about 5.1 GB (1x MI300X from $1.85 an hour) and Qwen3-TTS-12Hz-1.7B-CustomVoice about 4.6 GB (1x MI300X from $1.85 an hour).

Published metadata for 2 models, each read from its own repository.
Field MOSS-VoiceGenerator
OpenMOSS-Team/MOSS-VoiceGenerator
Qwen3-TTS-12Hz-1.7B-CustomVoice
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Publisher OpenMOSS Qwen
Task Text to speech Text to speech
Modality Audio Audio
Parameters, as reported 2.1B parameters 1.9B parameters
Architecture MossTTSDelayModel Qwen3TTSForConditionalGeneration
Library Not stated Not stated
Context length 40,960 tokens Not stated
Repository size 4.2 GB 4.5 GB
Artifact formats safetensors safetensors
License apache-2.0 apache-2.0
Access Open weights, no gate Open weights, no gate
Memory at 16-bit (weights and margin) 5.1 GB 4.6 GB
Cheapest GPUs at 16-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Memory at 4-bit (weights and margin) 1.3 GB 1.2 GB
Cheapest GPUs at 4-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Revision viewed 97521ec2b6f3 0c0e3051f131
Downloads reported by the hub 58.5k 2.6M
Last observed 2026-09-18 2026-09-21

An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.

SAVRN's Notes on Qwen3-TTS-12Hz-1.7B-CustomVoice

Qwen3-TTS speaks ten major languages, including Chinese, English, Japanese, Korean and the main European languages, with regional voice profiles. The CustomVoice version is built to clone a voice from a short reference sample, and it is by far the most downloaded of the family, at about 2.6 million a month.

It needs under 5 GB at 16-bit. Apache 2.0 allows commercial use. Voice cloning needs a consent process; build that in before you let anyone clone a voice through your product.

Questions

Which is larger, MOSS-VoiceGenerator or Qwen3-TTS-12Hz-1.7B-CustomVoice?

MOSS-VoiceGenerator (2.1B parameters) is larger than Qwen3-TTS-12Hz-1.7B-CustomVoice (1.9B parameters), by the parameter counts their publishers report.

Which is cheaper to run, MOSS-VoiceGenerator or Qwen3-TTS-12Hz-1.7B-CustomVoice?

At 4-bit, MOSS-VoiceGenerator fits on 1x MI300X from $1.85 an hour and Qwen3-TTS-12Hz-1.7B-CustomVoice on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use MOSS-VoiceGenerator commercially?

Yes. MOSS-VoiceGenerator is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Can I use Qwen3-TTS-12Hz-1.7B-CustomVoice commercially?

Yes. Qwen3-TTS-12Hz-1.7B-CustomVoice is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Related Comparisons