SAVRN Model Hub · Comparisons
MOSS-VoiceGenerator vs Qwen3-TTS-12Hz-1.7B-CustomVoice
MOSS-VoiceGenerator has 2.1B parameters and Qwen3-TTS-12Hz-1.7B-CustomVoice has 1.9B parameters; both are released under Apache License 2.0; at 16-bit, MOSS-VoiceGenerator needs about 5.1 GB (1x MI300X from $1.85 an hour) and Qwen3-TTS-12Hz-1.7B-CustomVoice about 4.6 GB (1x MI300X from $1.85 an hour).
| Field | MOSS-VoiceGenerator OpenMOSS-Team/MOSS-VoiceGenerator | Qwen3-TTS-12Hz-1.7B-CustomVoice Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice |
|---|---|---|
| Publisher | OpenMOSS | Qwen |
| Task | Text to speech | Text to speech |
| Modality | Audio | Audio |
| Parameters, as reported | 2.1B parameters | 1.9B parameters |
| Architecture | MossTTSDelayModel | Qwen3TTSForConditionalGeneration |
| Library | Not stated | Not stated |
| Context length | 40,960 tokens | Not stated |
| Repository size | 4.2 GB | 4.5 GB |
| Artifact formats | safetensors | safetensors |
| License | apache-2.0 | apache-2.0 |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 5.1 GB | 4.6 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Memory at 4-bit (weights and margin) | 1.3 GB | 1.2 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | 97521ec2b6f3 | 0c0e3051f131 |
| Downloads reported by the hub | 58.5k | 2.6M |
| Last observed | 2026-09-18 | 2026-09-21 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
SAVRN's Notes on Qwen3-TTS-12Hz-1.7B-CustomVoice
Qwen3-TTS speaks ten major languages, including Chinese, English, Japanese, Korean and the main European languages, with regional voice profiles. The CustomVoice version is built to clone a voice from a short reference sample, and it is by far the most downloaded of the family, at about 2.6 million a month.
It needs under 5 GB at 16-bit. Apache 2.0 allows commercial use. Voice cloning needs a consent process; build that in before you let anyone clone a voice through your product.
Questions
Which is larger, MOSS-VoiceGenerator or Qwen3-TTS-12Hz-1.7B-CustomVoice?
MOSS-VoiceGenerator (2.1B parameters) is larger than Qwen3-TTS-12Hz-1.7B-CustomVoice (1.9B parameters), by the parameter counts their publishers report.
Which is cheaper to run, MOSS-VoiceGenerator or Qwen3-TTS-12Hz-1.7B-CustomVoice?
At 4-bit, MOSS-VoiceGenerator fits on 1x MI300X from $1.85 an hour and Qwen3-TTS-12Hz-1.7B-CustomVoice on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use MOSS-VoiceGenerator commercially?
Yes. MOSS-VoiceGenerator is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Can I use Qwen3-TTS-12Hz-1.7B-CustomVoice commercially?
Yes. Qwen3-TTS-12Hz-1.7B-CustomVoice is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.