SAVRN
Search Contact SAVRN

Open-weight model · Text to speech

VoxCPM2

by OpenBMB openbmb/VoxCPM2

VoxCPM2 is a tokenizer-free, diffusion autoregressive Text-to-Speech model — 2B parameters, 30 languages, 48kHz audio output, trained on over 2 million hours of multilingual speech data.

Parameters2.3B
Context
Weights5.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads385.5k

Runs On

What it takes to serve VoxCPM2 (2.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 4.6 GB 5.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 2.3 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.1 GB 1.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on VoxCPM2

Where does a 2.3 billion parameter text-to-speech model belong in a facility built around 192 GB accelerators? Beside something bigger. VoxCPM2 needs 5.5 GB at 16-bit, 2.7 GB at 8-bit and 1.4 GB at 4-bit, and the cheapest priced slot, one MI300X at $1.85 an hour, has room for many copies. It outputs 48 kHz audio in 30 languages with no language tag, designs a voice from a written description, and clones one from a short clip.

Apache 2.0 keeps the deployment conversation short: commercial use, modification and redistribution are permitted, provided the license, copyright notices and any NOTICE file stay with the files and significant changes are stated. Two checks before committing: the runtime is its own voxcpm library rather than transformers, so plan the serving path around it, and no context length is recorded. The paper is arXiv:2509.24650; the files were last updated 2026-08-18.

Model Card

By OpenBMB, published under apache-2.0, revision 32279effe8c1.

VoxCPM2 is a tokenizer-free, diffusion autoregressive Text-to-Speech model — 2B parameters, 30 languages, 48kHz audio output, trained on over 2 million hours of multilingual speech data.

Highlights

Read the full model card (719 words)

Identity and Version

Repository
openbmb/VoxCPM2
Publisher
OpenBMB
Task
Text to speech
Modality
Audio
Library
voxcpm
Parameters
2.3B parameters
Languages
zh, en, ar, my, da, nl, fi, fr
Revision
32279effe8c19989596f05d353d1447f51d9e915
First published
2026-04-03
Last updated
2026-08-18

Files and Weights

9 files, 5.0 GB in total. The weights are 2 files totalling 5.0 GB in pth, safetensors.

Weights2 files · 5.0 GB
Configuration3 files · 8.9 KB
Tokenizer2 files · 3.7 MB
Documentation1 file · 7.9 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
audiovae.pthWeights377.0 MB 94b5d51e107e
model.safetensorsWeights4.6 GB f7f964cfa9da
config.jsonConfiguration4.3 KB
special_tokens_map.jsonConfiguration1.6 KB
tokenization_voxcpm2.pyConfiguration2.9 KB
README.mdDocumentation7.9 KB
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer3.7 MB
tokenizer_config.jsonTokenizer5.1 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
5.0 GB
Download from OpenBMB

Released by OpenBMB through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2509.24650

Memory Requirements

PrecisionWeights in memory
As published5.0 GB
16-bit4.6 GB
8-bit2.3 GB
4-bit1.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare VoxCPM2

Questions About VoxCPM2

How much GPU memory does VoxCPM2 need?

About 5.5 GB at 16-bit and 1.4 GB at 4-bit: the weights (2.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run VoxCPM2 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use VoxCPM2 commercially?

Yes. VoxCPM2 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to speech

MOSS-VoiceGenerator

OpenMOSS

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS. When a single piece of audio needs to sound like a real person, pronounce every word accurately, switch speaking styles across content, remain stable over tens of minutes, and support dialogue, role‑play, and real‑time interaction, a single TTS model is often not enough. The MOSS‑TTS Family breaks the workflow into five production‑ready models that can be…

Open weights apache-2.0 2.1B parameters 40,960 tokens

Model · Text to speech

Qwen3-TTS-12Hz-1.7B-CustomVoice

Qwen

Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs. In addition, the models feature strong contextual understanding, enabling adaptive control of tone, speaking rate, and emotional expression based on instructions and text semantics, and they show markedly improved robustness to noisy input text. Key features: Intelligent Text Understanding and Voice Control: Supports speech generation driven by natural language instructions, allowing for flexible control over multi-dimensional acoustic attributes such as timbre, emotion, and prosody.…

Open weights apache-2.0 1.9B parameters

Model · Text to speech

Qwen3-TTS-12Hz-1.7B-VoiceDesign

Qwen

We release Qwen3-TTS, a series of powerful speech generation models developed by Qwen, offering comprehensive support for voice cloning, voice design, ultra-high-quality human-like speech generation, and natural language-based voice control. Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles. Key features: Install the qwen-tts Python package from PyPI: Zero-shot speech generation on the Seed-TTS test set (Word Error Rate (WER, ↓)): If you find our paper and code useful in your research, please consider giving a star and citation

Open weights apache-2.0 1.9B parameters qwen-tts

Model · Text to speech

VibeVoice-1.5B

Microsoft

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a Large Language Model (LLM) to understand textual…

Open weights mit 2.7B parameters transformers

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a Large Language Model (LLM) to understand textual…

Open weights mit 2.7B parameters 65,536 tokens transformers

https://github.com/vibevoice-community/VibeVoice VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a…

Open weights mit 2.7B parameters transformers