SAVRN
Search Contact SAVRN

Open-weight model · Text to speech

VibeVoice-Realtime-0.5B

by Microsoft microsoft/VibeVoice-Realtime-0.5B

VibeVoice-Realtime is a lightweight real‑time text-to-speech model supporting streaming text input and robust long-form speech generation.

Parameters1B
Context
Weights2.0 GB
Licensemit
AccessOpen weights
Monthly Downloads367k

Runs On

What it takes to serve VibeVoice-Realtime-0.5B (1B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 2.0 GB 2.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.0 GB 1.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.5 GB 0.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Microsoft, published under mit, revision 6bce5f060448.

VibeVoice: A Frontier Open-Source Text-to-Speech Model

VibeVoice-Realtime is a lightweight real‑time text-to-speech model supporting streaming text input and robust long-form speech generation. It can be used to build realtime TTS services, narrate live data streams, and let different LLMs start speaking from their very first tokens (plug in your preferred model) long before a full answer is generated. It produces initial audible speech in ~300 ms (hardware dependent).

▶ Watch demo video (Launch your own realtime demo via the websocket example in Usage)

Although the model is primarily built for English, we found that it still exhibits a certain level of multilingual capability—and even performs reasonably well in some languages. We provide nine additional languages (German, French, Italian, Japanese, Korean, Dutch, Polish, Portuguese, and Spanish) for users to explore and share feedback.

Read the full model card (1,174 words)

Configuration

Architecture
VibeVoiceStreamingForConditionalGenerationInference
Stored precision
bfloat16
Model type
vibevoice_streaming

Identity and Version

Repository
microsoft/VibeVoice-Realtime-0.5B
Publisher
Microsoft
Task
Text to speech
Modality
Audio
Library
transformers
Parameters
1B parameters
Languages
en
Revision
6bce5f06044837fe6d2c5d7a71a84f0416bd57e4
First published
2025-12-04
Last updated
2025-12-12

Files and Weights

6 files, 2.0 GB in total. The weights are 1 file totalling 2.0 GB in safetensors.

Weights1 file · 2.0 GB
Configuration2 files · 2.5 KB
Documentation1 file · 10.2 KB
Other1 file · 123.5 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights2.0 GB 7758b150b813
config.jsonConfiguration2.1 KB
preprocessor_config.jsonConfiguration360 B
README.mdDocumentation10.2 KB
figures/Fig1.pngOther123.5 KB 0386a7f577a6
.gitattributesRepository1.6 KB

License and Download

License
mit
Access
Open weights, no gate
Download size
2.0 GB
Download from Microsoft

Released by Microsoft through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published2.0 GB
16-bit2.0 GB
8-bit1.0 GB
4-bit0.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About VibeVoice-Realtime-0.5B

How much GPU memory does VibeVoice-Realtime-0.5B need?

About 2.4 GB at 16-bit and 0.6 GB at 4-bit: the weights (1B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run VibeVoice-Realtime-0.5B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use VibeVoice-Realtime-0.5B commercially?

Yes. VibeVoice-Realtime-0.5B is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Model · Text to speech

indic-parler-tts

AI4Bharat

Indic Parler-TTS is a multilingual Indic extension of Parler-TTS Mini. It is a fine-tuned version of Indic Parler-TTS Pretrained, trained on a 1,806 hours multilingual Indic and English dataset. Indic Parler-TTS Mini can officially speak in 20 Indic languages, making it comprehensive for regional language technologies, and in English. The 21 languages supported are: Assamese, Bengali, Bodo, Dogri, English, Gujarati, Hindi, Kannada, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Odia, Sanskrit, Santali, Sindhi, Tamil, Telugu, and Urdu. Thanks to its better prompt tokenizer, it can easily be extended to other languages. This tokenizer has a larger vocabulary and handles byte…

Access requested at publisher apache-2.0 938M parameters transformers

Model · Text to speech

Qwen3-TTS-12Hz-0.6B-Base

Qwen

Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control. This specific checkpoint is the 0.6B Base model, which is capable of rapid voice cloning from a user-provided audio input. To clone a voice and synthesize new content using the Base model, you can use the following code snippet: Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application…

Open weights apache-2.0 915M parameters

Model · Text to speech

Qwen3-TTS-12Hz-0.6B-CustomVoice

Qwen

Qwen3-TTS is a series of advanced multilingual, controllable, robust, and streaming text-to-speech models developed by the Qwen team. This specific checkpoint is the 0.6B CustomVoice variant, based on the 12Hz tokenizer. It supports 9 premium timbres and allows for fine-grained style control over target voices via natural language instructions across 10 major languages. To use Qwen3-TTS, you can install the qwen-tts package: For Qwen3-TTS-12Hz-0.6B-CustomVoice, the following speakers are supported. We recommend using each speaker’s native language for the best results: If you find Qwen3-TTS useful for your research, please consider citing

Open weights apache-2.0 906M parameters

Model · Text to speech

OmniVoice

K2 FSA

OmniVoice is a massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model-style architecture, it delivers high-quality speech with superior inference speed, supporting voice cloning and voice design. - 600+ Languages Supported: The broadest language coverage among zero-shot TTS models. To get started, install the omnivoice library: You can use OmniVoice for zero-shot voice cloning as follows: For more generation modes (e.g., voice design), functions (e.g., non-verbal symbols, pronunciation correction) and comprehensive usage instructions, see our GitHub Repository. You can directly discuss on GitHub Issues. You can…

Open weights 613M parameters 40,960 tokens omnivoice

Model · Text to speech

gepard-1.0

NineNineSix

GEnerative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue Gepard is a text-to-speech model built for real-time conversation. It starts speaking the moment text begins arriving, generating audio piece by piece instead of waiting for a full sentence — so it feels like a live voice, not a recording. It's a single language model that learned text and speech together, so the output carries natural rhythm and timing rather than the flat, stitched tone of older pipelines. The name evokes "Gepard"(/geh-PART/), German for cheetah — a nod to the model's low-latency, high-throughput streaming. - One clean pass per frame — the whole audio frame (32 orthogonal FSQ channels) is…

Open weights apache-2.0 556M parameters 262,144 tokens transformers

Model · Text to speech

csm-1b

Sesame

2025/05/20 - CSM is availabile natively in Hugging Face Transformers as of version 4.52.1 2025/03/13 - We are releasing the 1B CSM variant. The checkpoint is hosted on Hugging Face. CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs. The model architecture employs a Llama backbone and a smaller audio decoder that produces Mimi audio codes. A fine-tuned variant of CSM powers the interactive voice demo shown in our blog post. A hosted HuggingFace space is also available for testing audio generation. CSM supports full-graph compilation with CUDA graphs! CSM can be fine-tuned using Transformers' Trainer. Does this…

Access requested at publisher apache-2.0 1.6B parameters transformers