SAVRN
Search Contact SAVRN

Open-weight model · Text to speech

svara-tts-v1

by Kenpath Labs kenpath/svara-tts-v1

svara-TTS is a developer-first multilingual TTS model for 19 languages (18 Indic + Indian English). Built on an Orpheus-style discrete audio token approach, it targets clarity, expressiveness, and low-latency on commodity GPUs/CPUs.

Parameters3.3B
Context131,072
Weights13.2 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads79.4k

Runs On

What it takes to serve svara-tts-v1 (3.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 6.6 GB 7.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 3.3 GB 4.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.7 GB 2.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Kenpath Labs, published under apache-2.0, revision db8a02fc1e4e.

svara-TTS is a developer-first multilingual TTS model for 19 languages (18 Indic + Indian English). Built on an Orpheus-style discrete audio token approach, it targets clarity, expressiveness, and low-latency on commodity GPUs/CPUs. It supports light-weight emotion/style control (e.g.,,,, ) and simple speaker identities (Language (Gender)), with zero-shot adaptation paths. Try it live on the Demo Space, or on Colab Deployment scripts and inference repo will be available soon. Watch our Github for updates - Place style/emotion tags at the end of the sentence: आज... सच में अच्छी खबर है — शाम को मिलते हैं! - Use punctuation to hint prosody (ellipses, commas, exclamation). - For technical or…

Read Kenpath Labs's full model card

svara-TTS v1 — Open Multilingual TTS for India’s Voices

svara-TTS is a developer-first multilingual TTS model for 19 languages (18 Indic + Indian English).
Built on an Orpheus-style discrete audio token approach, it targets clarity, expressiveness, and low-latency on commodity GPUs/CPUs.
It supports light-weight emotion/style control (e.g., <happy>, <sad>, <anger>, <fear>) and simple speaker identities (Language (Gender)), with zero-shot adaptation paths.


At a Glance

  • Languages (19): Hindi, Bengali, Marathi, Telugu, Kannada, Bhojpuri, Magahi, Chhattisgarhi, Maithili, Assamese, Bodo, Dogri, Gujarati, Malayalam, Punjabi, Tamil, Nepali, Sanskrit, Indian English.
  • Expressivity: End-of-utterance style tags; natural prosody; code-switch aware.
  • Latency & Deployment: Works well with GGUF exports; suitable for edge/CPU scenarios.
  • Adaptability: LoRA-friendly for quick speaker/domain specialization.

Try it live on the Demo Space, or on Colab Deployment scripts and inference repo will be available soon. Watch our Github for updates


Prompting (Orpheus-style)

  • Place style/emotion tags at the end of the sentence:
    आज... सच में अच्छी खबर है — शाम को मिलते हैं! <happy>
  • Use punctuation to hint prosody (ellipses, commas, exclamation).
  • For technical or dense text, end with <clear> to prioritize intelligibility.

Speaker IDs follow a simple convention: Language (Gender) (e.g., Marathi (Male)).


Training Data Summary

Trained on 2000+ hours of open, high-quality speech from SYSPIN, RASA, IndicTTS, and SPICOR, covering ~50 speakers (balanced male/female) across 19 languages.
Data was curated to encourage natural prosody, broad coverage, and stable multilingual transfer. See Acknowledgments for provenance.


Intended Uses

  • Multilingual assistants, IVR, learning apps, reading aids, accessibility tools
  • Content localization (education, public-information, civic services)
  • Research on Indic prosody, emotion control, cross-lingual transfer

Out-of-Scope / Not Intended

  • Impersonation of private individuals or public figures without consent
  • Deceptive content (fraud, harassment, misinformation)
  • Safety-critical deployments without human oversight

Limitations

  • Proper nouns & rare entities: may require spelling hints or <clear>.
  • Very long sentences: chunk or add punctuation for natural prosody.
  • Emotion strength: varies by language due to data density.
  • Code-mixing: common patterns work; it’s not a deterministic rules engine.

Many of these improve with targeted LoRA finetuning and better preprocessing.


Responsible Use

By using this model, you agree to follow applicable laws and ethical guidelines.
Avoid impersonation, harassment, targeted deception, or other harmful uses.
Where appropriate, disclose synthetic speech to end users.


Sources & Links

  • Model: https://huggingface.co/kenpath/svara-tts-v1
  • Demo Space: https://huggingface.co/spaces/kenpath/svara-tts
  • Inference repo: https://github.com/Kenpath/svara-tts-inference
  • Colab: https://colab.research.google.com/drive/15YxFo1DzdQNbFUIZ1HJA4AN4oHqKxGtg

Acknowledgments

This work was developed by Kenpath Technologies for the open-source community. We also thank RunPod for the startup credits that supported our GPU compute.

  • Canopy Labs — Orpheus: foundational ideas & open release
    Release: https://canopylabs.ai/releases/orpheus_can_speak_any_language
  • SPIRE Lab, IISc BangaloreSYSPIN (multilingual studio) and SPICOR (Indian English)
  • AI4BharatRASA expressive speech
  • IIT MadrasIndicTTS
  • Unsloth — helpful notes & tooling
  • RunPod — startup GPU credits that accelerated experiments

License

Apache-2.0


Versioning & Changelog

  • v1.0.0: Initial public release (19 languages)

Configuration

Architecture
LlamaForCausalLM
Context length (tokens)
131,072
Layers
28
Hidden size
3,072
Feed-forward size
8,192
Attention heads
24
Key/value heads
8
Head dimension
128
Vocabulary size
156,940
RoPE base
500000
Stored precision
bfloat16
Model type
llama

Identity and Version

Repository
kenpath/svara-tts-v1
Publisher
Kenpath Labs
Task
Text to speech
Modality
Audio
Library
transformers
Parameters
3.3B parameters
Languages
hi, bn, mr, te, kn, bho, mag, hne
Revision
db8a02fc1e4eab827ff6dda5bed3b56d4d2dd51e
First published
2025-10-26
Last updated
2025-10-27

Files and Weights

23 files, 13.2 GB in total. The weights are 3 files totalling 13.2 GB in safetensors.

Weights3 files · 13.2 GB
Configuration3 files · 22.2 KB
Tokenizer2 files · 28.3 MB
Documentation1 file · 5.7 KB
Other13 files · 2.3 MB
Repository1 file · 3.4 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00003.safetensorsWeights4.9 GB 5697a861fecd
model-00002-of-00003.safetensorsWeights4.9 GB a736acff0d88
model-00003-of-00003.safetensorsWeights3.3 GB 0602d21d09d8
config.jsonConfiguration899 B
model.safetensors.index.jsonConfiguration20.9 KB
special_tokens_map.jsonConfiguration385 B
README.mdDocumentation5.7 KB
chat_template.jinjaOther3.8 KB
examples/anger_english.wavOther118.8 KB 1406739bd12c
examples/chat_magahi.wavOther163.9 KB b1461d953a00
examples/clear_telugu.wavOther213.0 KB 6f684246f7b3
examples/fear_maithili.wavOther204.8 KB f1b5be7821b2
examples/formal_kannada.wavOther254.0 KB 1212437c71f4
examples/happy_hindi.wavOther131.1 KB 62ec49a637cf
examples/neutral_nepali.wavOther163.9 KB 6a4f2f7e80ba
examples/neutral_sanskrit.wavOther249.9 KB 764135d7c77e
examples/sad_malayalam.wavOther200.7 KB 302de9f6b986
examples/sad_marathi.wavOther270.4 KB 667bf2706d49
examples/surprise_punjabi.wavOther225.3 KB 627d1b187d18
examples/surprise_tamil.wavOther143.4 KB 4b89c7181e7a
.gitattributesRepository3.4 KB
tokenizer.jsonTokenizer22.8 MB 044e2a102017
tokenizer_config.jsonTokenizer5.4 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
13.2 GB
Download from Kenpath Labs

Released by Kenpath Labs through its official repository on Hugging Face. Read the license.

Built From

  • Adapter of canopylabs/3b-hi-ft-research_release
  • Derived from canopylabs/3b-hi-ft-research_release
  • Trained on (disclosed) IndicTTS
  • Trained on (disclosed) RASA
  • Trained on (disclosed) SPICOR
  • Trained on (disclosed) SYSPIN

Memory Requirements

PrecisionWeights in memory
As published13.2 GB
16-bit6.6 GB
8-bit3.3 GB
4-bit1.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About svara-tts-v1

How much GPU memory does svara-tts-v1 need?

About 7.9 GB at 16-bit and 2 GB at 4-bit: the weights (3.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run svara-tts-v1 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use svara-tts-v1 commercially?

Yes. svara-tts-v1 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is svara-tts-v1's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text to speech

orpheus-3b-0.1-ft

Unsloth AI

03/18/2025 – We are releasing our 3B Orpheus TTS model with additional finetunes. Code is available on GitHub: CanopyAI/Orpheus-TTS Orpheus TTS is a state-of-the-art, Llama-based Speech-LLM designed for high-quality, empathetic text-to-speech generation. This model has been finetuned to deliver human-level speech synthesis, achieving exceptional clarity, expressiveness, and real-time streaming performances. Check out our Colab (link to Colab) or GitHub (link to GitHub) on how to run easy inference on our finetuned models. Do not use our models for impersonation without consent, misinformation or deception (including fake news or fraudulent calls), or any illegal or harmful activity. By…

Open weights apache-2.0 3.3B parameters 131,072 tokens transformers

Model · Text to speech

orpheus-3b-0.1-ft

Canopy Labs

03/18/2025 – We are releasing our 3B Orpheus TTS model with additional finetunes. Code is available on GitHub: CanopyAI/Orpheus-TTS Orpheus TTS is a state-of-the-art, Llama-based Speech-LLM designed for high-quality, empathetic text-to-speech generation. This model has been finetuned to deliver human-level speech synthesis, achieving exceptional clarity, expressiveness, and real-time streaming performances. Check out our Colab (link to Colab) or GitHub (link to GitHub) on how to run easy inference on our finetuned models. Do not use our models for impersonation without consent, misinformation or deception (including fake news or fraudulent calls), or any illegal or harmful activity. By…

Access requested at publisher apache-2.0 3.8B parameters transformers

Model · Text to speech

VibeVoice-1.5B

Microsoft

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a Large Language Model (LLM) to understand textual…

Open weights mit 2.7B parameters transformers

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a Large Language Model (LLM) to understand textual…

Open weights mit 2.7B parameters 65,536 tokens transformers

https://github.com/vibevoice-community/VibeVoice VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. A core innovation of VibeVoice is its use of continuous speech tokenizers (Acoustic and Semantic) operating at an ultra-low frame rate of 7.5 Hz. These tokenizers efficiently preserve audio fidelity while significantly boosting computational efficiency for processing long sequences. VibeVoice employs a next-token diffusion framework, leveraging a…

Open weights mit 2.7B parameters transformers

Model · Text to speech

VoxCPM2

OpenBMB

VoxCPM2 is a tokenizer-free, diffusion autoregressive Text-to-Speech model — 2B parameters, 30 languages, 48kHz audio output, trained on over 2 million hours of multilingual speech data. - 30-Language Multilingual — No language tag needed; input text in any supported language directly - Voice Design — Generate a novel voice from a natural-language description alone (gender, age, tone, emotion, pace…); no reference audio required - Controllable Cloning — Clone any voice from a short clip, with optional style guidance to steer emotion, pace, and expression while preserving timbre - Ultimate Cloning — Provide reference audio + its transcript for audio-continuation cloning; every vocal nuance…

Open weights apache-2.0 2.3B parameters voxcpm