SAVRN
Search Contact SAVRN

Open-weight model · Text to audio

igbo-mms-tts

by Omeziri Zion nexusbert/igbo-mms-tts

igbo-mms-tts is an open-weight model for text to audio from Omeziri Zion. It has 36M parameters. At 16-bit it needs about 0.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model.

Parameters36M
Context—
Weights145.3 MB
License—
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve igbo-mms-tts (36M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 9, 2026.

igbo-mms-tts on every accelerator the SAVRN Index prices, at every precision

Model Card

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Excerpt from the card by Omeziri Zion.

Configuration

Architecture
VitsModel
Layers
6
Hidden size
192
Attention heads
2
Vocabulary size
74
Stored precision
float32
Model type
vits

Identity and Version

Repository
nexusbert/igbo-mms-tts
Publisher
Omeziri Zion
Task
Text to audio
Modality
Other
Library
transformers
Parameters
36M parameters
Languages
Not stated by the source
Revision
5fa282ec9e91b320c75ef7a3e7852bcf5b4bd19b
First published
2026-10-03
Last updated
2026-10-04

Files and Weights

8 files, 145.3 MB in total. The weights are 1 file totalling 145.3 MB in safetensors.

Weights1 file · 145.3 MB
Configuration3 files · 2.6 KB
Tokenizer2 files · 1.5 KB
Documentation1 file · 5.2 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights145.3 MB f8d287e50374
config.jsonConfiguration2.0 KB —
preprocessor_config.jsonConfiguration254 B —
special_tokens_map.jsonConfiguration275 B —
README.mdDocumentation5.2 KB —
.gitattributesRepository1.5 KB —
tokenizer_config.jsonTokenizer671 B —
vocab.jsonTokenizer848 B —

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
145.3 MB
Download from Omeziri Zion

Released by Omeziri Zion through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published145.3 MB
16-bit0.1 GB
8-bit0.0 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About igbo-mms-tts

How much GPU memory does igbo-mms-tts need?

About 0.1 GB at 16-bit and 0 GB at 4-bit: the weights (36M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run igbo-mms-tts on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Similar Models

Experimental calibration-based weight-only GPTQ variant of The Thinker transformer uses 4-bit weights for layers 0–26 and 8-bit weights for layer 27. Audio, vision, talker, token2wav, embeddings, and norms remain in the original precision. Calibration used eight short text samples with LLM Compressor 0.14.0 and compressed-tensors 0.19.0. This checkpoint includes a packaging repair: the compressor export contained invalid group scales, so scales were recomputed from the original BF16 weights per group before the vLLM test. Treat this as an experimental GPTQ-derived checkpoint and benchmark retrieval quality before production use. Tested with vLLM 0.30.0 on an 8-GiB RTX 3080 Laptop GPU: The…

Open weights 2,935 parameters transformers

Model · Text to audio

musicgen-medium

AI at Meta

MusicGen is a text-to-music model capable of genreating high-quality music samples conditioned on text descriptions or audio prompts. It is a single stage auto-regressive Transformer model trained over a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz. Unlike existing methods, like MusicLM, MusicGen doesn't require a self-supervised semantic representation, and it generates all 4 codebooks in one pass. By introducing a small delay between the codebooks, we show we can predict them in parallel, thus having only 50 auto-regressive steps per second of audio. MusicGen was published in Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant…

Open weights cc-by-nc-4.0 transformers

Model · Text to audio

MiniMax-Music3-GGUF

Audio.cpp

GGUF package for MiniMax Music 3 for audio.cpp. Star our repo so you don't miss important updates! https://github.com/0xShug0/audio.cpp Upstream license: https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE - The implementation is available on the main branch and release 0.6.1. - The current runtime uses model-local resource loading instead of treating the v1 spec as the runtime contract. This keeps component selection flexible while the package layout and option surface settle. - The default component mix favors Q40 for the large language model and flow transformer, with Q80 for the RVQ depth decoder. - BF16, Q80, and Q40 component variants are included for quality/performance…

Open weights other audio.cpp

Model · Text to audio

Yue2-3B-GGUF

Audio.cpp

This repository contains audio.cpp-native GGUF weights for Yue2-3B. The model has been merged into the main branch. LoRA support added in release 0.8.1. The demos below were generated with yue2-3b-bf16.gguf and yue2-vae-f32.gguf. The q8.wav comparison files were generated with yue2-3b-q80.gguf and yue2-vae-f16.gguf using the same prompts and seeds. The q40.wav comparison files were generated with yue2-3b-q40.gguf and yue2-vae-f16.gguf. Replace with your audio.cpp checkout and with this downloaded GGUF repo directory. To generate the Q40 versions, use the same commands and replace --session-option yue2.modelgguf=yue2-3b-bf16.gguf with --session-option yue2.modelgguf=yue2-3b-q40.gguf, and use…

Open weights cc-by-nc-4.0 audio.cpp

Four artist-style LoRAs that push YuE2-3B into modern militant roots reggae: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic. Each file patches both halves of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: mltnt. All demos use the same original lyric, seed 7, 32 steps dpm2 / sgmuniform, no post-processing. MLTNT Frontline — baseline recipe, prompt prompts/steppersbaseline.txt, dense lyric (verses written at ~17 words per…

Open weights cc-by-nc-4.0

Model · Text to audio

akan-twi-mms

Abdul Rashid Dickson

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights cc-by-nc-4.0 83M parameters transformers