SAVRN
Search Contact SAVRN

Open-weight model · Text to audio

akan-twi-mms

by Abdul Rashid Dickson Dickson32-cell/akan-twi-mms

akan-twi-mms is an open-weight model for text to audio from Abdul Rashid Dickson, released under Creative Commons Attribution-NonCommercial 4.0. It has 83M parameters. At 16-bit it needs about 0.2 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 44 downloads a month.

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model.

Parameters83M
Context—
Weights332.2 MB
Licensecc-by-nc-4.0
AccessOpen weights
Monthly Downloads44

Runs On

What it takes to serve akan-twi-mms (83M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 9, 2026.

akan-twi-mms on every accelerator the SAVRN Index prices, at every precision

Model Card

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Excerpt from the card by Abdul Rashid Dickson, licensed cc-by-nc-4.0.

Configuration

Architecture
VitsModel
Layers
6
Hidden size
192
Attention heads
2
Vocabulary size
30
Stored precision
float32
Model type
vits

Identity and Version

Repository
Dickson32-cell/akan-twi-mms
Publisher
Abdul Rashid Dickson
Task
Text to audio
Modality
Other
Library
transformers
Parameters
83M parameters
Languages
ak
Revision
2e52c366dd30f51d92f7f068a5cd690e762843e3
First published
2026-10-04
Last updated
2026-10-09

Files and Weights

15 files, 337.9 MB in total. The weights are 1 file totalling 332.2 MB in safetensors.

Weights1 file · 332.2 MB
Configuration7 files · 3.2 KB
Tokenizer2 files · 614 B
Documentation1 file · 5.7 KB
Other3 files · 5.7 MB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights332.2 MB 61f09c350ba6
added_tokens.jsonConfiguration18 B —
config.jsonConfiguration1.6 KB —
convergence_summary.jsonConfiguration47 B —
preprocessor_config.jsonConfiguration254 B —
special_tokens_map.jsonConfiguration47 B —
three_way_results.jsonConfiguration499 B —
three_way_results_v2.jsonConfiguration663 B —
README.mdDocumentation5.7 KB —
loss_curve.pngOther255.1 KB 34cc7144fb26
losses.csvOther130.9 KB —
train_akan.logOther5.3 MB —
.gitattributesRepository1.6 KB —
tokenizer_config.jsonTokenizer287 B —
vocab.jsonTokenizer327 B —

License and Download

License
cc-by-nc-4.0
Access
Open weights, no gate
Download size
332.2 MB
Download from Abdul Rashid Dickson

Released by Abdul Rashid Dickson through its official repository on Hugging Face. Read the license.

Built From

  • Derived from facebook/mms-tts-aka
  • Described by arXiv:1910.09700
  • Trained on (disclosed) ghananlpcommunity/ghana-speech

Memory Requirements

PrecisionWeights in memory
As published332.2 MB
16-bit0.2 GB
8-bit0.1 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About akan-twi-mms

How much GPU memory does akan-twi-mms need?

About 0.2 GB at 16-bit and 0 GB at 4-bit: the weights (83M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run akan-twi-mms on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use akan-twi-mms commercially?

Not without separate permission. akan-twi-mms is released under Creative Commons Attribution-NonCommercial 4.0. CC BY-NC 4.0 permits sharing and adapting with credit for non-commercial purposes only. Commercial use needs separate permission from the rights holder.

Similar Models

Model · Text to audio

igbo-mms-tts

Omeziri Zion

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 36M parameters transformers

Experimental calibration-based weight-only GPTQ variant of The Thinker transformer uses 4-bit weights for layers 0–26 and 8-bit weights for layer 27. Audio, vision, talker, token2wav, embeddings, and norms remain in the original precision. Calibration used eight short text samples with LLM Compressor 0.14.0 and compressed-tensors 0.19.0. This checkpoint includes a packaging repair: the compressor export contained invalid group scales, so scales were recomputed from the original BF16 weights per group before the vLLM test. Treat this as an experimental GPTQ-derived checkpoint and benchmark retrieval quality before production use. Tested with vLLM 0.30.0 on an 8-GiB RTX 3080 Laptop GPU: The…

Open weights 2,935 parameters transformers

Model · Text to audio

musicgen-medium

AI at Meta

MusicGen is a text-to-music model capable of genreating high-quality music samples conditioned on text descriptions or audio prompts. It is a single stage auto-regressive Transformer model trained over a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz. Unlike existing methods, like MusicLM, MusicGen doesn't require a self-supervised semantic representation, and it generates all 4 codebooks in one pass. By introducing a small delay between the codebooks, we show we can predict them in parallel, thus having only 50 auto-regressive steps per second of audio. MusicGen was published in Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant…

Open weights cc-by-nc-4.0 transformers

Model · Text to audio

MiniMax-Music3-GGUF

Audio.cpp

GGUF package for MiniMax Music 3 for audio.cpp. Star our repo so you don't miss important updates! https://github.com/0xShug0/audio.cpp Upstream license: https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE - The implementation is available on the main branch and release 0.6.1. - The current runtime uses model-local resource loading instead of treating the v1 spec as the runtime contract. This keeps component selection flexible while the package layout and option surface settle. - The default component mix favors Q40 for the large language model and flow transformer, with Q80 for the RVQ depth decoder. - BF16, Q80, and Q40 component variants are included for quality/performance…

Open weights other audio.cpp

Model · Text to audio

Yue2-3B-GGUF

Audio.cpp

This repository contains audio.cpp-native GGUF weights for Yue2-3B. The model has been merged into the main branch. LoRA support added in release 0.8.1. The demos below were generated with yue2-3b-bf16.gguf and yue2-vae-f32.gguf. The q8.wav comparison files were generated with yue2-3b-q80.gguf and yue2-vae-f16.gguf using the same prompts and seeds. The q40.wav comparison files were generated with yue2-3b-q40.gguf and yue2-vae-f16.gguf. Replace with your audio.cpp checkout and with this downloaded GGUF repo directory. To generate the Q40 versions, use the same commands and replace --session-option yue2.modelgguf=yue2-3b-bf16.gguf with --session-option yue2.modelgguf=yue2-3b-q40.gguf, and use…

Open weights cc-by-nc-4.0 audio.cpp

Four artist-style LoRAs that push YuE2-3B into modern militant roots reggae: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic. Each file patches both halves of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: mltnt. All demos use the same original lyric, seed 7, 32 steps dpm2 / sgmuniform, no post-processing. MLTNT Frontline — baseline recipe, prompt prompts/steppersbaseline.txt, dense lyric (verses written at ~17 words per…

Open weights cc-by-nc-4.0