SAVRN
Search Contact SAVRN

Open-weight model · Text to audio

MiniMax-Music3-GGUF

by Audio.cpp audio-cpp/MiniMax-Music3-GGUF

MiniMax-Music3-GGUF is an open-weight model for text to audio from Audio.cpp, released under other. Its published files total 58.1 GB. It draws 579.7k downloads a month.

GGUF package for MiniMax Music 3 for audio.cpp. Star our repo so you don't miss important updates!

Parameters—
Context—
Weights58.1 GB
Licenseother
AccessOpen weights
Monthly Downloads579.7k

Model Card

GGUF package for MiniMax Music 3 for audio.cpp. Star our repo so you don't miss important updates! https://github.com/0xShug0/audio.cpp Upstream license: https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE - The implementation is available on the main branch and release 0.6.1. - The current runtime uses model-local resource loading instead of treating the v1 spec as the runtime contract. This keeps component selection flexible while the package layout and option surface settle. - The default component mix favors Q40 for the large language model and flow transformer, with Q80 for the RVQ depth decoder. - BF16, Q80, and Q40 component variants are included for quality/performance…

Excerpt from the card by Audio.cpp, licensed other.

Configuration

Architecture
MiniMaxMusic3ForConditionalGeneration
Model type
minimax_music3

Identity and Version

Repository
audio-cpp/MiniMax-Music3-GGUF
Publisher
Audio.cpp
Task
Text to audio
Modality
Other
Library
audio.cpp
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
46fe77f8bb395ccdbe797d784225f7b84866e7ce
First published
2026-08-14
Last updated
2026-10-06

Files and Weights

24 files, 58.1 GB in total. The weights are 13 files totalling 58.1 GB in gguf.

Weights13 files · 58.1 GB
Configuration6 files · 2.8 KB
Tokenizer2 files · 11.4 MB
Documentation2 files · 10.9 KB
Repository1 file · 2.4 KB
Every file
FileTypeSizeSHA-256
condition_encoder.ggufWeights100.7 MB ae4be8b61546
language_model_bf16.ggufWeights17.2 GB 61692fa58e77
language_model_q4_0.ggufWeights6.0 GB 6f621dd63632
language_model_q4_k.ggufWeights7.2 GB 019300ae45ae
language_model_q8_0.ggufWeights9.9 GB cf8a9b01bb0e
rvq_depth_decoder_bf16.ggufWeights1.3 GB 986ca48662fc
rvq_depth_decoder_q4_k.ggufWeights405.8 MB 4c5d41b27418
rvq_depth_decoder_q8_0.ggufWeights714.0 MB d5b7495fb1a7
transformer_bf16.ggufWeights9.7 GB bd6f9cd59568
transformer_q4_0.ggufWeights1.4 GB 18d3b2461a29
transformer_q4_k.ggufWeights1.4 GB 5ab8a79d251d
transformer_q8_0.ggufWeights2.6 GB 40ed8a14884e
vocoder.ggufWeights216.7 MB 5d40e6d52349
config.jsonConfiguration107 B —
config/condition_encoder.jsonConfiguration292 B —
config/language_model.jsonConfiguration1.6 KB —
config/rvq_depth_decoder.jsonConfiguration274 B —
config/transformer.jsonConfiguration294 B —
config/vocoder.jsonConfiguration251 B —
LICENSEDocumentation7.4 KB —
README.mdDocumentation3.6 KB —
.gitattributesRepository2.4 KB —
tokenizer/tokenizer.jsonTokenizer11.4 MB b1537fa9e59a
tokenizer/tokenizer_config.jsonTokenizer377 B —

License and Download

License
other
Access
Open weights, no gate
Download size
58.1 GB
Download from Audio.cpp

Released by Audio.cpp through its official repository on Hugging Face.

Built From

  • Derived from MiniMaxAI/MiniMax-Music3
  • Quantized from MiniMaxAI/MiniMax-Music3

Memory Requirements

PrecisionWeights in memory
As published58.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About MiniMax-Music3-GGUF

What license is MiniMax-Music3-GGUF released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Text to audio

musicgen-medium

AI at Meta

MusicGen is a text-to-music model capable of genreating high-quality music samples conditioned on text descriptions or audio prompts. It is a single stage auto-regressive Transformer model trained over a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz. Unlike existing methods, like MusicLM, MusicGen doesn't require a self-supervised semantic representation, and it generates all 4 codebooks in one pass. By introducing a small delay between the codebooks, we show we can predict them in parallel, thus having only 50 auto-regressive steps per second of audio. MusicGen was published in Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant…

Open weights cc-by-nc-4.0 transformers

Model · Text to audio

Yue2-3B-GGUF

Audio.cpp

This repository contains audio.cpp-native GGUF weights for Yue2-3B. The model has been merged into the main branch. LoRA support added in release 0.8.1. The demos below were generated with yue2-3b-bf16.gguf and yue2-vae-f32.gguf. The q8.wav comparison files were generated with yue2-3b-q80.gguf and yue2-vae-f16.gguf using the same prompts and seeds. The q40.wav comparison files were generated with yue2-3b-q40.gguf and yue2-vae-f16.gguf. Replace with your audio.cpp checkout and with this downloaded GGUF repo directory. To generate the Q40 versions, use the same commands and replace --session-option yue2.modelgguf=yue2-3b-bf16.gguf with --session-option yue2.modelgguf=yue2-3b-q40.gguf, and use…

Open weights cc-by-nc-4.0 audio.cpp

Four artist-style LoRAs that push YuE2-3B into modern militant roots reggae: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic. Each file patches both halves of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: mltnt. All demos use the same original lyric, seed 7, 32 steps dpm2 / sgmuniform, no post-processing. MLTNT Frontline — baseline recipe, prompt prompts/steppersbaseline.txt, dense lyric (verses written at ~17 words per…

Open weights cc-by-nc-4.0

Experimental calibration-based weight-only GPTQ variant of The Thinker transformer uses 4-bit weights for layers 0–26 and 8-bit weights for layer 27. Audio, vision, talker, token2wav, embeddings, and norms remain in the original precision. Calibration used eight short text samples with LLM Compressor 0.14.0 and compressed-tensors 0.19.0. This checkpoint includes a packaging repair: the compressor export contained invalid group scales, so scales were recomputed from the original BF16 weights per group before the vLLM test. Treat this as an experimental GPTQ-derived checkpoint and benchmark retrieval quality before production use. Tested with vLLM 0.30.0 on an 8-GiB RTX 3080 Laptop GPU: The…

Open weights 2,935 parameters transformers

Model · Text to audio

igbo-mms-tts

Omeziri Zion

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 36M parameters transformers

Model · Text to audio

akan-twi-mms

Abdul Rashid Dickson

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights cc-by-nc-4.0 83M parameters transformers