SAVRN
Search Contact SAVRN

Open-weight model · Text to audio

Yue2-3B-GGUF

by Audio.cpp audio-cpp/Yue2-3B-GGUF

Yue2-3B-GGUF is an open-weight model for text to audio from Audio.cpp, released under Creative Commons Attribution-NonCommercial 4.0. Its published files total 17.6 GB. It draws 166.3k downloads a month.

This repository contains audio.cpp-native GGUF weights for Yue2-3B. The model has been merged into the main branch. LoRA support added in release 0.8.1. The demos below were generated with yue2-3b-bf16.gguf and yue2-vae-f32.gguf.

Parameters—
Context—
Weights17.3 GB
Licensecc-by-nc-4.0
AccessOpen weights
Monthly Downloads166.3k

Model Card

This repository contains audio.cpp-native GGUF weights for Yue2-3B. The model has been merged into the main branch. LoRA support added in release 0.8.1. The demos below were generated with yue2-3b-bf16.gguf and yue2-vae-f32.gguf. The q8.wav comparison files were generated with yue2-3b-q80.gguf and yue2-vae-f16.gguf using the same prompts and seeds. The q40.wav comparison files were generated with yue2-3b-q40.gguf and yue2-vae-f16.gguf. Replace with your audio.cpp checkout and with this downloaded GGUF repo directory. To generate the Q40 versions, use the same commands and replace --session-option yue2.modelgguf=yue2-3b-bf16.gguf with --session-option yue2.modelgguf=yue2-3b-q40.gguf, and use…

Excerpt from the card by Audio.cpp, licensed cc-by-nc-4.0.

Identity and Version

Repository
audio-cpp/Yue2-3B-GGUF
Publisher
Audio.cpp
Task
Text to audio
Modality
Other
Library
audio.cpp
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
cca74d8980c00f30ab8d52248a367c01ad633ec2
First published
2026-09-10
Last updated
2026-10-06

Files and Weights

34 files, 17.6 GB in total. The weights are 6 files totalling 17.3 GB in gguf.

Weights6 files · 17.3 GB
Configuration4 files · 6.2 KB
Tokenizer1 file · 2.6 MB
Documentation2 files · 12.8 KB
Other20 files · 278.4 MB
Repository1 file · 2.9 KB
Every file
FileTypeSizeSHA-256
yue2-3b-bf16.ggufWeights7.3 GB 4cf3fdffbaae
yue2-3b-ios-q4_0.ggufWeights2.3 GB 83a2dbc57f8d
yue2-3b-q4_0.ggufWeights2.7 GB b80215b59a7f
yue2-3b-q8_0.ggufWeights4.3 GB 52be345b5155
yue2-vae-f16.ggufWeights265.2 MB d4f4a05d8f29
yue2-vae-f32.ggufWeights530.5 MB cf3b71c6ef13
lora/training-metrics.jsonConfiguration3.4 KB —
sidecars/yue2-generation-config.jsonConfiguration466 B —
sidecars/yue2-model-config.jsonConfiguration959 B —
sidecars/yue2-vae-config.jsonConfiguration1.4 KB —
README.mdDocumentation10.0 KB —
lora/README.mdDocumentation2.9 KB —
examples/demo1.wavOther7.7 MB 00d3947c1d07
examples/demo1_q4_0.wavOther7.2 MB 0d2a4fbe6902
examples/demo1_q8.wavOther7.6 MB 777f1089c166
examples/demo2.wavOther7.1 MB 96b2f509b407
examples/demo2_q4_0.wavOther7.6 MB c1c31107eb57
examples/demo2_q8.wavOther7.9 MB d4f96dfa8077
examples/demo3.wavOther7.4 MB 9152fa7b8a90
examples/demo3_q4_0.wavOther4.2 MB f5e8901cc4e3
examples/demo3_q8.wavOther4.2 MB b28eeda64ba3
examples/demo4.wavOther4.2 MB 8ac632779605
examples/demo4_q4_0.wavOther4.3 MB aef2bb37d9f2
examples/demo4_q8.wavOther10.7 MB ddd388a0ed0d
examples/tonight-awake-audiocpp-current.wavOther43.2 MB 64cd582e51d3
examples/tonight-awake-audiocpp-current_q4_0.wavOther42.5 MB c83e7b1d0b55
examples/tonight-awake-audiocpp-current_q8.wavOther37.4 MB 45e9b6c97dbe
examples/yue2_demo.mp4Other958.6 KB 268a1245cd3f
lora/SHA256SUMSOther435 B —
lora/samples/harbor-keeps-the-dawn-cot-off.wavOther74.3 MB 188e91700abf
lora/samples/harbor-keeps-the-dawn-lyrics.txtOther1.9 KB —
lora/samples/harbor-keeps-the-dawn-style.txtOther156 B —
.gitattributesRepository2.9 KB —
sidecars/yue2-qwen.tiktokenTokenizer2.6 MB —

License and Download

License
cc-by-nc-4.0
Access
Open weights, no gate
Download size
17.3 GB
Download from Audio.cpp

Released by Audio.cpp through its official repository on Hugging Face. Read the license.

Built From

  • Derived from m-a-p/YuE2-3B
  • Quantized from m-a-p/YuE2-3B

Memory Requirements

PrecisionWeights in memory
As published17.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Yue2-3B-GGUF

Can I use Yue2-3B-GGUF commercially?

Not without separate permission. Yue2-3B-GGUF is released under Creative Commons Attribution-NonCommercial 4.0. CC BY-NC 4.0 permits sharing and adapting with credit for non-commercial purposes only. Commercial use needs separate permission from the rights holder.

Similar Models

Model · Text to audio

musicgen-medium

AI at Meta

MusicGen is a text-to-music model capable of genreating high-quality music samples conditioned on text descriptions or audio prompts. It is a single stage auto-regressive Transformer model trained over a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz. Unlike existing methods, like MusicLM, MusicGen doesn't require a self-supervised semantic representation, and it generates all 4 codebooks in one pass. By introducing a small delay between the codebooks, we show we can predict them in parallel, thus having only 50 auto-regressive steps per second of audio. MusicGen was published in Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant…

Open weights cc-by-nc-4.0 transformers

Model · Text to audio

MiniMax-Music3-GGUF

Audio.cpp

GGUF package for MiniMax Music 3 for audio.cpp. Star our repo so you don't miss important updates! https://github.com/0xShug0/audio.cpp Upstream license: https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE - The implementation is available on the main branch and release 0.6.1. - The current runtime uses model-local resource loading instead of treating the v1 spec as the runtime contract. This keeps component selection flexible while the package layout and option surface settle. - The default component mix favors Q40 for the large language model and flow transformer, with Q80 for the RVQ depth decoder. - BF16, Q80, and Q40 component variants are included for quality/performance…

Open weights other audio.cpp

Four artist-style LoRAs that push YuE2-3B into modern militant roots reggae: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic. Each file patches both halves of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: mltnt. All demos use the same original lyric, seed 7, 32 steps dpm2 / sgmuniform, no post-processing. MLTNT Frontline — baseline recipe, prompt prompts/steppersbaseline.txt, dense lyric (verses written at ~17 words per…

Open weights cc-by-nc-4.0

Experimental calibration-based weight-only GPTQ variant of The Thinker transformer uses 4-bit weights for layers 0–26 and 8-bit weights for layer 27. Audio, vision, talker, token2wav, embeddings, and norms remain in the original precision. Calibration used eight short text samples with LLM Compressor 0.14.0 and compressed-tensors 0.19.0. This checkpoint includes a packaging repair: the compressor export contained invalid group scales, so scales were recomputed from the original BF16 weights per group before the vLLM test. Treat this as an experimental GPTQ-derived checkpoint and benchmark retrieval quality before production use. Tested with vLLM 0.30.0 on an 8-GiB RTX 3080 Laptop GPU: The…

Open weights 2,935 parameters transformers

Model · Text to audio

igbo-mms-tts

Omeziri Zion

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 36M parameters transformers

Model · Text to audio

akan-twi-mms

Abdul Rashid Dickson

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights cc-by-nc-4.0 83M parameters transformers