MusicGen is a text-to-music model capable of genreating high-quality music samples conditioned on text descriptions or audio prompts. It is a single stage auto-regressive Transformer model trained over a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz. Unlike existing methods, like MusicLM, MusicGen doesn't require a self-supervised semantic representation, and it generates all 4 codebooks in one pass. By introducing a small delay between the codebooks, we show we can predict them in parallel, thus having only 50 auto-regressive steps per second of audio. MusicGen was published in Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant…
Open weights
cc-by-nc-4.0
transformers
This repository contains audio.cpp-native GGUF weights for Yue2-3B. The model has been merged into the main branch. LoRA support added in release 0.8.1. The demos below were generated with yue2-3b-bf16.gguf and yue2-vae-f32.gguf. The q8.wav comparison files were generated with yue2-3b-q80.gguf and yue2-vae-f16.gguf using the same prompts and seeds. The q40.wav comparison files were generated with yue2-3b-q40.gguf and yue2-vae-f16.gguf. Replace with your audio.cpp checkout and with this downloaded GGUF repo directory. To generate the Q40 versions, use the same commands and replace --session-option yue2.modelgguf=yue2-3b-bf16.gguf with --session-option yue2.modelgguf=yue2-3b-q40.gguf, and use…
Open weights
cc-by-nc-4.0
audio.cpp
Four artist-style LoRAs that push YuE2-3B into modern militant roots reggae: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic. Each file patches both halves of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: mltnt. All demos use the same original lyric, seed 7, 32 steps dpm2 / sgmuniform, no post-processing. MLTNT Frontline — baseline recipe, prompt prompts/steppersbaseline.txt, dense lyric (verses written at ~17 words per…
Open weights
cc-by-nc-4.0
Experimental calibration-based weight-only GPTQ variant of The Thinker transformer uses 4-bit weights for layers 0–26 and 8-bit weights for layer 27. Audio, vision, talker, token2wav, embeddings, and norms remain in the original precision. Calibration used eight short text samples with LLM Compressor 0.14.0 and compressed-tensors 0.19.0. This checkpoint includes a packaging repair: the compressor export contained invalid group scales, so scales were recomputed from the original BF16 weights per group before the vLLM test. Treat this as an experimental GPTQ-derived checkpoint and benchmark retrieval quality before production use. Tested with vLLM 0.30.0 on an 8-GiB RTX 3080 Laptop GPU: The…
Open weights
2,935 parameters
transformers
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Open weights
36M parameters
transformers
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Open weights
cc-by-nc-4.0
83M parameters
transformers