SAVRN
Search Contact SAVRN

Open-weight model · Text to audio

musicgen-medium

by AI at Meta facebook/musicgen-medium

MusicGen is a text-to-music model capable of genreating high-quality music samples conditioned on text descriptions or audio prompts.

Parameters
Context
Weights12.0 GB
Licensecc-by-nc-4.0
AccessOpen weights
Monthly Downloads2M

Model Card

MusicGen is a text-to-music model capable of genreating high-quality music samples conditioned on text descriptions or audio prompts. It is a single stage auto-regressive Transformer model trained over a 32kHz EnCodec tokenizer with 4 codebooks sampled at 50 Hz. Unlike existing methods, like MusicLM, MusicGen doesn't require a self-supervised semantic representation, and it generates all 4 codebooks in one pass. By introducing a small delay between the codebooks, we show we can predict them in parallel, thus having only 50 auto-regressive steps per second of audio. MusicGen was published in Simple and Controllable Music Generation by Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant…

Excerpt from the card by AI at Meta, licensed cc-by-nc-4.0.

Configuration

Architecture
MusicgenForConditionalGeneration
Stored precision
float32
Model type
musicgen

Identity and Version

Repository
facebook/musicgen-medium
Publisher
AI at Meta
Task
Text to audio
Modality
Other
Library
transformers
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
d3bd7b00761b78ad7a8a05145ee31e7832e9916c
First published
2023-06-08
Last updated
2023-11-17

Files and Weights

12 files, 12.0 GB in total. The weights are 3 files totalling 12.0 GB in bin.

Weights3 files · 12.0 GB
Configuration4 files · 10.6 KB
Tokenizer3 files · 3.2 MB
Documentation1 file · 12.3 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
compression_state_dict.binWeights236.0 MB 37d256b525d4
pytorch_model.binWeights8.0 GB b51b0ec25096
state_dict.binWeights3.7 GB fbc9f401db2a
config.jsonConfiguration7.9 KB
generation_config.jsonConfiguration224 B
preprocessor_config.jsonConfiguration275 B
special_tokens_map.jsonConfiguration2.2 KB
README.mdDocumentation12.3 KB
.gitattributesRepository1.5 KB
spiece.modelTokenizer791.7 KB d60acb128cf7
tokenizer.jsonTokenizer2.4 MB
tokenizer_config.jsonTokenizer2.4 KB

License and Download

License
cc-by-nc-4.0
Access
Open weights, no gate
Download size
12.0 GB
Download from AI at Meta

Released by AI at Meta through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2306.05284

Memory Requirements

PrecisionWeights in memory
As published12.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About musicgen-medium

Can I use musicgen-medium commercially?

Not without separate permission. musicgen-medium is released under Creative Commons Attribution-NonCommercial 4.0. CC BY-NC 4.0 permits sharing and adapting with credit for non-commercial purposes only. Commercial use needs separate permission from the rights holder.

Similar Models

Four artist-style LoRAs that push YuE2-3B into modern militant roots reggae: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic. Each file patches both halves of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: mltnt. All demos use the same original lyric, seed 7, 32 steps dpm2 / sgmuniform, no post-processing. MLTNT Frontline — baseline recipe, prompt prompts/steppersbaseline.txt, dense lyric (verses written at ~17 words per…

Open weights cc-by-nc-4.0