SAVRN
Search Contact SAVRN

Open-weight model · Audio classification

audiobox-aesthetics

by AI at Meta facebook/audiobox-aesthetics

This model has been pushed to the Hub using the PytorchModelHubMixin integration: Unified automatic quality assessment for speech, music, and sound. Paper arXiv / MetaAI. Blogpost ai.meta.com This repository requires Python 3.9 and Pytorch 2.2 or greater.

Parameters104M
Context
Weights831.0 MB
Licensecc-by-4.0
AccessOpen weights
Monthly Downloads635.2k

Runs On

What it takes to serve audiobox-aesthetics (104M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on audiobox-aesthetics

When a pipeline produces or ingests audio at volume and someone needs a score for how each clip sounds, this 104M-parameter model from AI at Meta returns one across speech, music and sound in a single pass. It is small enough to disappear into a rack: 0.2 GB of weights and 0.2 GB of memory at 16-bit, so the cheapest Index option, one MI300X with 192 GB at $1.85 an hour, is priced for the generator beside it, not for this. The repository is 6 files and 831 MB, safetensors only, and the code wants Python 3.9 and PyTorch 2.2 or newer.

CC BY 4.0 permits commercial use and adaptation with two obligations: credit the creator and state what you changed. Read the paper it is described by, arXiv:2502.05139, because the page lists no architecture, library or reported evaluations. Last update was March 4, 2025, so treat it as fixed.

Model Card

By AI at Meta, published under cc-by-4.0, revision 9b1dd8e5df9a.

This model has been pushed to the Hub using the PytorchModelHubMixin integration: - Code: https://github.com/facebookresearch/audiobox-aesthetics - Paper: https://huggingface.co/papers/2502.05139

--- README below copied from https://github.com/facebookresearch/audiobox-aesthetics

audiobox-aesthetics

Unified automatic quality assessment for speech, music, and sound.

Installation

  1. Install via pip pip install audiobox_aesthetics

  2. Install directly from source

This repository requires Python 3.9 and Pytorch 2.2 or greater. To install, you can clone this repo and run: pip install -e .

Pre-trained Models

Model | S3 | HuggingFace |---|---|---| All axes | checkpoint.pt | HF Repo

Usage

How to run prediction using CLI:

Read the full model card (577 words)

Identity and Version

Repository
facebook/audiobox-aesthetics
Publisher
AI at Meta
Task
Audio classification
Modality
Audio
Library
Not stated by the source
Parameters
104M parameters
Languages
Not stated by the source
Revision
9b1dd8e5df9af7216e836a98974fe3b82c56ded6
First published
2025-02-13
Last updated
2025-03-04

Files and Weights

6 files, 831.4 MB in total. The weights are 2 files totalling 831.0 MB in pt, safetensors.

Weights2 files · 831.0 MB
Configuration1 file · 492 B
Documentation1 file · 5.9 KB
Other1 file · 389.1 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
checkpoint.ptWeights415.5 MB a4931a7a01c3
model.safetensorsWeights415.5 MB a5a3c2412649
config.jsonConfiguration492 B
README.mdDocumentation5.9 KB
assets/aes_model.pngOther389.1 KB bcc57fff5a71
.gitattributesRepository1.6 KB

License and Download

License
cc-by-4.0
Access
Open weights, no gate
Download size
831.0 MB
Download from AI at Meta

Released by AI at Meta through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2502.05139

Memory Requirements

PrecisionWeights in memory
As published831.0 MB
16-bit0.2 GB
8-bit0.1 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare audiobox-aesthetics

Questions About audiobox-aesthetics

How much GPU memory does audiobox-aesthetics need?

About 0.2 GB at 16-bit and 0.1 GB at 4-bit: the weights (104M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run audiobox-aesthetics on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use audiobox-aesthetics commercially?

Yes. audiobox-aesthetics is released under Creative Commons Attribution 4.0. CC BY 4.0 permits sharing and adapting the work, including commercially, provided the creator is credited and changes are indicated.

Similar Models

Model · Audio classification

wav2vec2-base-drum-kit

Andrew Keig

Fine-tuned facebook/wav2vec2-base for audio classification of single drum/percussion sounds into 10 classes. - clap, conga, crash, cymbal, hat, kick, ride, rim, snare, tom - Trained on short, single-hit drum sounds. Performance may drop on long mixes, multiple overlapping sounds, or very different recording conditions.

Open weights mit 95M parameters

Model · Audio classification

wav2vec2-emotion-recognition

Deepan Gautam

This model is a fine-tuned version of facebook/wav2vec2-base-960h for Speech Emotion Recognition (SER). It has been trained using a Frozen Feature Extractor strategy to preserve the model's acoustic understanding while adapting to emotion detection. This approach ensures stable performance and prevents "Catastrophic Forgetting," achieving nearly 80% accuracy on the validation set. Update: The "Calm" and "Neutral" classes have been merged to improve classification consistency, resulting in 7 distinct emotion classes. The model was trained on a combined dataset of ~12,000 audio files from: The model classifies audio into one of the following emotions: 1. Angry 2. Disgust 3. Fear 4. Happy 5.…

Open weights mit 95M parameters transformers

Model · Audio classification

Common-Voice-Gender-Detection

Prithiv Sakthi

Wav2Vec2: Self-Supervised Learning for Speech Recognition: https://arxiv.org/pdf/2006.11477 male female Common-Voice-Gender-Detection is designed for: Speech Analytics – Assist in analyzing speaker demographics in call centers or customer service recordings. Conversational AI Personalization – Adjust tone or dialogue based on gender detection for more personalized voice assistants. Voice Dataset Curation – Automatically tag or filter voice datasets by speaker gender for better dataset management. Research Applications – Enable linguistic and acoustic research involving gender-specific speech patterns. Multimedia Content Tagging – Automate metadata generation for gender identification in…

Open weights apache-2.0 95M parameters transformers

Model · Audio classification

Deepfake-audio-detection

Mohammed Abdeldayem

This model is a fine-tuned version of mo-thecreator/wav2vec2-base-finetuned on the None dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 3e-05 - trainbatchsize: 8 - evalbatchsize: 8 - gradientaccumulationsteps: 4 - totaltrainbatchsize: 32 - lrschedulertype: linear - lrschedulerwarmupratio: 0.1 - numepochs: 5 - Transformers 4.39.3 - Pytorch 2.1.2 - Datasets 2.18.0 - Tokenizers 0.15.2 - mo-thecreator

Open weights apache-2.0 95M parameters transformers

Model · Audio classification

Deepfake-audio-detection-V2

Melody Machine

This model is a fine-tuned version of motheecreator/Deepfake-audio-detection on the audiofolder dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 3e-05 - trainbatchsize: 32 - evalbatchsize: 32 - gradientaccumulationsteps: 4 - totaltrainbatchsize: 128 - lrschedulertype: cosine - lrschedulerwarmupratio: 0.1 - numepochs: 5 - Transformers 4.41.2 - Pytorch 2.1.2 - Datasets 2.19.2 - Tokenizers 0.19.1

Open weights apache-2.0 95M parameters transformers

The model expects a raw audio signal as input and outputs predictions for age in a range of approximately 0...1 (0...100 years) and gender expressing the probababilty for being child, female, or male. In addition, it also provides the pooled states of the last transformer layer. The model was created by fine-tuning Wav2Vec2-Large-Robust Timit and For this version of the model we only trained the first six transformer layers. An ONNX export of the model is available from Further details are given in the associated paper and tutorial.

Open weights cc-by-nc-sa-4.0 91M parameters transformers