[[Paper]](https://arxiv.org/abs/2206.04658) - [[Code]](https://github.com/NVIDIA/BigVGAN) - [[Showcase]](https://bigvgan-demo.github.io/) - [[Project Page]](https://research.nvidia.com/labs/adlr/projects/bigvgan/) - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan-66959df3d97fd7d98d97dc9a) - [[Demo]](https://huggingface.co/spaces/nvidia/BigVGAN) - General refactor and code improvements for improved readability. - Fully fused CUDA kernel of anti-alised activation (upsampling + activation + downsampling) with inference speed benchmark. - We provide pretrained checkpoints of BigVGAN-v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling…
Open weights
mit
PyTorch
[[Paper]](https://arxiv.org/abs/2206.04658) - [[Code]](https://github.com/NVIDIA/BigVGAN) - [[Showcase]](https://bigvgan-demo.github.io/) - [[Project Page]](https://research.nvidia.com/labs/adlr/projects/bigvgan/) - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan-66959df3d97fd7d98d97dc9a) - [[Demo]](https://huggingface.co/spaces/nvidia/BigVGAN) - General refactor and code improvements for improved readability. - Fully fused CUDA kernel of anti-alised activation (upsampling + activation + downsampling) with inference speed benchmark. - We provide pretrained checkpoints of BigVGAN-v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling…
Open weights
mit
PyTorch
Descript Audio Codec running on-device on the LiteRT CompiledModel GPU (ML Drift). The convolutional encoder/decoder run on the GPU; the RVQ runs on CPU. 43:1 compression (1 s → 12×50 codes), RTF ≈ 0.82 (faster than real-time) on Pixel 8a. - dac16khzencoderfp16.tflite (43 MB) — audio[1,1,16000] → latent[1,1024,50], GPU. - dac16khzdeconlyzsfp16.tflite (105 MB) — latent[1,1024,50] → audio, GPU. - dacrvq.bin (1.2 MB) — RVQ weights (12 codebooks) for the CPU quantizer (float32 LE). encoder 367/367 + decoder 398/398 nodes on the LiteRT GPU delegate (LITERTCL, 1 partition, no CPU fallback); warm RTF ~0.82; reconstruction corr 1.0 vs PyTorch DAC. The decoder's ConvTranspose1d are rewritten to a…
Open weights
mit
litert
A small causal Qwen3 model that predicts 32 MOSS v2 audio-codebook indices from the 12 indices of the current frame and up to ten previously generated HQ frames. It is trained from scratch on German, English, Spanish and French speech. This is an experimental enhancement model. Validation cross entropy measures token prediction under teacher forcing. It does not establish an improvement in perceived audio quality. Use the generated-history listening evaluation to judge speech content, speaker identity, artifacts and long-utterance stability. Root inference weights: step 64,699, selected by held-out extra-20-codebook cross entropy 5.818123. - Step 12,940: 100 originals / 200 input cases…
Open weights
cc-by-nc-4.0
pytorch
Synthesise music from MIDI, or add, remove, and modify notes in a recording by revising its MIDI. SpanSynth-Edit generates the selected region, using surrounding audio for timbre guidance. Try it in your browser: upload audio, transcribe with YourMT3+, edit the piano roll, and generate. See the local web setup to run the app yourself. - CPU inference is also supported. Measured peak VRAM was about 3.9 GiB on a GH200 for synthesis and both editing methods with default settings (20.48 s crop, 16 steps, CFG 2.0). Install PyTorch for your GPU first, then install the CLI without downloading the demo audio: Model and codec weights download automatically on first use. No token is required.…
Open weights
apache-2.0
481M parameters
diffusers