SAVRN
Search Contact SAVRN

Open-weight model · Audio to audio

DAC-16kHz-LiteRT

by LiteRT Community (FKA TFLite) litert-community/DAC-16kHz-LiteRT

DAC-16kHz-LiteRT is an open-weight model for audio to audio from LiteRT Community (FKA TFLite), released under MIT License. Its published files total 149.5 MB.

Descript Audio Codec running on-device on the LiteRT CompiledModel GPU (ML Drift). The convolutional encoder/decoder run on the GPU; the RVQ runs on CPU. 43:1 compression (1 s → 12×50 codes), RTF ≈ 0.82 (faster than real-time) on Pixel 8a.

Parameters—
Context—
Weights149.5 MB
Licensemit
AccessOpen weights
Monthly Downloads—

Model Card

By LiteRT Community (FKA TFLite), published under mit, revision 07e52bf19210.

Descript Audio Codec running on-device on the LiteRT CompiledModel GPU (ML Drift). The convolutional encoder/decoder run on the GPU; the RVQ runs on CPU. 43:1 compression (1 s → 12×50 codes), RTF ≈ 0.82 (faster than real-time) on Pixel 8a. - dac16khzencoderfp16.tflite (43 MB) — audio[1,1,16000] → latent[1,1024,50], GPU. - dac16khzdeconlyzsfp16.tflite (105 MB) — latent[1,1024,50] → audio, GPU. - dacrvq.bin (1.2 MB) — RVQ weights (12 codebooks) for the CPU quantizer (float32 LE). encoder 367/367 + decoder 398/398 nodes on the LiteRT GPU delegate (LITERTCL, 1 partition, no CPU fallback); warm RTF ~0.82; reconstruction corr 1.0 vs PyTorch DAC. The decoder's ConvTranspose1d are rewritten to a…

Read LiteRT Community (FKA TFLite)'s full model card

DAC (Descript Audio Codec) 16 kHz — LiteRT (CompiledModel GPU)

Descript Audio Codec running on-device on the LiteRT CompiledModel GPU (ML Drift). The convolutional encoder/decoder run on the GPU; the RVQ runs on CPU. 43:1 compression (1 s → 12×50 codes), RTF ≈ 0.82 (faster than real-time) on Pixel 8a.

Files

  • dac_16khz_encoder_fp16.tflite (43 MB) — audio[1,1,16000] → latent[1,1024,50], GPU.
  • dac_16khz_deconly_zs_fp16.tflite (105 MB) — latent[1,1024,50] → audio, GPU.
  • dac_rvq.bin (1.2 MB) — RVQ weights (12 codebooks) for the CPU quantizer (float32 LE).

Pipeline

audio -> encoder.tflite (GPU) -> z -> RVQ.encode (CPU) -> codes[12,50]
      -> RVQ.decode (CPU) -> z_q -> decoder.tflite (GPU) -> audio

On-device (Pixel 8a, Tensor G3 — verified)

encoder 367/367 + decoder 398/398 nodes on the LiteRT GPU delegate (LITERT_CL, 1 partition, no CPU fallback); warm RTF ~0.82; reconstruction corr 1.0 vs PyTorch DAC.

Why the split

The decoder's ConvTranspose1d are rewritten to a GPU-clean zero-stuff form (the real DAC's odd stride-5 transposed conv fails converter legalization, and TRANSPOSE_CONV is rejected by Mali). The RVQ uses EMBEDDING_LOOKUP + int64 indices (Mali-rejected) so it runs on CPU. So the float conv graph stays fully on the GPU.

Android sample + conversion/validation scripts: https://github.com/john-rocky/LiteRT-Models/tree/main/dac

License: MIT (Descript DAC).

Identity and Version

Repository
litert-community/DAC-16kHz-LiteRT
Publisher
LiteRT Community (FKA TFLite)
Task
Audio to audio
Modality
Audio
Library
litert
Parameters
Not stated by the source
Languages
dac, on-device, gpu
Revision
07e52bf19210ee172f180c898563f13dfc8c6e40
First published
2026-09-22
Last updated
2026-09-22

Files and Weights

5 files, 149.5 MB in total. The weights are 3 files totalling 149.5 MB in bin, tflite.

Weights3 files · 149.5 MB
Documentation1 file · 1.7 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
dac_16khz_deconly_zs_fp16.tfliteWeights105.0 MB 052ebafe8f4a
dac_16khz_encoder_fp16.tfliteWeights43.2 MB b32e86c67085
dac_rvq.binWeights1.2 MB d3e20e3ec3d6
README.mdDocumentation1.7 KB —
.gitattributesRepository1.5 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
149.5 MB
Download from LiteRT Community (FKA TFLite)

Released by LiteRT Community (FKA TFLite) through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published149.5 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DAC-16kHz-LiteRT

Can I use DAC-16kHz-LiteRT commercially?

Yes. DAC-16kHz-LiteRT is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Model · Audio to audio

bigvgan_v2_22khz_80band_256x

NVIDIA

[[Paper]](https://arxiv.org/abs/2206.04658) - [[Code]](https://github.com/NVIDIA/BigVGAN) - [[Showcase]](https://bigvgan-demo.github.io/) - [[Project Page]](https://research.nvidia.com/labs/adlr/projects/bigvgan/) - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan-66959df3d97fd7d98d97dc9a) - [[Demo]](https://huggingface.co/spaces/nvidia/BigVGAN) - General refactor and code improvements for improved readability. - Fully fused CUDA kernel of anti-alised activation (upsampling + activation + downsampling) with inference speed benchmark. - We provide pretrained checkpoints of BigVGAN-v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling…

Open weights mit PyTorch

Model · Audio to audio

bigvgan_v2_44khz_128band_512x

NVIDIA

[[Paper]](https://arxiv.org/abs/2206.04658) - [[Code]](https://github.com/NVIDIA/BigVGAN) - [[Showcase]](https://bigvgan-demo.github.io/) - [[Project Page]](https://research.nvidia.com/labs/adlr/projects/bigvgan/) - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan-66959df3d97fd7d98d97dc9a) - [[Demo]](https://huggingface.co/spaces/nvidia/BigVGAN) - General refactor and code improvements for improved readability. - Fully fused CUDA kernel of anti-alised activation (upsampling + activation + downsampling) with inference speed benchmark. - We provide pretrained checkpoints of BigVGAN-v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling…

Open weights mit PyTorch

Model · Audio to audio

ArkEcho-RVC-M3-C-F

Ethernos

本仓库(Repository)所包含的所有人工神经网络(Artificial Neural Network)权重文件(.pth/.index)、训练日志及相关代码,均为计算声学(Computational Acoustics)与深度学习(Deep Learning)领域的技术研究实验产物。 这些文件本质上是高维张量(High-dimensional Tensors)的数值序列,通过随机梯度下降(SGD)与反向传播算法(Backpropagation)对公开可获取的音频数据进行统计建模(Statistical Modeling)得到。其技术形态与数字图像处理中的卷积核(Convolution Kernels)、自然语言处理中的词向量(Word Embeddings)并无本质差异。 本仓库不构成对任何第三方知识产权的故意侵犯,所有代码遵循 MIT License 开源协议,模型权重文件仅作为技术实现的副产品(By-products)存在。 本技术实验所使用的训练数据集(Training Dataset)包含以下角色的语音样本: - Mon3tr(モンスター)、Mon3tr 中文语音 等 - 声源版权归属:上海鹰角网络科技有限公司(Hypergryph Network Technology Co., Ltd.)及其关联公司 - 原始作品:《明日方舟》(Arknights) 上述角色的声音版权、肖像权、姓名权及相关知识产权均完全归属于鹰角网络及其合法授权方。本实验仅基于已公开发布的游戏内语音资源进行技术层面的信号处理(Signal Processing)与特征提取(Feature…

Open weights cc-by-nc-4.0

Model · Audio to audio

MOSS-v2-12to32-Enhancer-60M

LAION eV

A small causal Qwen3 model that predicts 32 MOSS v2 audio-codebook indices from the 12 indices of the current frame and up to ten previously generated HQ frames. It is trained from scratch on German, English, Spanish and French speech. This is an experimental enhancement model. Validation cross entropy measures token prediction under teacher forcing. It does not establish an improvement in perceived audio quality. Use the generated-history listening evaluation to judge speech content, speaker identity, artifacts and long-utterance stability. Root inference weights: step 64,699, selected by held-out extra-20-codebook cross entropy 5.818123. - Step 12,940: 100 originals / 200 input cases…

Open weights cc-by-nc-4.0 pytorch

Model · Audio to audio

spansynth-edit

Sungkyun Chang

Synthesise music from MIDI, or add, remove, and modify notes in a recording by revising its MIDI. SpanSynth-Edit generates the selected region, using surrounding audio for timbre guidance. Try it in your browser: upload audio, transcribe with YourMT3+, edit the piano roll, and generate. See the local web setup to run the app yourself. - CPU inference is also supported. Measured peak VRAM was about 3.9 GiB on a GH200 for synthesis and both editing methods with default settings (20.48 s crop, 16 steps, CFG 2.0). Install PyTorch for your GPU first, then install the CLI without downloading the demo audio: Model and codec weights download automatically on first use. No token is required.…

Open weights apache-2.0 481M parameters diffusers