SAVRN
Search Contact SAVRN

Open-weight model · Audio to audio

ArkEcho-RVC-M3-C-F

by Ethernos Ethernos/ArkEcho-RVC-M3-C-F

ArkEcho-RVC-M3-C-F is an open-weight model for audio to audio from Ethernos, released under Creative Commons Attribution-NonCommercial 4.0. Its published files total 95.2 MB.

本仓库(Repository)所包含的所有人工神经网络(Artificial Neural Network)权重文件(.pth/.index)、训练日志及相关代码,均为计算声学(Computational Acoustics)与深度学习(Deep Learning)领域的技术研究实验产物。 这些文件本质上是高维张量(High-dimensional Tensors)的数值序列,通过随机梯度下降(SGD)与反向传播算法(Backpropagation)对公开可获取的音频数据进行统计建模(Statistical…

Parameters—
Context—
Weights55.2 MB
Licensecc-by-nc-4.0
AccessOpen weights
Monthly Downloads—

Model Card

本仓库(Repository)所包含的所有人工神经网络(Artificial Neural Network)权重文件(.pth/.index)、训练日志及相关代码,均为计算声学(Computational Acoustics)与深度学习(Deep Learning)领域的技术研究实验产物。 这些文件本质上是高维张量(High-dimensional Tensors)的数值序列,通过随机梯度下降(SGD)与反向传播算法(Backpropagation)对公开可获取的音频数据进行统计建模(Statistical Modeling)得到。其技术形态与数字图像处理中的卷积核(Convolution Kernels)、自然语言处理中的词向量(Word Embeddings)并无本质差异。 本仓库不构成对任何第三方知识产权的故意侵犯,所有代码遵循 MIT License 开源协议,模型权重文件仅作为技术实现的副产品(By-products)存在。 本技术实验所使用的训练数据集(Training Dataset)包含以下角色的语音样本: - Mon3tr(モンスター)、Mon3tr 中文语音 等 - 声源版权归属:上海鹰角网络科技有限公司(Hypergryph Network Technology Co., Ltd.)及其关联公司 - 原始作品:《明日方舟》(Arknights) 上述角色的声音版权、肖像权、姓名权及相关知识产权均完全归属于鹰角网络及其合法授权方。本实验仅基于已公开发布的游戏内语音资源进行技术层面的信号处理(Signal Processing)与特征提取(Feature…

Excerpt from the card by Ethernos, licensed cc-by-nc-4.0.

Identity and Version

Repository
Ethernos/ArkEcho-RVC-M3-C-F
Publisher
Ethernos
Task
Audio to audio
Modality
Audio
Library
Not stated by the source
Parameters
Not stated by the source
Languages
zh
Revision
b02267602a63131746178b7caa3375a5f8286dd4
First published
2026-09-26
Last updated
2026-09-26

Files and Weights

4 files, 95.2 MB in total. The weights are 1 file totalling 55.2 MB in pth.

Weights1 file · 55.2 MB
Documentation1 file · 6.1 KB
Other1 file · 40.0 MB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
ArkEcho-RVC-M3-C-F.pthWeights55.2 MB 9796cd3ef286
README.mdDocumentation6.1 KB —
Index-ArkEcho-RVC-M3-C-F.indexOther40.0 MB 29388ffbdf1f
.gitattributesRepository1.6 KB —

License and Download

License
cc-by-nc-4.0
Access
Open weights, no gate
Download size
55.2 MB
Download from Ethernos

Released by Ethernos through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published55.2 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ArkEcho-RVC-M3-C-F

Can I use ArkEcho-RVC-M3-C-F commercially?

Not without separate permission. ArkEcho-RVC-M3-C-F is released under Creative Commons Attribution-NonCommercial 4.0. CC BY-NC 4.0 permits sharing and adapting with credit for non-commercial purposes only. Commercial use needs separate permission from the rights holder.

Similar Models

Model · Audio to audio

bigvgan_v2_22khz_80band_256x

NVIDIA

[[Paper]](https://arxiv.org/abs/2206.04658) - [[Code]](https://github.com/NVIDIA/BigVGAN) - [[Showcase]](https://bigvgan-demo.github.io/) - [[Project Page]](https://research.nvidia.com/labs/adlr/projects/bigvgan/) - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan-66959df3d97fd7d98d97dc9a) - [[Demo]](https://huggingface.co/spaces/nvidia/BigVGAN) - General refactor and code improvements for improved readability. - Fully fused CUDA kernel of anti-alised activation (upsampling + activation + downsampling) with inference speed benchmark. - We provide pretrained checkpoints of BigVGAN-v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling…

Open weights mit PyTorch

Model · Audio to audio

bigvgan_v2_44khz_128band_512x

NVIDIA

[[Paper]](https://arxiv.org/abs/2206.04658) - [[Code]](https://github.com/NVIDIA/BigVGAN) - [[Showcase]](https://bigvgan-demo.github.io/) - [[Project Page]](https://research.nvidia.com/labs/adlr/projects/bigvgan/) - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan-66959df3d97fd7d98d97dc9a) - [[Demo]](https://huggingface.co/spaces/nvidia/BigVGAN) - General refactor and code improvements for improved readability. - Fully fused CUDA kernel of anti-alised activation (upsampling + activation + downsampling) with inference speed benchmark. - We provide pretrained checkpoints of BigVGAN-v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling…

Open weights mit PyTorch

Descript Audio Codec running on-device on the LiteRT CompiledModel GPU (ML Drift). The convolutional encoder/decoder run on the GPU; the RVQ runs on CPU. 43:1 compression (1 s → 12×50 codes), RTF ≈ 0.82 (faster than real-time) on Pixel 8a. - dac16khzencoderfp16.tflite (43 MB) — audio[1,1,16000] → latent[1,1024,50], GPU. - dac16khzdeconlyzsfp16.tflite (105 MB) — latent[1,1024,50] → audio, GPU. - dacrvq.bin (1.2 MB) — RVQ weights (12 codebooks) for the CPU quantizer (float32 LE). encoder 367/367 + decoder 398/398 nodes on the LiteRT GPU delegate (LITERTCL, 1 partition, no CPU fallback); warm RTF ~0.82; reconstruction corr 1.0 vs PyTorch DAC. The decoder's ConvTranspose1d are rewritten to a…

Open weights mit litert

Model · Audio to audio

MOSS-v2-12to32-Enhancer-60M

LAION eV

A small causal Qwen3 model that predicts 32 MOSS v2 audio-codebook indices from the 12 indices of the current frame and up to ten previously generated HQ frames. It is trained from scratch on German, English, Spanish and French speech. This is an experimental enhancement model. Validation cross entropy measures token prediction under teacher forcing. It does not establish an improvement in perceived audio quality. Use the generated-history listening evaluation to judge speech content, speaker identity, artifacts and long-utterance stability. Root inference weights: step 64,699, selected by held-out extra-20-codebook cross entropy 5.818123. - Step 12,940: 100 originals / 200 input cases…

Open weights cc-by-nc-4.0 pytorch

Model · Audio to audio

spansynth-edit

Sungkyun Chang

Synthesise music from MIDI, or add, remove, and modify notes in a recording by revising its MIDI. SpanSynth-Edit generates the selected region, using surrounding audio for timbre guidance. Try it in your browser: upload audio, transcribe with YourMT3+, edit the piano roll, and generate. See the local web setup to run the app yourself. - CPU inference is also supported. Measured peak VRAM was about 3.9 GiB on a GH200 for synthesis and both editing methods with default settings (20.48 s crop, 16 steps, CFG 2.0). Install PyTorch for your GPU first, then install the CLI without downloading the demo audio: Model and codec weights download automatically on first use. No token is required.…

Open weights apache-2.0 481M parameters diffusers