SAVRN
Search Contact SAVRN

Open-weight model · Audio to audio

spansynth-edit

by Sungkyun Chang mimbres/spansynth-edit

spansynth-edit is an open-weight model for audio to audio from Sungkyun Chang, released under Apache License 2.0. It has 481M parameters. At 16-bit it needs about 1.2 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Synthesise music from MIDI, or add, remove, and modify notes in a recording by revising its MIDI. SpanSynth-Edit generates the selected region, using surrounding audio for timbre guidance.

Parameters481M
Context—
Weights5.2 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve spansynth-edit (481M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1.0 GB 1.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.5 GB 0.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 2, 2026.

spansynth-edit on every accelerator the SAVRN Index prices, at every precision

Model Card

By Sungkyun Chang, published under apache-2.0, revision 5f0a7e3aca72.

Synthesise music from MIDI, or add, remove, and modify notes in a recording by revising its MIDI. SpanSynth-Edit generates the selected region, using surrounding audio for timbre guidance. Try it in your browser: upload audio, transcribe with YourMT3+, edit the piano roll, and generate. See the local web setup to run the app yourself. - CPU inference is also supported. Measured peak VRAM was about 3.9 GiB on a GH200 for synthesis and both editing methods with default settings (20.48 s crop, 16 steps, CFG 2.0). Install PyTorch for your GPU first, then install the CLI without downloading the demo audio: Model and codec weights download automatically on first use. No token is required.…

Read Sungkyun Chang's full model card

MIDI-guided synthesis and editing of multi-instrument audio mixtures

Synthesise music from MIDI, or add, remove, and modify notes in a recording by revising its MIDI. SpanSynth-Edit generates the selected region, using surrounding audio for timbre guidance.

Try it in your browser: upload audio, transcribe with YourMT3+, edit the piano roll, and generate. See the local web setup to run the app yourself.

Install  ·  Examples  ·  Options  ·  Timing & MIDI  ·  Results  ·  Citation

Install

System requirements

  • Python: 3.11–3.13 with PyTorch.
  • NVIDIA GPU: CUDA with bfloat16 support. 6 GB VRAM recommended.
  • Apple Silicon GPU: supported through Metal (PyTorch MPS). Tested on an M1 Pro with 16 GB unified memory and PyTorch 2.13.
  • CPU inference is also supported.

Measured peak VRAM was about 3.9 GiB on a GH200 for synthesis and both editing methods with default settings (20.48 s crop, 16 steps, CFG 2.0).

Install PyTorch for your GPU first, then install the CLI without downloading the demo audio:

git clone --depth 1 --filter=blob:none --sparse https://github.com/mimbres/spansynth-edit.git
cd spansynth-edit
git sparse-checkout set spansynth
python -m pip install .

Model and codec weights download automatically on first use. No token is required.

Diffusers

Install the optional integration:

python -m pip install '.[diffusers]'

Download the example files, then edit the selected region:

import soundfile as sf
from diffusers import DiffusionPipeline

pipe = DiffusionPipeline.from_pretrained(
    'mimbres/spansynth-edit',
    custom_pipeline='mimbres/spansynth-edit',
    trust_remote_code=True,
).to('cuda')
result = pipe(
    audio='../spansynth-inputs/early-slakh-track00006-original.mp3',
    midi='../spansynth-inputs/early-slakh-track00006-after.mid',
    num_inference_steps=16,
    guidance_scale=2.0,
)
sf.write('edited.wav', result.audios[0, 0], result.sample_rate, subtype='FLOAT')

Use .to('mps') for Apple Silicon or .to('cpu') for CPU. For FlowEdit, add method='flowedit' and source_midi pointing to the original score. The pipeline also accepts the crop, region, MIDI offset, and context options described below. This custom pipeline uses the installed spansynth package. pipe.save_pretrained(path) saves the pipeline and both components for later loading.

Try an example

Defaults: spansynth-edit · 16 Euler steps · CFG 2.0
Audio context enabled · Context MIDI omitted

Example files

From the repository directory, download the sample audio and its original and revised MIDI files (about 0.5 MB total):

mkdir -p ../spansynth-inputs
for name in early-slakh-track00006-original.mp3 \
            early-slakh-track00006-before.mid \
            early-slakh-track00006-after.mid; do
  curl --fail --location --output "../spansynth-inputs/$name" \
    "https://raw.githubusercontent.com/mimbres/spansynth-edit/main/demo/assets/$name"
done

The examples below regenerate 6.40–14.08 s within the first 20.48 s of the recording.

spansynth-edit

Provide the full revised MIDI, including notes that should remain unchanged within the selected region:

spansynth-edit edit \
  --audio ../spansynth-inputs/early-slakh-track00006-original.mp3 \
  --midi ../spansynth-inputs/early-slakh-track00006-after.mid \
  --cfg 2.0 \
  --output ../spansynth-results/slakh-edit

spansynth-edit + flowedit

Provide both the original and revised MIDI:

spansynth-edit edit --method flowedit \
  --audio ../spansynth-inputs/early-slakh-track00006-original.mp3 \
  --source-midi ../spansynth-inputs/early-slakh-track00006-before.mid \
  --midi ../spansynth-inputs/early-slakh-track00006-after.mid \
  --output ../spansynth-results/slakh-flowedit

Synthesis

Synthesise the original score in the selected region, using the surrounding recording as audio context:

spansynth-edit synthesize \
  --audio ../spansynth-inputs/early-slakh-track00006-original.mp3 \
  --midi ../spansynth-inputs/early-slakh-track00006-before.mid \
  --output ../spansynth-results/synthesis

Add --check-inputs to validate audio, MIDI, and timing without loading the model.

Options

Common options are listed below. Flags marked off are enabled by adding them to the command. For all options, run spansynth-edit edit --help or spansynth-edit synthesize --help.

Generation

Option Default Use
--method ordinary ordinary selects spansynth-edit; flowedit selects spansynth-edit + flowedit. Available with edit only.
--cfg 2.0 MIDI classifier-free guidance scale (0 or higher).
--steps 16 Number of Euler steps.
--context-midi off Use original MIDI outside the generated region. Requires --source-midi.
--drop-context-audio off Drop the audio-context condition. Audio outside the generated region is still preserved.

Inputs and timing

All times are in seconds.

Option Default Use
--audio required Source recording or audio context.
--midi required Target MIDI: the score to synthesise or the revised score.
--source-midi none Original MIDI, required for spansynth-edit + flowedit or --context-midi.
--crop-start 0.0 Crop start on the audio timeline.
--duration 20.48 Crop length, up to 20.48 seconds.
--edit-start, --edit-end 6.40, 14.08 Generated region relative to the crop.
--midi-offset 0.0 Offset added to target MIDI times to obtain audio times.
--source-midi-offset 0.0 Offset added to original MIDI times to obtain audio times.

Execution and output

Option Default Use
--output required Folder for generated audio and run settings.
--device auto Choose cpu, mps (Apple GPU), cuda, or cuda:N. auto tries CUDA, then MPS, then CPU.
--overwrite off Replace existing results in the output folder.

Timing and instruments

Provide aligned audio and MIDI, or use the offsets to align their timelines.

With --crop-start 30, a MIDI note at 36.4 s appears at 6.4 s in the crop. If MIDI time 0 corresponds to the start of that crop, use --midi-offset 30. Set --source-midi-offset independently for the original MIDI. Each run processes one crop, with edit boundaries rounded outward to 40 ms.

The instrument vocabulary lists supported MIDI programs and merged instrument groups. Programs in a group share one model category: for example, 0, 1, 3, 6, and 7 map to Acoustic Piano. Program numbers are zero-based, and MIDI channel 10 selects drums (internal program 128).

Results

File Contents
output.wav Full crop with the synthesised or edited region
generated.wav Generated region only
input.wav Source crop converted to 48 kHz mono
run.json Settings and timing

All audio outputs are 48 kHz mono. Outside the generated region, output.wav matches input.wav exactly. Use --overwrite to replace existing results.

License and credits

Code and model weights are released under Apache-2.0. See LICENSE and NOTICE for terms and credits, including YourMT3+ and HeartCodec. Demo recordings retain their original rights.

Training data retain their own licenses, including MAESTRO's CC BY-NC-SA 4.0 terms. The repository's Apache-2.0 license does not replace these dataset terms.

Citation

If you use SpanSynth-Edit in your research, please cite our paper:

@misc{chang2026spansynthedit,
  title={Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with {MIDI Span} conditioning},
  author={Sungkyun Chang and Keshav Bhandari and Simon Dixon and Emmanouil Benetos},
  year={2026},
  eprint={2609.25546},
  archivePrefix={arXiv},
  primaryClass={cs.SD},
  url={https://arxiv.org/abs/2609.25546}
}

Identity and Version

Repository
mimbres/spansynth-edit
Publisher
Sungkyun Chang
Task
Audio to audio
Modality
Audio
Library
diffusers
Parameters
481M parameters
Languages
Not stated by the source
Revision
5f0a7e3aca7267b3152ae9794a3f1bf01d0cc274
First published
2026-09-22
Last updated
2026-09-23

Files and Weights

16 files, 5.2 GB in total. The weights are 4 files totalling 5.2 GB in safetensors.

Weights4 files · 5.2 GB
Configuration8 files · 13.6 KB
Documentation3 files · 22.9 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
codec/diffusion_pytorch_model.safetensorsWeights672.4 MB d2a46ff20bfc
heartcodec-sq/scalar_model.safetensorsWeights672.4 MB 766b42871363
transformer/diffusion_pytorch_model.safetensorsWeights1.9 GB 5a6b4b420e60
v3-sq-v8-127750step/model.safetensorsWeights1.9 GB 5a6b4b420e60
codec/config.jsonConfiguration778 B —
codec/spansynth.diffusers_models.pyConfiguration55 B —
heartcodec-sq/scalar_model_config.jsonConfiguration592 B —
model_index.jsonConfiguration359 B —
pipeline.pyConfiguration10.3 KB —
transformer/config.jsonConfiguration641 B —
transformer/spansynth.diffusers_models.pyConfiguration65 B —
v3-sq-v8-127750step/config.jsonConfiguration838 B —
LICENSEDocumentation11.4 KB —
NOTICEDocumentation781 B —
README.mdDocumentation10.7 KB —
.gitattributesRepository1.5 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
5.2 GB
Download from Sungkyun Chang

Released by Sungkyun Chang through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2609.25546

Memory Requirements

PrecisionWeights in memory
As published5.2 GB
16-bit1.0 GB
8-bit0.5 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About spansynth-edit

How much GPU memory does spansynth-edit need?

About 1.2 GB at 16-bit and 0.3 GB at 4-bit: the weights (481M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run spansynth-edit on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use spansynth-edit commercially?

Yes. spansynth-edit is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Audio to audio

bigvgan_v2_22khz_80band_256x

NVIDIA

[[Paper]](https://arxiv.org/abs/2206.04658) - [[Code]](https://github.com/NVIDIA/BigVGAN) - [[Showcase]](https://bigvgan-demo.github.io/) - [[Project Page]](https://research.nvidia.com/labs/adlr/projects/bigvgan/) - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan-66959df3d97fd7d98d97dc9a) - [[Demo]](https://huggingface.co/spaces/nvidia/BigVGAN) - General refactor and code improvements for improved readability. - Fully fused CUDA kernel of anti-alised activation (upsampling + activation + downsampling) with inference speed benchmark. - We provide pretrained checkpoints of BigVGAN-v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling…

Open weights mit PyTorch

Model · Audio to audio

bigvgan_v2_44khz_128band_512x

NVIDIA

[[Paper]](https://arxiv.org/abs/2206.04658) - [[Code]](https://github.com/NVIDIA/BigVGAN) - [[Showcase]](https://bigvgan-demo.github.io/) - [[Project Page]](https://research.nvidia.com/labs/adlr/projects/bigvgan/) - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan-66959df3d97fd7d98d97dc9a) - [[Demo]](https://huggingface.co/spaces/nvidia/BigVGAN) - General refactor and code improvements for improved readability. - Fully fused CUDA kernel of anti-alised activation (upsampling + activation + downsampling) with inference speed benchmark. - We provide pretrained checkpoints of BigVGAN-v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling…

Open weights mit PyTorch

Descript Audio Codec running on-device on the LiteRT CompiledModel GPU (ML Drift). The convolutional encoder/decoder run on the GPU; the RVQ runs on CPU. 43:1 compression (1 s → 12×50 codes), RTF ≈ 0.82 (faster than real-time) on Pixel 8a. - dac16khzencoderfp16.tflite (43 MB) — audio[1,1,16000] → latent[1,1024,50], GPU. - dac16khzdeconlyzsfp16.tflite (105 MB) — latent[1,1024,50] → audio, GPU. - dacrvq.bin (1.2 MB) — RVQ weights (12 codebooks) for the CPU quantizer (float32 LE). encoder 367/367 + decoder 398/398 nodes on the LiteRT GPU delegate (LITERTCL, 1 partition, no CPU fallback); warm RTF ~0.82; reconstruction corr 1.0 vs PyTorch DAC. The decoder's ConvTranspose1d are rewritten to a…

Open weights mit litert

Model · Audio to audio

ArkEcho-RVC-M3-C-F

Ethernos

本仓库(Repository)所包含的所有人工神经网络(Artificial Neural Network)权重文件(.pth/.index)、训练日志及相关代码,均为计算声学(Computational Acoustics)与深度学习(Deep Learning)领域的技术研究实验产物。 这些文件本质上是高维张量(High-dimensional Tensors)的数值序列,通过随机梯度下降(SGD)与反向传播算法(Backpropagation)对公开可获取的音频数据进行统计建模(Statistical Modeling)得到。其技术形态与数字图像处理中的卷积核(Convolution Kernels)、自然语言处理中的词向量(Word Embeddings)并无本质差异。 本仓库不构成对任何第三方知识产权的故意侵犯,所有代码遵循 MIT License 开源协议,模型权重文件仅作为技术实现的副产品(By-products)存在。 本技术实验所使用的训练数据集(Training Dataset)包含以下角色的语音样本: - Mon3tr(モンスター)、Mon3tr 中文语音 等 - 声源版权归属:上海鹰角网络科技有限公司(Hypergryph Network Technology Co., Ltd.)及其关联公司 - 原始作品:《明日方舟》(Arknights) 上述角色的声音版权、肖像权、姓名权及相关知识产权均完全归属于鹰角网络及其合法授权方。本实验仅基于已公开发布的游戏内语音资源进行技术层面的信号处理(Signal Processing)与特征提取(Feature…

Open weights cc-by-nc-4.0

Model · Audio to audio

MOSS-v2-12to32-Enhancer-60M

LAION eV

A small causal Qwen3 model that predicts 32 MOSS v2 audio-codebook indices from the 12 indices of the current frame and up to ten previously generated HQ frames. It is trained from scratch on German, English, Spanish and French speech. This is an experimental enhancement model. Validation cross entropy measures token prediction under teacher forcing. It does not establish an improvement in perceived audio quality. Use the generated-history listening evaluation to judge speech content, speaker identity, artifacts and long-utterance stability. Root inference weights: step 64,699, selected by held-out extra-20-codebook cross entropy 5.818123. - Step 12,940: 100 originals / 200 input cases…

Open weights cc-by-nc-4.0 pytorch