Papers
arxiv:2609.25546

Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

Published on Sep 22
Authors:
,
,

Abstract

Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.25546
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.25546 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 1