Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
Haowen Li, Tianxiang Li, Yi Yang, Boyu Cao, Qi Liu
Abstract
The advancement of diffusion-based text-to-music generation has opened new avenues for zero-shot music editing. However, existing methods fail to achieve stem-specific timbre transfer, which requires altering specific stems while strictly preserving the background accompaniment. This limitation severely hinders practical application, since real-world production necessitates precise manipulation of components within dense mixtures. Our key finding is that, while vanilla cross-attention captures semantic features of stems, it lacks the spectral resolution to strictly localize targets in dense mixtures, leading to boundary leakage. To resolve this dilemma, we propose Polyphonia , a zero-shot editing framework with Acoustic-Informed Attention Calibration. Rather than relying solely on diffuse semantic attention, Polyphonia leverages a probabilistic acoustic prior to establish coarse boundaries, enabling non-target stems preserved precise semantic synthesis. For evaluation, we propose PolyEvalPrompts , a standardized prompt set with 1,170 timbre transfer tasks in polyphonic music. Specifically, Polyphonia achieves an increase of 15.5% in target alignment compared to baselines, while maintaining competitive music fidelity and non-target integrity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3cda3d2-077c-470b-96b1-c4867b502118Builds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan et al.ICCV 2023 · 770 citations
Related papers
- MelodyEdit: Zero-shot Music Editing with Disentangled Inversion ControlHuadai Liu, Jialei Wang, Xiangtai Li, Wen Wang et al.ACM MM 2025
- Prompt-guided Precise Audio Editing with Diffusion ModelsManjie Xu, Chenxing Li, Duzhen Zhang, Dan Su et al.ICML 2024 · 15 citations
- Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion ModelsYi Yang, Haowen Li, Tianxiang Li, Boyu Cao et al.AAAI 2026 · 1 citation
- SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music EditingXinlei Niu, Kin Wai Cheuk, Jing Zhang, Naoki Murata et al.AAAI 2026 · 5 citations
- Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM InversionHila Manor, Tomer MichaeliICML 2024 · 61 citations
