Prompt-guided Precise Audio Editing with Diffusion Models
Manjie Xu, Chenxing Li, Duzhen Zhang, Dan Su, Wei Liang, Dong Yu
Abstract
Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as Prompt-guided Precise Audio Editing (PPAE), which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely trainingfree. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b1397a5-02af-4c5f-b3ad-80e641e86b4eCited by top-tier papers3
- SmartDJ: Declarative Audio Editing with Audio Language ModelZitong Lan, Yiduo Hao, Mingmin ZhaoICLR 2026 · 11 citations
- Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion ModelShuai Yu, Xiaoliang He, Kangjie Dong, Yi YuAAAI 2026
- Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention CalibrationHaowen Li, Tianxiang Li, Yi Yang, Boyu Cao et al.ICML 2026
Builds on15
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
Related papers
- AUDIT: Audio Editing by Following Instructions with Latent Diffusion ModelsYuancheng Wang, Zeqian Ju, Xu Tan, Lei He et al.NeurIPS 2023 · 120 citations
- MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted GuidanceQi Mao, Lan Chen, Yuchao Gu, Zhen Fang et al.ACM MM 2024 · 7 citations
- SAO-Instruct: Free-form Audio Editing using Natural Language InstructionsMichael Ungersböck, Florian Grötschla, Luca A. Lanzendörfer, June Young Yi et al.NeurIPS 2025 · 9 citations
- An Item Is Worth a Prompt: Versatile Image Editing with Disentangled ControlAosong Feng, Weikang Qiu, Jinbin Bai, Zhen Dong et al.AAAI 2025 · 9 citations
- Training-Free Text-Guided Image Editing with Visual Autoregressive ModelYufei Wang, Lanqing Guo, Zhihao Li, Jiaxing Huang et al.ICCV 2025
