Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
Hila Manor, Tomer Michaeli
Abstract
Editing signals using large pre-trained models, in a zero-shot manner, has recently seen rapid advancements in the image domain. However, this wave has yet to reach the audio domain. In this paper, we explore two zero-shot editing techniques for audio signals, which use DDPM inversion with pre-trained diffusion models. The first, which we coin ZEro-shot Text-based Audio (ZETA) editing, is adopted from the image domain. The second, named ZEro-shot UnSupervized (ZEUS) editing, is a novel approach for discovering semantically meaningful editing directions without supervision. When applied to music signals, this method exposes a range of musically interesting modifications, from controlling the participation of specific instruments to improvisations on the melody. Samples and code can be found in https://hilamanor.github.io/AudioEditing/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b9a0fd2-6094-4e98-a56e-1489ff7b38b0Cited by top-tier papers20
- Flowedit: Inversion-Free Text-Based Editing Using Pre-Trained Flow ModelsVladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas, Tomer MichaeliICCV 2025 · 30 citations
- Exploring Low-Dimensional Subspace in Diffusion Models for Controllable Image EditingSiyi Chen, Huijie Zhang, Minzhe Guo, Yifu Lu et al.NeurIPS 2024 · 29 citations
- SonicMaster: Towards Controllable All-in-One Music Restoration and MasteringJan Melechovsky, Ambuj Mehrish, Abhinaba Roy, Dorien HerremansICML 2026 · 11 citations
- SmartDJ: Declarative Audio Editing with Audio Language ModelZitong Lan, Yiduo Hao, Mingmin ZhaoICLR 2026 · 11 citations
- SAO-Instruct: Free-form Audio Editing using Natural Language InstructionsMichael Ungersböck, Florian Grötschla, Luca A. Lanzendörfer, June Young Yi et al.NeurIPS 2025 · 9 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- AUDIT: Audio Editing by Following Instructions with Latent Diffusion ModelsYuancheng Wang, Zeqian Ju, Xu Tan, Lei He et al.NeurIPS 2023 · 120 citations
- MelodyEdit: Zero-shot Music Editing with Disentangled Inversion ControlHuadai Liu, Jialei Wang, Xiangtai Li, Wen Wang et al.ACM MM 2025
- SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music EditingXinlei Niu, Kin Wai Cheuk, Jing Zhang, Naoki Murata et al.AAAI 2026 · 5 citations
- ZeroSep: Separate Anything in Audio with Zero TrainingChao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang et al.NeurIPS 2025 · 8 citations
- Prompt-guided Precise Audio Editing with Diffusion ModelsManjie Xu, Chenxing Li, Duzhen Zhang, Dan Su et al.ICML 2024 · 15 citations
