SAO-Instruct: Free-form Audio Editing using Natural Language Instructions
Michael Ungersböck, Florian Grötschla, Luca A. Lanzendörfer, June Young Yi, Changho Choi, Roger Wattenhofer
摘要
Generative models have made significant progress in synthesizing high-fidelity audio from short textual descriptions. However, editing existing audio using natural language has remained largely underexplored. Current approaches either require the complete description of the edited audio or are constrained to predefined edit instructions that lack flexibility. In this work, we introduce SAO-Instruct, a model based on Stable Audio Open capable of editing audio clips using any free-form natural language instruction. To train our model, we create a dataset of audio editing triplets (input audio, edit instruction, output audio) using Prompt-to-Prompt, DDPM inversion, and a manual editing pipeline. Although partially trained on synthetic data, our model generalizes well to real in-the-wild audio clips and unseen edit instructions. We demonstrate that SAO-Instruct achieves competitive performance on objective metrics and outperforms other audio editing approaches in a subjective listening study. To encourage future research, we release our code and model weights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
相关 Paper
- AUDIT: Audio Editing by Following Instructions with Latent Diffusion ModelsYuancheng Wang, Zeqian Ju, Xu Tan, Lei He 等NeurIPS 2023 · 被引用 120 次
- Prompt-guided Precise Audio Editing with Diffusion ModelsManjie Xu, Chenxing Li, Duzhen Zhang, Dan Su 等ICML 2024 · 被引用 15 次
- AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaQifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan 等CVPR 2025
- SmartDJ: Declarative Audio Editing with Audio Language ModelZitong Lan, Yiduo Hao, Mingmin ZhaoICLR 2026 · 被引用 11 次
- InstructSpeech: Following Speech Editing Instructions via Large Language ModelsRongjie Huang, Ruofan Hu, Yongqi Wang, Zehan Wang 等ICML 2024 · 被引用 10 次
