AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
Xinyue Guo, Xiaoran Yang, Lipan Zhang, Jianxuan Yang, Zhao Wang, Jian Luan
摘要
Sound effect editing—modifying audio by adding, removing, or replacing elements—remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and suboptimal audio quality. To address this, we propose AV-Edit, a generative sound effect editing framework that enables fine-grained editing of existing audio tracks in videos by jointly leveraging visual, audio, and text semantics. Specifically, the proposed method employs a specially designed contrastive audio-visual masking autoencoder (CAV-MAE-Edit) for multimodal pre-training, learning aligned cross-modal representations. These representations are then used to train an editorial Multimodal Diffusion Transformer (MM-DiT) capable of removing visually irrelevant sounds and generating missing audio elements consistent with video content through a correlation-based feature gating training strategy. Furthermore, we construct a dedicated video-based sound editing dataset as an evaluation benchmark. Experiments demonstrate that the proposed AV-Edit generates high-quality audio with precise modifications based on visual content, achieving state-of-the-art performance in the field of sound effect editing and exhibiting strong competitiveness in the domain of audio generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
- Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion ModelsSimian Luo, Chuanhao Yan, Chenxu Hu, Hang ZhaoNeurIPS 2023 · 被引用 192 次
相关 Paper
- AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video GenerationKai Wang, Shijian Deng, Jing Shi, Dimitrios Hatzinakos 等ACM MM 2025 · 被引用 2 次
- Animate and Sound an ImageXihua Wang, Ruihua Song, Chongxuan Li, Xin Cheng 等CVPR 2025
- Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and EditingZeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang 等SIGGRAPH 2026
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin 等ICCV 2025 · 被引用 2 次
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationJiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang 等NeurIPS 2025 · 被引用 6 次
