Visual Sound Localization in the Wild by Cross-Modal Interference Erasing
Xian Liu, Rui Qian, Hang Zhou, Di Hu, Weiyao Lin, Ziwei Liu, Bolei Zhou, Xiaowei Zhou
摘要
The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios are usually contaminated by off-screen sound and background noise. They will interfere with the procedure of identifying desired sources and building visual-sound connections, making previous studies non-applicable. In this work, we propose the Interference Eraser (IEr) framework, which tackles the problem of audio-visual sound source localization in the wild. The key idea is to eliminate the interference by redefining and carving discriminative audio representations. Specifically, we observe that the previous practice of learning only a single audio representation is insufficient due to the additive nature of audio signals. We thus extend the audio representation with our Audio-Instance-Identifier module, which clearly distinguishes sounding instances when audio signals of different volumes are unevenly mixed. Then we erase the influence of the audible but off-screen sounds and the silent but visible objects by a Cross-modal Referrer module with cross-modality distillation. Quantitative and qualitative evaluations demonstrate that our proposed framework achieves superior results on sound localization tasks, especially under real-world scenarios. Code is available at https://github.com/ alvinliu0/Visual-Sound-Localization-in-the-Wild .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Learning Hierarchical Cross-Modal Association for Co-Speech Gesture GenerationXian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu 等CVPR 2022 · 被引用 118 次
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 等CVPR 2022 · 被引用 101 次
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 被引用 92 次
- Audio-Driven Co-Speech Gesture Video GenerationXian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 等NeurIPS 2022 · 被引用 77 次
- Connecting Multi-modal Contrastive RepresentationsZehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang 等NeurIPS 2023 · 被引用 60 次
它引用的顶会 Paper10
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani 等NeurIPS 2020 · 被引用 483 次
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan 等NeurIPS 2020 · 被引用 156 次
- Vision-Infused Deep Audio InpaintingHang Zhou, Ziwei Liu, Xudong Xu, Ping Luo 等ICCV 2019 · 被引用 92 次
相关 Paper
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 等ICCV 2023 · 被引用 39 次
- Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio SemanticsChen Liu, Liying Yang, Peike Li, Dadong Wang 等CVPR 2025
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang 等ICLR 2025
- Visual Scene Graphs for Audio Source SeparationMoitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop CherianICCV 2021 · 被引用 45 次
- AV-RISE: Hierarchical Cross-Modal Denoising for Learning Robust Audio-Visual Speech RepresentationZhishuo Zhao, Yi Lin, Dongyue Guo, Junyu FanACM MM 2025 · 被引用 1 次
