Binaural Audio-Visual Localization
Xinyi Wu, Zhenyao Wu, Lili Ju, Song Wang
摘要
Localizing sound sources in a visual scene has many important applications and quite a few traditional or learning-based methods have been proposed for this task. Humans have the ability to roughly localize sound sources within or beyond the range of the vision using their binaural system. However most existing methods use monaural audio, instead of binaural audio, as a modality to help the localization. In addition, prior works usually localize sound sources in the form of object-level bounding boxes in images or videos and evaluate the localization accuracy by examining the overlap between the ground-truth and predicted bounding boxes. This is too rough since a real sound source is often only a part of an object. In this paper, we propose a deep learning method for pixel-level sound source localization by leveraging both binaural recordings and the corresponding videos. Specifically, we design a novel Binaural Audio-Visual Network (BAVNet), which concurrently extracts and integrates features from binaural recordings and videos. We also propose a point-annotation strategy to construct pixel-level ground truth for network training and performance evaluation. Experimental results on Fair-Play and YT-Music datasets demonstrate the effectiveness of the proposed method and show that binaural audio can greatly improve the performance of localizing the sound sources, especially when the quality of the visual information is limited.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Hyperbolic Audio-visual Zero-shot LearningJie Hong, Zeeshan Hayder, Junlin Han, Pengfei Fang 等ICCV 2023 · 被引用 27 次
- How Would it Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor ScenesMahnoor Fatima Saad, Ziad Al-HalahICCV 2025 · 被引用 4 次
- Cyclic Learning for Binaural Audio Generation and LocalizationZhaojian Li, Bin Zhao, Yuan YuanCVPR 2024
- Unraveling Instance Associations: A Closer Look for Audio-Visual SegmentationYuanhong Chen, Yuyuan Liu, Hu Wang, Fengbei Liu 等CVPR 2024
- Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality PerspectivesZeliang Zhang, Susan Liang, Daiki Shimada, Chenliang XuICLR 2025
它引用的顶会 Paper5
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 被引用 233 次
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox 等ICCV 2019 · 被引用 157 次
- Recursive Visual Sound Separation Using Minus-Plus NetXudong Xu, Bo Dai, Dahua LinICCV 2019 · 被引用 95 次
- Vision-Infused Deep Audio InpaintingHang Zhou, Ziwei Liu, Xudong Xu, Ping Luo 等ICCV 2019 · 被引用 92 次
相关 Paper
- Localize to Binauralize: Audio Spatialization from Visual Sound Source LocalizationKranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. RajagopalanICCV 2021 · 被引用 27 次
- Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source LocalizationSung Jin Um, Dongjin Kim, Jung Uk KimACM MM 2023 · 被引用 4 次
- Audio-Visual Localization by Synthetic Acoustic Image GenerationValentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio MurinoAAAI 2021 · 被引用 9 次
- Supervising Sound Localization by In-the-wild EgomotionAnna Min, Ziyang Chen, Hang Zhao, Andrew OwensCVPR 2025
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 被引用 92 次
