Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source Localization
Sung Jin Um, Dongjin Kim, Jung Uk Kim
Abstract
The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches only use audio as an auxiliary role to compare spatial regions of the visual modality. Humans, on the other hand, utilize both audio and visual modalities as spatial cues to locate sound sources. In this paper, we propose an audio-visual spatial integration network that integrates spatial cues from both modalities to mimic human behavior when detecting sound-making objects. Additionally, we introduce a recursive attention network to mimic human behavior of iterative focusing on objects, resulting in more accurate attention regions. To effectively encode spatial information from both modalities, we propose audio-visual pair matching loss and spatial region alignment loss. By utilizing the spatial cues of audio-visual modalities and recursively focusing objects, our method can perform more robust sound source localization. Comprehensive experimental results on the Flickr SoundNet and VGG-Sound Source datasets demonstrate the superiority of our proposed method over existing approaches. Our code is available at: https://github.com/VisualAIKHU/SIRA-SSL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9939cd9c-c0e9-4307-b044-e32533b43ebfCited by top-tier papers5
- Sonic4D: Spatial Audio Generation for Immersive 4D Scene ExplorationSiyi Xie, Hanxin Zhu, Xinyi Chen, Tianyu He et al.AAAI 2026 · 4 citations
- CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional VideoZhaolin Wan, Han Qin, Zhiyang Li, Xiaopeng Fan et al.CVPR 2025
- Object-aware Sound Source Localization via Audio-Visual Scene UnderstandingSung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk KimCVPR 2025
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
- Learning to Visually Localize Sound Sources from Mixtures without Prior Source KnowledgeDongjin Kim, Sung Jin Um, Sangmin Lee, Jung Uk KimCVPR 2024
Builds on5
- Robust Small-scale Pedestrian Detection with Cued Recall via Memory LearningJung Uk Kim, Sungjune Park, Yong Man RoICCV 2021 · 61 citations
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng et al.ACM MM 2022 · 34 citations
- A Proposal-based Paradigm for Self-supervised Sound Source Localization in VideosHanyu Xuan, Zhiliang Wu, Jian Yang, Yan Yan et al.CVPR 2022 · 27 citations
- Localizing Visual Sounds the Hard WayHonglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani et al.CVPR 2021
- Recursive Least-Squares Estimator-Aided Online Learning for Visual TrackingJin Gao, Weiming Hu, Yan LuCVPR 2020
Related papers
- Binaural Audio-Visual LocalizationXinyi Wu, Zhenyao Wu, Lili Ju, Song WangAAAI 2021 · 32 citations
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 92 citations
- Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source LocalizationTianyu Liu, Peng Zhang, Wei Huang, Yufei Zha et al.ACM MM 2023 · 4 citations
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
- Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationHanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang et al.AAAI 2020 · 110 citations
