Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source Localization
Sung Jin Um, Dongjin Kim, Jung Uk Kim
摘要
The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches only use audio as an auxiliary role to compare spatial regions of the visual modality. Humans, on the other hand, utilize both audio and visual modalities as spatial cues to locate sound sources. In this paper, we propose an audio-visual spatial integration network that integrates spatial cues from both modalities to mimic human behavior when detecting sound-making objects. Additionally, we introduce a recursive attention network to mimic human behavior of iterative focusing on objects, resulting in more accurate attention regions. To effectively encode spatial information from both modalities, we propose audio-visual pair matching loss and spatial region alignment loss. By utilizing the spatial cues of audio-visual modalities and recursively focusing objects, our method can perform more robust sound source localization. Comprehensive experimental results on the Flickr SoundNet and VGG-Sound Source datasets demonstrate the superiority of our proposed method over existing approaches. Our code is available at: https://github.com/VisualAIKHU/SIRA-SSL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Sonic4D: Spatial Audio Generation for Immersive 4D Scene ExplorationSiyi Xie, Hanxin Zhu, Xinyi Chen, Tianyu He 等AAAI 2026 · 被引用 4 次
- CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional VideoZhaolin Wan, Han Qin, Zhiyang Li, Xiaopeng Fan 等CVPR 2025
- Object-aware Sound Source Localization via Audio-Visual Scene UnderstandingSung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk KimCVPR 2025
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
- Learning to Visually Localize Sound Sources from Mixtures without Prior Source KnowledgeDongjin Kim, Sung Jin Um, Sangmin Lee, Jung Uk KimCVPR 2024
它引用的顶会 Paper5
- Robust Small-scale Pedestrian Detection with Cued Recall via Memory LearningJung Uk Kim, Sungjune Park, Yong Man RoICCV 2021 · 被引用 61 次
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng 等ACM MM 2022 · 被引用 34 次
- A Proposal-based Paradigm for Self-supervised Sound Source Localization in VideosHanyu Xuan, Zhiliang Wu, Jian Yang, Yan Yan 等CVPR 2022 · 被引用 27 次
- Localizing Visual Sounds the Hard WayHonglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani 等CVPR 2021
- Recursive Least-Squares Estimator-Aided Online Learning for Visual TrackingJin Gao, Weiming Hu, Yan LuCVPR 2020
相关 Paper
- Binaural Audio-Visual LocalizationXinyi Wu, Zhenyao Wu, Lili Ju, Song WangAAAI 2021 · 被引用 32 次
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 被引用 92 次
- Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source LocalizationTianyu Liu, Peng Zhang, Wei Huang, Yufei Zha 等ACM MM 2023 · 被引用 4 次
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao 等ACM MM 2024 · 被引用 8 次
- Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationHanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang 等AAAI 2020 · 被引用 110 次
