Object-aware Sound Source Localization via Audio-Visual Scene Understanding
Sung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk Kim
Abstract
Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objects coexist. This limitation arises primarily from their reliance on simple audio-visual correspondence, which does not capture fine-grained semantic differences between sound-making and silent objects. To address these challenges, we propose a novel sound source localization framework leveraging Multimodal Large Language Models (MLLMs) to generate detailed contextual information that explicitly distinguishes between sound-making foreground objects and silent background objects. To effectively integrate this detailed information, we introduce two novel loss functions: Object-aware Contrastive Alignment (OCA) loss and Object Region Isolation (ORI) loss. Extensive experimental results on MUSIC and VGGSound datasets demonstrate the effectiveness of our approach, significantly outperforming existing methods in both singlesource and multi-source localization scenarios. Code and generated detailed contextual information are available at: https://github.com/VisualAIKHU/OA-SSL . * Equal contribution † Corresponding author Visual Modality Audio Modality Visual Modality Background Foreground Foreground There is a playing guitar. Background There are guitars hanging on the wall, drum set, and •••.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4364fa98-3006-41c0-841e-09ba60e77f65Cited by top-tier papers3
- Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationJinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang et al.AAAI 2026 · 5 citations
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
- How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?Arda Senocak, Sooyoung Park, Tae-Hyun Oh, Joon Son ChungCVPR 2026
Builds on19
- With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual RepresentationsDebidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet et al.ICCV 2021 · 542 citations
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- Weakly Supervised Contrastive LearningMingkai Zheng, Fei Wang, Shan You, Chen Qian et al.ICCV 2021 · 153 citations
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 92 citations
- V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation ModelsHeng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright et al.AAAI 2024 · 84 citations
Related papers
- Learning to Visually Localize Sound Sources from Mixtures without Prior Source KnowledgeDongjin Kim, Sung Jin Um, Sangmin Lee, Jung Uk KimCVPR 2024
- Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentChen Liu, Peike Li, Liying Yang, Dadong Wang et al.CVPR 2025
- SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt ContextsKhanh Binh Nguyen, Chae Jung ParkCVPR 2026
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh et al.ICCV 2023 · 39 citations
- Gotta Hear Them All: Towards Sound Source Aware Audio GenerationWei Guo, Heng Wang, Jianbo Ma, Weidong CaiAAAI 2026 · 2 citations
