Object-aware Sound Source Localization via Audio-Visual Scene Understanding
Sung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk Kim
摘要
Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objects coexist. This limitation arises primarily from their reliance on simple audio-visual correspondence, which does not capture fine-grained semantic differences between sound-making and silent objects. To address these challenges, we propose a novel sound source localization framework leveraging Multimodal Large Language Models (MLLMs) to generate detailed contextual information that explicitly distinguishes between sound-making foreground objects and silent background objects. To effectively integrate this detailed information, we introduce two novel loss functions: Object-aware Contrastive Alignment (OCA) loss and Object Region Isolation (ORI) loss. Extensive experimental results on MUSIC and VGGSound datasets demonstrate the effectiveness of our approach, significantly outperforming existing methods in both singlesource and multi-source localization scenarios. Code and generated detailed contextual information are available at: https://github.com/VisualAIKHU/OA-SSL . * Equal contribution † Corresponding author Visual Modality Audio Modality Visual Modality Background Foreground Foreground There is a playing guitar. Background There are guitars hanging on the wall, drum set, and •••.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationJinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang 等AAAI 2026 · 被引用 5 次
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
- How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?Arda Senocak, Sooyoung Park, Tae-Hyun Oh, Joon Son ChungCVPR 2026
它引用的顶会 Paper19
- With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual RepresentationsDebidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet 等ICCV 2021 · 被引用 542 次
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan 等NeurIPS 2020 · 被引用 156 次
- Weakly Supervised Contrastive LearningMingkai Zheng, Fei Wang, Shan You, Chen Qian 等ICCV 2021 · 被引用 153 次
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 被引用 92 次
- V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation ModelsHeng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright 等AAAI 2024 · 被引用 84 次
相关 Paper
- Learning to Visually Localize Sound Sources from Mixtures without Prior Source KnowledgeDongjin Kim, Sung Jin Um, Sangmin Lee, Jung Uk KimCVPR 2024
- Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentChen Liu, Peike Li, Liying Yang, Dadong Wang 等CVPR 2025
- SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt ContextsKhanh Binh Nguyen, Chae Jung ParkCVPR 2026
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 等ICCV 2023 · 被引用 39 次
- Gotta Hear Them All: Towards Sound Source Aware Audio GenerationWei Guo, Heng Wang, Jianbo Ma, Weidong CaiAAAI 2026 · 被引用 2 次
