Improving Sound Source Localization with Joint Slot Attention on Image and Audio
Inho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim, Suha Kwak
摘要
Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embedding vector each, and use them to learn SSL via contrastive learning. To this end, previous work samples one of local image features as the image embedding and aggregates all local audio features to obtain the audio embedding, which is far from optimal due to the presence of noise and background irrelevant to the actual target in the input. We present a novel SSL method that addresses this chronic issue by joint slot attention on image and audio. To be specific, two slots competitively attend image and audio features to decompose them into target and off-target representations, and only target representations of image and audio are used for contrastive learning. Also, we introduce cross-modal attention matching to further align local features of image and audio. Our method achieved the best in almost all settings on three public benchmarks for SSL, and substantially outperformed all the prior work in cross-modal retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 被引用 92 次
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 等ICCV 2023 · 被引用 39 次
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 被引用 37 次
相关 Paper
- Learning Spatially-Aware Language and Audio EmbeddingsBhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso 等NeurIPS 2024 · 被引用 31 次
- Unsupervised Sounding Pixel LearningYining Zhang, Yanli Ji, Yang YangEMNLP 2023 · 被引用 2 次
- A Proposal-based Paradigm for Self-supervised Sound Source Localization in VideosHanyu Xuan, Zhiliang Wu, Jian Yang, Yan Yan 等CVPR 2022 · 被引用 27 次
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
- Object-aware Sound Source Localization via Audio-Visual Scene UnderstandingSung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk KimCVPR 2025
