Improving Sound Source Localization with Joint Slot Attention on Image and Audio
Inho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim, Suha Kwak
Abstract
Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embedding vector each, and use them to learn SSL via contrastive learning. To this end, previous work samples one of local image features as the image embedding and aggregates all local audio features to obtain the audio embedding, which is far from optimal due to the presence of noise and background irrelevant to the actual target in the input. We present a novel SSL method that addresses this chronic issue by joint slot attention on image and audio. To be specific, two slots competitively attend image and audio features to decompose them into target and off-target representations, and only target representations of image and audio are used for contrastive learning. Also, we introduce cross-modal attention matching to further align local features of image and audio. Our method achieved the best in almost all settings on three public benchmarks for SSL, and substantially outperformed all the prior work in cross-modal retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7690d15e-1e2c-48a3-af33-d9631458ac7fCited by top-tier papers1
Ask how each one uses itBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 92 citations
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh et al.ICCV 2023 · 39 citations
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 37 citations
Related papers
- Learning Spatially-Aware Language and Audio EmbeddingsBhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso et al.NeurIPS 2024 · 31 citations
- Unsupervised Sounding Pixel LearningYining Zhang, Yanli Ji, Yang YangEMNLP 2023 · 2 citations
- A Proposal-based Paradigm for Self-supervised Sound Source Localization in VideosHanyu Xuan, Zhiliang Wu, Jian Yang, Yan Yan et al.CVPR 2022 · 27 citations
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
- Object-aware Sound Source Localization via Audio-Visual Scene UnderstandingSung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk KimCVPR 2025
