Mix and Localize: Localizing Sound Sources in Mixtures
Xixi Hu, Ziyang Chen, Andrew Owens
Abstract
We present a method for simultaneously localizing multiple sound sources within a visual scene. This task requires a model to both group a sound mixture into individual sources, and to associate them with a visual signal. Our method jointly solves both tasks at once, using a formulation inspired by the contrastive random walk of Jabri et al. We create a graph in which images and separated sounds correspond to nodes, and train a random walker to transition between nodes from different modalities with high return probability. The transition probabilities for this walk are determined by an audio-visual similarity metric that is learned by our model. We show through experiments with musical instruments and human speech that our model can successfully localize multiple sounds, outperforming other self-supervised methods. Project site: https://hxixixh.github.io/mix-and-localize .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbc631d3-c411-4e3c-bfe1-d68a71bc2543Cited by top-tier papers37
- Improving Audio-Visual Segmentation with Bidirectional GenerationDawei Hao, Yuxin Mao, Bowen He, Xiaodong Han et al.AAAI 2024 · 54 citations
- Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsFengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park et al.CVPR 2024 · 47 citations
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 44 citations
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 34 citations
- Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event ParserYung-Hsuan Lai, Yen-Chun Chen, Frank WangNeurIPS 2023 · 27 citations
Builds on14
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss et al.NeurIPS 2020 · 227 citations
- Labelling unlabelled videos from scratch with multi-modal self-supervisionYuki Markus Asano, Mandela Patrick, Christian Rupprecht, Andrea VedaldiNeurIPS 2020 · 169 citations
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- Move2Hear: Active Audio-Visual Source SeparationSagnik Majumder, Ziad Al-Halah, Kristen GraumanICCV 2021 · 48 citations
Related papers
- Audio-Visual Grouping Network for Sound Localization from MixturesShentong Mo, Yapeng TianCVPR 2023
- Improving Sound Source Localization with Joint Slot Attention on Image and AudioInho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim et al.CVPR 2025
- Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual ScenesHyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda SenocakCVPR 2025
- Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentChen Liu, Peike Li, Liying Yang, Dadong Wang et al.CVPR 2025
- Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source LocalizationTianyu Liu, Peng Zhang, Wei Huang, Yufei Zha et al.ACM MM 2023 · 4 citations
