Mix and Localize: Localizing Sound Sources in Mixtures
Xixi Hu, Ziyang Chen, Andrew Owens
摘要
We present a method for simultaneously localizing multiple sound sources within a visual scene. This task requires a model to both group a sound mixture into individual sources, and to associate them with a visual signal. Our method jointly solves both tasks at once, using a formulation inspired by the contrastive random walk of Jabri et al. We create a graph in which images and separated sounds correspond to nodes, and train a random walker to transition between nodes from different modalities with high return probability. The transition probabilities for this walk are determined by an audio-visual similarity metric that is learned by our model. We show through experiments with musical instruments and human speech that our model can successfully localize multiple sounds, outperforming other self-supervised methods. Project site: https://hxixixh.github.io/mix-and-localize .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Improving Audio-Visual Segmentation with Bidirectional GenerationDawei Hao, Yuxin Mao, Bowen He, Xiaodong Han 等AAAI 2024 · 被引用 54 次
- Binding Touch to Everything: Learning Unified Multimodal Tactile RepresentationsFengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park 等CVPR 2024 · 被引用 47 次
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 被引用 44 次
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 被引用 34 次
- Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event ParserYung-Hsuan Lai, Yen-Chun Chen, Frank WangNeurIPS 2023 · 被引用 27 次
它引用的顶会 Paper14
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss 等NeurIPS 2020 · 被引用 227 次
- Labelling unlabelled videos from scratch with multi-modal self-supervisionYuki Markus Asano, Mandela Patrick, Christian Rupprecht, Andrea VedaldiNeurIPS 2020 · 被引用 169 次
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan 等NeurIPS 2020 · 被引用 156 次
- Move2Hear: Active Audio-Visual Source SeparationSagnik Majumder, Ziad Al-Halah, Kristen GraumanICCV 2021 · 被引用 48 次
相关 Paper
- Audio-Visual Grouping Network for Sound Localization from MixturesShentong Mo, Yapeng TianCVPR 2023
- Improving Sound Source Localization with Joint Slot Attention on Image and AudioInho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim 等CVPR 2025
- Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual ScenesHyeonggon Ryu, Seongyu Kim, Joon Son Chung, Arda SenocakCVPR 2025
- Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentChen Liu, Peike Li, Liying Yang, Dadong Wang 等CVPR 2025
- Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source LocalizationTianyu Liu, Peng Zhang, Wei Huang, Yufei Zha 等ACM MM 2023 · 被引用 4 次
