Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization
Tianyu Liu, Peng Zhang, Wei Huang, Yufei Zha, Tao You, Yanning Zhang
摘要
Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and sound sources in visual scenarios. Unfortunately, the insufficient attention to the heterogeneity influence in the different modality features still limits this scheme to be further improved, which also becomes the motivation of our work. In this study, an Induction Network is proposed to bridge the modality gap more effectively. By decoupling the gradients of visual and audio modalities, the discriminative visual representations of sound sources can be learned with the designed Induction Vector in a bootstrap manner, which also enables the audio modality to be aligned with the visual modality consistently. In addition to a visual weighted contrastive loss, an adaptive threshold selection strategy is introduced to enhance the robustness of the Induction Network. Substantial experiments conducted on SoundNet-Flickr and VGG-Sound Source datasets have demonstrated a superior performance compared to other state-of-the-art works in different challenging scenarios. The code is available at https://github.com/Tahy1/AVIN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Prompting Segmentation with Sound Is Generalizable Audio-Visual Source LocalizerYaoting Wang, Weisong Liu, Guangyao Li, Jian Ding 等AAAI 2024 · 被引用 42 次
- Mixtures of Experts for Audio-Visual LearningYing Cheng, Yang Li, Junjie He, Rui FengNeurIPS 2024 · 被引用 22 次
- Cyclic Learning for Binaural Audio Generation and LocalizationZhaojian Li, Bin Zhao, Yuan YuanCVPR 2024
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan 等NeurIPS 2020 · 被引用 156 次
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan 等CVPR 2022 · 被引用 143 次
相关 Paper
- Learning Audio-Visual Source Localization via False Negative Aware Contrastive LearningWeixuan Sun, Jiayi Zhang, Jianyuan Wang, Zheyuan Liu 等CVPR 2023
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 被引用 37 次
- SCLAV: Supervised Cross-modal Contrastive Learning for Audio-Visual CodingChao Sun, Min Chen, Jialiang Cheng, Han Liang 等ACM MM 2023 · 被引用 3 次
- Unsupervised Sounding Pixel LearningYining Zhang, Yanli Ji, Yang YangEMNLP 2023 · 被引用 2 次
- Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source LocalizationSung Jin Um, Dongjin Kim, Jung Uk KimACM MM 2023 · 被引用 4 次
