A Closer Look at Weakly-Supervised Audio-Visual Source Localization
Shentong Mo, Pedro Morgado
Abstract
Audio-visual source localization is a challenging task that aims to predict the location of visual sound sources in a video. Since collecting ground-truth annotations of sounding objects can be costly, a plethora of weakly-supervised localization methods that can learn from datasets with no bounding-box annotations have been proposed in recent years, by leveraging the natural co-occurrence of audio and visual signals. Despite significant interest, popular evaluation protocols have two major flaws. First, they allow for the use of a fully annotated dataset to perform early stopping, thus significantly increasing the annotation effort required for training. Second, current evaluation metrics assume the presence of sound sources at all times. This is of course an unrealistic assumption, and thus better metrics are necessary to capture the model's performance on (negative) samples with no visible sound sources. To accomplish this, we extend the test set of popular benchmarks, Flickr SoundNet and VGG-Sound Sources, in order to include negative samples, and measure performance using metrics that balance localization accuracy and recall. Using the new protocol, we conducted an extensive evaluation of prior methods, and found that most prior works are not capable of identifying negatives and suffer from significant overfitting problems (rely heavily on early stopping for best results). We also propose a new approach for visual sound source localization that addresses both these problems. In particular, we found that, through extreme visual dropout and the use of momentum encoders, the proposed approach combats overfitting effectively, and establishes a new state-of-the-art performance on both Flickr SoundNet and VGG-Sound Source. Code and pre-trained models are available at https://github.com/stoneMo/SLAVC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63dbdb47-ddf1-48b2-9270-c3a3f9dddaf1Cited by top-tier papers38
- Connecting Multi-modal Contrastive RepresentationsZehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang et al.NeurIPS 2023 · 60 citations
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 44 citations
- Sound Source Localization is All about Cross-Modal AlignmentArda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh et al.ICCV 2023 · 39 citations
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 34 citations
- Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsChen Liu, Peike Patrick Li, Xingqun Qi, Hu Zhang et al.ACM MM 2023 · 33 citations
Builds on14
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 149 citations
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen SoundsEfthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey et al.ICLR 2021 · 83 citations
Related papers
- Learning Audio-Visual Source Localization via False Negative Aware Contrastive LearningWeixuan Sun, Jiayi Zhang, Jianyuan Wang, Zheyuan Liu et al.CVPR 2023
- Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source LocalizationYuxin Guo, Shijie Ma, Hu Su, Zhiqing Wang et al.NeurIPS 2023 · 19 citations
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 37 citations
- Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source LocalizationSung Jin Um, Dongjin Kim, Jung Uk KimACM MM 2023 · 4 citations
- Binaural Audio-Visual LocalizationXinyi Wu, Zhenyao Wu, Lili Ju, Song WangAAAI 2021 · 32 citations
