Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization
Yuxin Guo, Shijie Ma, Hu Su, Zhiqing Wang, Yuhao Zhao, Wei Zou, Siyang Sun, Yun Zheng
Abstract
Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise localization, especially for small objects, and suffer from blurry boundaries and false positives. Moreover, the naive semi-supervised method is poor in fully leveraging the information of abundant unlabeled data. In this paper, we propose a novel semi-supervised learning framework for AVSL, namely Dual Mean-Teacher (DMT), comprising two teacher-student structures to circumvent the confirmation bias issue. Specifically, two teachers, pre-trained on limited labeled data, are employed to filter out noisy samples via the consensus between their predictions, and then generate high-quality pseudo-labels by intersecting their confidence maps. The sufficient utilization of both labeled and unlabeled data and the proposed unbiased framework enable DMT to outperform current state-of-the-art methods by a large margin, with CIoU of 90.4% and 48.8% on Flickr-SoundNet and VGG-Sound Source, obtaining 8.9%, 9.6% and 4.6%, 6.4% improvements over self-and semi-supervised methods respectively, given only < 3% positional-annotations. We also extend our framework to some existing AVSL methods and consistently boost their performance. Our code is available at https://github.com/gyx-gloria/DMT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0922f768-beb3-4536-a48b-b5655ee6f316Cited by top-tier papers5
- Happy: A Debiased Learning Framework for Continual Generalized Category DiscoveryShijie Ma, Fei Zhu, Zhun Zhong, Wenzhuo Liu et al.NeurIPS 2024 · 28 citations
- Active Generalized Category DiscoveryShijie Ma, Fei Zhu, Zhun Zhong, Xu-Yao Zhang et al.CVPR 2024 · 13 citations
- What's Making That Sound Right Now? Video-Centric Audio-Visual LocalizationHahyeon Choi, Junhoo Lee, Nojun KwakICCV 2025 · 1 citation
- CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-TrainingYuxin Guo, Siyang Sun, Shuailei Ma, Kecheng Zheng et al.CVPR 2024
- Aligned Better, Listen Better for Audio-Visual Large Language ModelsYuxin Guo, Shuailei Ma, Shijie Ma, Xiaoyi Bao et al.ICLR 2025
Builds on23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
Related papers
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 37 citations
- Learning Audio-Visual Source Localization via False Negative Aware Contrastive LearningWeixuan Sun, Jiayi Zhang, Jianyuan Wang, Zheyuan Liu et al.CVPR 2023
- A Closer Look at Weakly-Supervised Audio-Visual Source LocalizationShentong Mo, Pedro MorgadoNeurIPS 2022 · 92 citations
- Weakly-Supervised Audio-Visual SegmentationShentong Mo, Bhiksha RajNeurIPS 2023 · 26 citations
- Unsupervised Sounding Pixel LearningYining Zhang, Yanli Ji, Yang YangEMNLP 2023 · 2 citations
