Enhanced Audio Tagging via Multi- to Single-Modal Teacher-Student Mutual Learning
Yifang Yin, Harsh Shrivastava, Ying Zhang, Zhenguang Liu, Rajiv Ratn Shah, Roger Zimmermann
Abstract
Recognizing ongoing events based on acoustic clues has been a critical yet challenging problem that has attracted significant research attention in recent years. Joint audio-visual analysis can improve the event detection accuracy but may not always be feasible as under many circumstances only audio recordings are available in real-world scenarios. To solve the challenges, we present a novel visual-assisted teacher-student mutual learning framework for robust sound event detection from audio recordings. Our model adopts a multi-modal teacher network based on both acoustic and visual clues, and a single-modal student network based on acoustic clues only. Conventional teacher-student learning performs unsatisfactorily for knowledge transfer from a multi-modality network to a single-modality network. We thus present a mutual learning framework by introducing a single-modal transfer loss and a cross-modal transfer loss to collaboratively learn the audio-visual correlations between the two networks. Our proposed solution takes the advantages of joint audio-visual analysis in training while maximizing the feasibility of the model in use cases. Our extensive experiments on the DCASE17 and the DCASE18 sound event detection datasets show that our proposed method outperforms the state-of-the-art audio tagging approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- AnyFace: Free-style Text-to-Face Synthesis and ManipulationJianxin Sun, Qiyao Deng, Qi Li, Muyi Sun et al.CVPR 2022 · 46 citations
- DFIL: Deepfake Incremental Learning by Exploiting Domain-invariant Forgery CluesKun Pan, Yifang Yin, Yao Wei, Feng Lin et al.ACM MM 2023 · 35 citations
- CrossMatch: Source-Free Domain Adaptive Semantic Segmentation via Cross-Modal Consistency TrainingYifang Yin, Wenmiao Hu, Zhenguang Liu, Guanfeng Wang et al.ICCV 2023 · 21 citations
- Prototypical Cross-domain Knowledge Transfer for Cervical Dysplasia Visual InspectionYichen Zhang, Yifang Yin, Ying Zhang, Zhenguang Liu et al.ACM MM 2023 · 4 citations
- Video Infringement Detection via Feature Disentanglement and Mutual Information MaximizationZhenguang Liu, Xinyang Yu, Ruili Wang, Shuai Ye et al.ACM MM 2023 · 1 citation
Builds on4
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 233 citations
- Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationHanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang et al.AAAI 2020 · 110 citations
- What Makes Training Multi-Modal Classification Networks Hard?Weiyao Wang, Du Tran, Matt FeiszliCVPR 2020
- Listen to Look: Action Recognition by Previewing AudioRuohan Gao, Tae-Hyun Oh, Kristen Grauman, Lorenzo TorresaniCVPR 2020
Related papers
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
- Enhancing Audio-Visual Association with Self-Supervised Curriculum LearningJingran Zhang, Xing Xu, Fumin Shen, Huimin Lu et al.AAAI 2021 · 22 citations
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng et al.ACM MM 2020 · 93 citations
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath et al.ICLR 2023 · 17 citations
