SCLAV: Supervised Cross-modal Contrastive Learning for Audio-Visual Coding
Chao Sun, Min Chen, Jialiang Cheng, Han Liang, Chuanbo Zhu, Jincai Chen
摘要
Audio and vision are important senses for high-level cognition, and their special strong correlation makes audio-visual coding a crucial factor in many multimodal tasks. However, there are two challenges in audio-visual coding. First, the heterogeneity of multimodal data often leads to misalignment of cross-modal features under the same sample, which reduces their representation quality. Second, most self-supervised learning frameworks are constructed based on instance semantics, and the generated pseudo labels introduce additional classification noise. To address these challenges, we propose a Supervised Cross-modal Contrastive Learning Framework for Audio-Visual Coding (SCLAV). Our framework includes an audio-visual coding network composed of an inter-modal attention interaction module and an intra-modal self-integration module, which leverage multimodal complementary and hidden information for better representation. Additionally, we introduce a supervised cross-modal contrastive loss to minimize the distance between audio and vision features of the same instance, and use weak labels of multimodal data to eliminate the feature-oriented classification noise. Extensive experiments on the AVE and XD-Violence datasets demonstrate that SCLAV outperforms the state-of-the-art results, even with limited computational resources.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Robust Audio-Visual Instance DiscriminationPedro Morgado, Ishan Misra, Nuno VasconcelosCVPR 2021
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath 等ICLR 2023 · 被引用 17 次
- Enhancing Audio-Visual Association with Self-Supervised Curriculum LearningJingran Zhang, Xing Xu, Fumin Shen, Huimin Lu 等AAAI 2021 · 被引用 22 次
- Audio-Visual Instance Discrimination with Cross-Modal AgreementPedro Morgado, Nuno Vasconcelos, Ishan MisraCVPR 2021
- Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source LocalizationTianyu Liu, Peng Zhang, Wei Huang, Yufei Zha 等ACM MM 2023 · 被引用 4 次
