Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event Parser
Yung-Hsuan Lai, Yen-Chun Chen, Frank Wang
摘要
Audio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its modality-aligned setting, i.e., the audio and visual modality are both assumed to signal the prediction target. With the Look, Listen, and Parse dataset (LLP), we investigate the under-explored unaligned setting, where the goal is to recognize audio and visual events in a video with only weak labels observed. Such weak video-level labels only tell what events happen without knowing the modality they are perceived (audio, visual, or both). To enhance learning in this challenging setting, we incorporate large-scale contrastively pre-trained models as the modality teachers. A simple, effective, and generic method, termed Visual-Audio Label Elaboration (VALOR), is innovated to harvest modality labels for the training events. Empirical studies show that the harvested labels significantly improve an attentional baseline by 8.0 in average F-score (Type@AV). Surprisingly, we found that modality-independent teachers outperform their modality-fused counterparts since they are noise-proof from the other potentially unaligned modality. Moreover, our best model achieves the new state-of-the-art on all metrics of LLP by a substantial margin (+5.4 F-score for Type@AV). VALOR is further generalized to Audio-Visual Event Localization and achieves the new state-of-the-art as well. Code is available at: https://github.com/Franklin905/VALOR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video ParsingPengcheng Zhao, Jinxing Zhou, Yang Zhao, Dan Guo 等AAAI 2025 · 被引用 19 次
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationJinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao 等AAAI 2026 · 被引用 4 次
- Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity CollaborationZiheng Zhou, Jinxing Zhou, Wei Qian, Shengeng Tang 等AAAI 2025 · 被引用 4 次
- PreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingXiao Yu, Yan Fang, Yao Zhao, Yunchao WeiNeurIPS 2025 · 被引用 4 次
- MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video ParsingLangyu Wang, Bingke Zhu, Yingying Chen, Yiyuan Zhang 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 被引用 687 次
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 被引用 395 次
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
相关 Paper
- Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video ParsingYu Wu, Yi YangCVPR 2021
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao 等ACM MM 2024 · 被引用 8 次
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er 等AAAI 2023 · 被引用 13 次
- Weakly-Supervised Audio-Visual Video Parsing with Prototype-Based Pseudo-LabelingKranthi Kumar Rachavarapu, Kalyan Ramakrishnan, A. N. RajagopalanCVPR 2024 · 被引用 3 次
- Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action LocalizationJun-Tae Lee, Mihir Jain, Hyoungwoo Park, Sungrack YunICLR 2021 · 被引用 73 次
