Weakly-Supervised Audio-Visual Video Parsing with Prototype-Based Pseudo-Labeling
Kranthi Kumar Rachavarapu, Kalyan Ramakrishnan, A. N. Rajagopalan
摘要
In this paper, we address the weakly-supervised Audio-Visual Video Parsing (AVVP) problem, which aims at labeling events in a video as audible, visible, or both, and temporally localizing and classifying them into known categories. This is challenging since we only have access to video-level (weak) event labels when training but need to predict event labels at the segment (frame) level at test time. Recent methods employ multiple-instance learning (MIL) techniques that tend to focus solely on the most discriminative segments, resulting in frequent misclassifications. Our idea is to first construct several "prototype" features for each event class by clustering key segments identified for the event in the training data. We then assign pseudo labels to all training segments based on their feature similarities with these prototypes and re-train the model under weak and strong supervision. We facilitate this by structuring the feature space with contrastive learning using pseudo labels. Experiments show that we outperform existing methods for weakly-supervised AVVP. We also show that learning with weak and iteratively re-estimated pseudo labels can be interpreted as an expectation-maximization (EM) algorithm, providing further insight for our training procedure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment LocalizationCailing Han, Zhangbin Li, Jinxing Zhou, Wei Qian 等CVPR 2026 · 被引用 1 次
- UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video ParsingYung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang, François G. Germain 等CVPR 2025
- Adaptive Momentum and EMA-weighted Modeling for Imbalanced Label Distribution LearningYongbiao Gao, Xiangcheng Sun, Chao Tan, Chunyu Hu 等AAAI 2026
- Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic ThresholdsEitan Shaar, Ariel Shaulov, Gal Chechik, Lior WolfCVPR 2025
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou 等ICCV 2019 · 被引用 1,404 次
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 被引用 956 次
- Attribute Prototype Network for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 等NeurIPS 2020 · 被引用 392 次
- Rethinking Semantic Segmentation: A Prototype ViewTianfei Zhou, Wenguan Wang, Ender Konukoglu, Luc Van GoolCVPR 2022 · 被引用 353 次
相关 Paper
- Boosting Positive Segments for Weakly-Supervised Audio-Visual Video ParsingKranthi Kumar Rachavarapu, A. N. RajagopalanICCV 2023 · 被引用 13 次
- Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video ParsingYu Wu, Yi YangCVPR 2021
- MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video ParsingLangyu Wang, Bingke Zhu, Yingying Chen, Yiyuan Zhang 等ICCV 2025 · 被引用 1 次
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er 等AAAI 2023 · 被引用 13 次
- Weakly-Supervised Audio-Visual SegmentationShentong Mo, Bhiksha RajNeurIPS 2023 · 被引用 26 次
