Boosting Positive Segments for Weakly-Supervised Audio-Visual Video Parsing
Kranthi Kumar Rachavarapu, A. N. Rajagopalan
Abstract
In this paper, we address the problem of weakly supervised Audio-Visual Video Parsing (AVVP), where the goal is to temporally localize events that are audible or visible and simultaneously classify them into known event categories. This is a challenging task, as we only have access to the video-level event labels during training but need to predict event labels at the segment level during evaluation. Existing multiple-instance learning (MIL) based methods use a form of attentive pooling over segment-level predictions. These methods only optimize for a subset of most discriminative segments that satisfy the weak-supervision constraints, which miss identifying positive segments. To address this, we focus on improving the proportion of positive segments detected in a video. To this end, we model the number of positive segments in a video as a latent variable and show that it can be modeled as Poisson binomial distribution over segment-level predictions, which can be computed exactly. Given the absence of fine-grained supervision, we propose an Expectation-Maximization approach to learn the model parameters by maximizing the evidence lower bound (ELBO). We iteratively estimate the minimum positive segments in a video and refine them to capture more positive segments. We conducted extensive experiments on AVVP tasks to evaluate the effectiveness of our proposed approach, and the results clearly demonstrate that it increases the number of positive segments captured compared to existing methods. Additionally, our experiments on Temporal Action Localization (TAL) demonstrate the potential of our method for generalization to similar MIL tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b6e01ed-3a3b-4680-aedf-f0c3840573bcCited by top-tier papers3
- Weakly-Supervised Audio-Visual Video Parsing with Prototype-Based Pseudo-LabelingKranthi Kumar Rachavarapu, Kalyan Ramakrishnan, A. N. RajagopalanCVPR 2024 · 3 citations
- UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video ParsingYung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang, François G. Germain et al.CVPR 2025
- Adaptive Momentum and EMA-weighted Modeling for Imbalanced Label Distribution LearningYongbiao Gao, Xiangcheng Sun, Chao Tan, Chunyu Hu et al.AAAI 2026
Builds on16
- Weakly-Supervised Action Localization With Background ModelingPhuc Xuan Nguyen, Deva Ramanan, Charless C. FowlkesICCV 2019 · 176 citations
- A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action LocalizationAshraful Islam, Chengjiang Long, Richard J. RadkeAAAI 2021 · 145 citations
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng et al.CVPR 2022 · 104 citations
- Cross-Modal Relation-Aware Networks for Audio-Visual Event LocalizationHaoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan et al.ACM MM 2020 · 97 citations
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin et al.NeurIPS 2021 · 94 citations
Related papers
- Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action LocalizationJun-Tae Lee, Mihir Jain, Hyoungwoo Park, Sungrack YunICLR 2021 · 73 citations
- Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language PerspectiveYingying Fan, Yu Wu, Bo Du, Yutian LinNeurIPS 2023 · 20 citations
- DHHN: Dual Hierarchical Hybrid Network for Weakly-Supervised Audio-Visual Video ParsingXun Jiang, Xing Xu, Zhiguo Chen, Jingran Zhang et al.ACM MM 2022 · 35 citations
- Multi-modal Grouping Network for Weakly-Supervised Audio-Visual Video ParsingShentong Mo, Yapeng TianNeurIPS 2022 · 73 citations
- Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention ModelingGuiqin Wang, Peng Zhao, Cong Zhao, Shusen Yang et al.ICCV 2023 · 7 citations
