Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative Learning
Yuan Ji, Xu Jia, Huchuan Lu, Xiang Ruan
Abstract
Weakly supervised temporal action localization (WTAL) is a challenging task as only video-level category labels are available during training stage. Without precise temporal annotations, most approaches rely on complementary RGB and optical flow features to predict the start and end frame of each action category in a video. However, existing approaches simply resort to either concatenation or weighted sum to learn how to take advantages of these two modalities for accurate action localization, which ignore the substantial variance between such two modalities. In this paper, we present Cross-Stream Collaborative Learning (CSCL) to address these issues. The proposed CSCL introduce a cross-stream weighting module to identify which modality is more robust during training and take advantage of the robust modality to guide the weaker one. Furthermore, we suppress the snippets which has high action-ness scores in both modalities to further exploiting the complementary property between two modalities. In addition, we bring the concept of co-training for WTAL and take both modalities into account for pseudo label generation to help training a stronger model. Extensive experiments conducted on THUMOS14 and ActivityNet dataset demonstrate that CSCL achieves a favorable performance against state-of-the-arts methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9bbbf245-4b33-4c00-8e0f-5ab1f38c0fa4Cited by top-tier papers5
- Quality-Agnostic Deepfake Detection with Intra-model Collaborative LearningBinh Minh Le, Simon S. WooICCV 2023 · 50 citations
- Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object SegmentationXiaoqi Zhao, Youwei Pang, Jiaxing Yang, Lihe Zhang et al.ACM MM 2021 · 35 citations
- Temporal Sentiment Localization: Listen and Look in Untrimmed VideosZhicheng Zhang, Jufeng YangACM MM 2022 · 19 citations
- DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action LocalizationXiaojun Tang, Junsong Fan, Chuanchen Luo, Zhaoxiang Zhang et al.ICCV 2023 · 16 citations
- Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention ModelingGuiqin Wang, Peng Zhao, Cong Zhao, Shusen Yang et al.ICCV 2023 · 7 citations
Related papers
- Uncertainty Guided Collaborative Training for Weakly Supervised Temporal Action DetectionWenfei Yang, Tianzhu Zhang, Xiaoyuan Yu, Qi Tian et al.CVPR 2021
- Cross-modal Consensus Network for Weakly Supervised Temporal Action LocalizationFa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan et al.ACM MM 2021 · 104 citations
- Distilling Vision-Language Pre-Training to Collaborate with Weakly-Supervised Temporal Action LocalizationChen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao et al.CVPR 2023
- Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextZiyi Liu, Le Wang, Wei Tang, Junsong Yuan et al.AAAI 2021 · 28 citations
- Learning Temporal Co-Attention Models for Unsupervised Video Action LocalizationGuoqiang Gong, Xinghan Wang, Yadong Mu, Qi TianCVPR 2020
