Cross-modal Consensus Network for Weakly Supervised Temporal Action Localization
Fa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan, Wei-Shi Zheng
Abstract
Weakly supervised temporal action localization (WS-TAL) is a challenging task that aims to localize action instances in the given video with video-level categorical supervision. Previous works use the appearance and motion features extracted from pre-trained feature encoder directly, e.g., feature concatenation or score-level fusion. In this work, we argue that the features extracted from the pre-trained extractors, e.g., I3D, which are trained for trimmed video action classification, but not specific for WS-TAL task, leading to inevitable redundancy and sub-optimization . Therefore, the feature re-calibration is needed for reducing the task-irrelevant information redundancy. Here, we propose a cross-modal consensus network (CO 2 -Net) to tackle this problem. In CO 2 -Net, we mainly introduce two identical proposed cross-modal consensus modules (CCM) that design a cross-modal attention mechanism to filter out the task-irrelevant information redundancy using the global information from the main modality and the cross-modal local information from the auxiliary modality. Moreover, we further explore inter-modality consistency, where we treat the attention weights derived from each CCM as the pseudo targets of the attention weights derived from another CCM to maintain the consistency between the predictions derived from two CCMs, forming a mutual learning manner. Finally, we conduct extensive experiments on two commonly used temporal action localization datasets, THUMOS14 and ActivityNet1.2, to verify our method, which we achieve the stateof-the-art results. The experimental results show that our proposed cross-modal consensus module can produce more representative features for temporal action localization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 87 citations
- Weakly Supervised Temporal Action Localization via Representative Snippet Knowledge PropagationLinjiang Huang, Liang Wang, Hongsheng LiCVPR 2022 · 84 citations
- Likert Scoring with Grade Decoupling for Long-term Action AssessmentAngchi Xu, Ling-An Zeng, Wei-Shi ZhengCVPR 2022 · 41 citations
- Forcing the Whole Video as Background: An Adversarial Learning Strategy for Weakly Temporal Action LocalizationZiqiang Li, Yongxin Ge, Jiaruo Yu, Zhongming ChenACM MM 2022 · 24 citations
- Temporal Sentiment Localization: Listen and Look in Untrimmed VideosZhicheng Zhang, Jufeng YangACM MM 2022 · 19 citations
Builds on16
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Background Suppression Network for Weakly-Supervised Temporal Action LocalizationPilhyeon Lee, Youngjung Uh, Hyeran ByunAAAI 2020 · 234 citations
- 3C-Net: Category Count and Center Loss for Weakly-Supervised Action LocalizationSanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan, Ling ShaoICCV 2019 · 174 citations
- A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action LocalizationAshraful Islam, Chengjiang Long, Richard J. RadkeAAAI 2021 · 145 citations
- Hybrid Dynamic-static Context-aware Attention Network for Action Assessment in Long VideosLing-An Zeng, Fa-Ting Hong, Wei-Shi Zheng, Qi-Zhi Yu et al.ACM MM 2020 · 84 citations
Related papers
- Similar Modality Enhancement and Action Consistency Learning for Weakly Supervised Temporal Action LocalizationMaodong Li, Chao Zheng, Jian Wang, Bing LiAAAI 2025 · 2 citations
- Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksZiyi Liu, Le Wang, Qilin Zhang, Zhanning Gao et al.ICCV 2019 · 122 citations
- Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative LearningYuan Ji, Xu Jia, Huchuan Lu, Xiang RuanACM MM 2021 · 27 citations
- Learning Temporal Co-Attention Models for Unsupervised Video Action LocalizationGuoqiang Gong, Xinghan Wang, Yadong Mu, Qi TianCVPR 2020
- ACGNet: Action Complement Graph Network for Weakly-Supervised Temporal Action LocalizationZichen Yang, Jie Qin, Di HuangAAAI 2022 · 72 citations
