Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed Videos
Xiaoyu Zhang, Haichao Shi, Changsheng Li, Peng Li
摘要
Weakly supervised action recognition and localization for untrimmed videos is a challenging problem with extensive applications. The overwhelming irrelevant background contents in untrimmed videos severely hamper effective identification of actions of interest. In this paper, we propose a novel multi-instance multi-label modeling network based on spatio-temporal pre-trimming to recognize actions and locate corresponding frames in untrimmed videos. Motivated by the fact that person is the key factor in a human action, we spatially and temporally segment each untrimmed video into person-centric clips with pose estimation and tracking techniques. Given the bag-of-instances structure associated with video-level labels, action recognition is naturally formulated as a multi-instance multi-label learning problem. The network is optimized iteratively with selective coarse-to-fine pre-trimming based on instance-label activation. After convergence, temporal localization is further achieved with local-global temporal class activation map. Extensive experiments are conducted on two benchmark datasets, i.e. THUMOS14 and ActivityNet1.3, and experimental results clearly corroborate the efficacy of our method when compared with the state-of-the-arts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Weakly-supervised Temporal Action Localization by Uncertainty ModelingPilhyeon Lee, Jinglu Wang, Yan Lu, Hyeran ByunAAAI 2021 · 被引用 141 次
- Cross-modal Consensus Network for Weakly Supervised Temporal Action LocalizationFa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan 等ACM MM 2021 · 被引用 104 次
- Learning Action Completeness from Points for Weakly-supervised Temporal Action LocalizationPilhyeon Lee, Hyeran ByunICCV 2021 · 被引用 81 次
- Multi-label Self Knowledge DistillationXucong Wang, Pengkun Wang, Shurui Zhang, Miao Fang 等AAAI 2025 · 被引用 2 次
- Two-Stream Networks for Weakly-Supervised Temporal Action Localization with Semantic-Aware MechanismsYu Wang, Yadong Li, Hongbin WangCVPR 2023
它引用的顶会 Paper1
相关 Paper
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng 等CVPR 2022 · 被引用 104 次
- Relational Prototypical Network for Weakly Supervised Temporal Action LocalizationLinjiang Huang, Yan Huang, Wanli Ouyang, Liang WangAAAI 2020 · 被引用 71 次
- PivoTAL: Prior-Driven Supervision for Weakly-Supervised Temporal Action LocalizationMamshad Nayeem Rizve, Gaurav Mittal, Ye Yu, Matthew Hall 等CVPR 2023
- Learning Temporal Co-Attention Models for Unsupervised Video Action LocalizationGuoqiang Gong, Xinghan Wang, Yadong Mu, Qi TianCVPR 2020
- Action Unit Memory Network for Weakly Supervised Temporal Action LocalizationWang Luo, Tianzhu Zhang, Wenfei Yang, Jingen Liu 等CVPR 2021
