Intentional Evolutionary Learning for Untrimmed Videos with Long Tail Distribution
Yuxi Zhou, Xiujie Wang, Jianhua Zhang, Jiajia Wang, Jie Yu, Hao Zhou, Yi Gao, Shengyong Chen
Abstract
Human intention understanding in untrimmed videos aims to watch a natural video and predict what the person’s intention is. Currently, exploration of predicting human intentions in untrimmed videos is far from enough. On the one hand, untrimmed videos with mixed actions and backgrounds have a significant long-tail distribution with concept drift characteristics. On the other hand, most methods can only perceive instantaneous intentions, but cannot determine the evolution of intentions. To solve the above challenges, we propose a loss based on Instance Confidence and Class Accuracy (ICCA), which aims to alleviate the prediction bias caused by the long-tail distribution with concept drift characteristics in video streams. In addition, we propose an intention-oriented evolutionary learning method to determine the intention evolution pattern (from what action to what action) and the time of evolution (when the action evolves). We conducted extensive experiments on two untrimmed video datasets (THUMOS14 and ActivityNET v1.3), and our method has achieved excellent results compared to SOTA methods. The code and supplementary materials are available at https://github.com/Jennifer123www/UntrimmedVideo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Adaptive Graph Convolutional Recurrent Network for Traffic ForecastingLei Bai, Lina Yao, Can Li, Xianzhi Wang et al.NeurIPS 2020 · 2,206 citations
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 270 citations
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng et al.CVPR 2022 · 104 citations
- Equalization Loss for Long-Tailed Object RecognitionJingru Tan, Changbao Wang, Buyu Li, Quanquan Li et al.CVPR 2020
- Action Unit Memory Network for Weakly Supervised Temporal Action LocalizationWang Luo, Tianzhu Zhang, Wenfei Yang, Jingen Liu et al.CVPR 2021
Related papers
- Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed VideosXiaoyu Zhang, Haichao Shi, Changsheng Li, Peng LiAAAI 2020 · 37 citations
- CAG-QIL: Context-Aware Actionness Grouping via Q Imitation Learning for Online Temporal Action LocalizationHyolim Kang, Kyungmin Kim, Yumin Ko, Seon Joo KimICCV 2021 · 18 citations
- Modeling Temporal Concept Receptive Field Dynamically for Untrimmed Video AnalysisZhaobo Qi, Shuhui Wang, Chi Su, Li Su et al.ACM MM 2020 · 10 citations
- Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class ImbalanceSanchayan Santra, Vishal M. Chudasama, Pankaj Wasnik, Vineeth N. BalasubramanianCVPR 2025
- Concept Drift Detection for Multivariate Data Streams and Temporal Segmentation of Daylong Egocentric VideosPravin Nagar, Mansi Khemka, Chetan AroraACM MM 2020 · 9 citations
