Intentional Evolutionary Learning for Untrimmed Videos with Long Tail Distribution
Yuxi Zhou, Xiujie Wang, Jianhua Zhang, Jiajia Wang, Jie Yu, Hao Zhou, Yi Gao, Shengyong Chen
摘要
Human intention understanding in untrimmed videos aims to watch a natural video and predict what the person’s intention is. Currently, exploration of predicting human intentions in untrimmed videos is far from enough. On the one hand, untrimmed videos with mixed actions and backgrounds have a significant long-tail distribution with concept drift characteristics. On the other hand, most methods can only perceive instantaneous intentions, but cannot determine the evolution of intentions. To solve the above challenges, we propose a loss based on Instance Confidence and Class Accuracy (ICCA), which aims to alleviate the prediction bias caused by the long-tail distribution with concept drift characteristics in video streams. In addition, we propose an intention-oriented evolutionary learning method to determine the intention evolution pattern (from what action to what action) and the time of evolution (when the action evolves). We conducted extensive experiments on two untrimmed video datasets (THUMOS14 and ActivityNET v1.3), and our method has achieved excellent results compared to SOTA methods. The code and supplementary materials are available at https://github.com/Jennifer123www/UntrimmedVideo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Adaptive Graph Convolutional Recurrent Network for Traffic ForecastingLei Bai, Lina Yao, Can Li, Xianzhi Wang 等NeurIPS 2020 · 被引用 2,206 次
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 被引用 270 次
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng 等CVPR 2022 · 被引用 104 次
- Equalization Loss for Long-Tailed Object RecognitionJingru Tan, Changbao Wang, Buyu Li, Quanquan Li 等CVPR 2020
- Action Unit Memory Network for Weakly Supervised Temporal Action LocalizationWang Luo, Tianzhu Zhang, Wenfei Yang, Jingen Liu 等CVPR 2021
相关 Paper
- Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed VideosXiaoyu Zhang, Haichao Shi, Changsheng Li, Peng LiAAAI 2020 · 被引用 37 次
- CAG-QIL: Context-Aware Actionness Grouping via Q Imitation Learning for Online Temporal Action LocalizationHyolim Kang, Kyungmin Kim, Yumin Ko, Seon Joo KimICCV 2021 · 被引用 18 次
- Modeling Temporal Concept Receptive Field Dynamically for Untrimmed Video AnalysisZhaobo Qi, Shuhui Wang, Chi Su, Li Su 等ACM MM 2020 · 被引用 10 次
- Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class ImbalanceSanchayan Santra, Vishal M. Chudasama, Pankaj Wasnik, Vineeth N. BalasubramanianCVPR 2025
- Concept Drift Detection for Multivariate Data Streams and Temporal Segmentation of Daylong Egocentric VideosPravin Nagar, Mansi Khemka, Chetan AroraACM MM 2020 · 被引用 9 次
