Annotation-Efficient Untrimmed Video Action Recognition
Yixiong Zou, Shanghang Zhang, Guangyao Chen, Yonghong Tian, Kurt Keutzer, José M. F. Moura
Abstract
Deep learning has achieved great success in recognizing video actions, but the collection and annotation of training data are still quite laborious, which mainly lies in two aspects: (1) the amount of required annotated data is large; (2) temporally annotating the location of each action is time-consuming. Works such as few-shot learning or untrimmed video recognition have been proposed to handle either one aspect or the other. However, very few existing works can handle both issues simultaneously. In this paper, we target a new problem, Annotation-Efficient Video Recognition, to reduce the requirement of annotations for both large amount of samples and the action location. Such problem is challenging due to two aspects: (1) the untrimmed videos only have weak supervision; (2) video segments not relevant to current actions of interests (background, BG) could contain actions of interests (foreground, FG) in novel classes, which is a widely existing phenomenon but has rarely been studied in few-shot untrimmed video recognition. To achieve this goal, by analyzing the property of BG, we categorize BG into informative BG (IBG) and non-informative BG (NBG), and we propose (1) an open-set detection based method to find the NBG and FG, (2) a contrastive learning method to learn IBG and distinguish NBG in a self-supervised way, and (3) a self-weighting mechanism for the better distinguishing of IBG and FG. Extensive experiments on ActivityNet v1.2 and ActivityNet v1.3 verify the rationale and effectiveness of the proposed methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- OpenTAL: Towards Open Set Temporal Action LocalizationWentao Bao, Qi Yu, Yu KongCVPR 2022 · 30 citations
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging BenchmarkDeheng Zhang, Yuqian Fu, Runyi Yang, Yang Miao et al.ICLR 2026 · 19 citations
- Revisiting Mid-Level Patterns for Cross-Domain Few-Shot RecognitionYixiong Zou, Shanghang Zhang, Jianpeng Yu, Yonghong Tian et al.ACM MM 2021 · 12 citations
- Learning Unknowns from Unknowns: Diversified Negative Prototypes Generator for Few-shot Open-Set RecognitionZhenyu Zhang, Guangyao Chen, Yixiong Zou, Yuhua Li et al.ACM MM 2024 · 7 citations
- MICM: Rethinking Unsupervised Pretraining for Enhanced Few-shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Zhimeng Huang et al.ACM MM 2024 · 5 citations
Builds on6
- Background Suppression Network for Weakly-Supervised Temporal Action LocalizationPilhyeon Lee, Youngjung Uh, Hyeran ByunAAAI 2020 · 234 citations
- Weakly-Supervised Action Localization With Background ModelingPhuc Xuan Nguyen, Deva Ramanan, Charless C. FowlkesICCV 2019 · 176 citations
- Learning Compositional Representations for Few-Shot RecognitionPavel Tokmakov, Yu-Xiong Wang, Martial HebertICCV 2019 · 133 citations
- Compositional Few-Shot Recognition with Primitive Discovery and EnhancingYixiong Zou, Shanghang Zhang, Ke Chen, Yonghong Tian et al.ACM MM 2020 · 30 citations
- METAL: Minimum Effort Temporal Activity Localization in Untrimmed VideosDa Zhang, Xiyang Dai, Yuan-Fang WangCVPR 2020
Related papers
- Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed VideosXiaoyu Zhang, Haichao Shi, Changsheng Li, Peng LiAAAI 2020 · 37 citations
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu et al.ACM MM 2020 · 97 citations
- Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action LocalizationJingjing Li, Tianyu Yang, Wei Ji, Jue Wang et al.CVPR 2022 · 57 citations
- Relational Prototypical Network for Weakly Supervised Temporal Action LocalizationLinjiang Huang, Yan Huang, Wanli Ouyang, Liang WangAAAI 2020 · 71 citations
- CDFSL-V: Cross-Domain Few-Shot Learning for VideosSarinda Samarasinghe, Mamshad Nayeem Rizve, Navid Kardan, Mubarak ShahICCV 2023 · 17 citations
