Annotation-Efficient Untrimmed Video Action Recognition
Yixiong Zou, Shanghang Zhang, Guangyao Chen, Yonghong Tian, Kurt Keutzer, José M. F. Moura
摘要
Deep learning has achieved great success in recognizing video actions, but the collection and annotation of training data are still quite laborious, which mainly lies in two aspects: (1) the amount of required annotated data is large; (2) temporally annotating the location of each action is time-consuming. Works such as few-shot learning or untrimmed video recognition have been proposed to handle either one aspect or the other. However, very few existing works can handle both issues simultaneously. In this paper, we target a new problem, Annotation-Efficient Video Recognition, to reduce the requirement of annotations for both large amount of samples and the action location. Such problem is challenging due to two aspects: (1) the untrimmed videos only have weak supervision; (2) video segments not relevant to current actions of interests (background, BG) could contain actions of interests (foreground, FG) in novel classes, which is a widely existing phenomenon but has rarely been studied in few-shot untrimmed video recognition. To achieve this goal, by analyzing the property of BG, we categorize BG into informative BG (IBG) and non-informative BG (NBG), and we propose (1) an open-set detection based method to find the NBG and FG, (2) a contrastive learning method to learn IBG and distinguish NBG in a self-supervised way, and (3) a self-weighting mechanism for the better distinguishing of IBG and FG. Extensive experiments on ActivityNet v1.2 and ActivityNet v1.3 verify the rationale and effectiveness of the proposed methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- OpenTAL: Towards Open Set Temporal Action LocalizationWentao Bao, Qi Yu, Yu KongCVPR 2022 · 被引用 30 次
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging BenchmarkDeheng Zhang, Yuqian Fu, Runyi Yang, Yang Miao 等ICLR 2026 · 被引用 19 次
- Revisiting Mid-Level Patterns for Cross-Domain Few-Shot RecognitionYixiong Zou, Shanghang Zhang, Jianpeng Yu, Yonghong Tian 等ACM MM 2021 · 被引用 12 次
- Learning Unknowns from Unknowns: Diversified Negative Prototypes Generator for Few-shot Open-Set RecognitionZhenyu Zhang, Guangyao Chen, Yixiong Zou, Yuhua Li 等ACM MM 2024 · 被引用 7 次
- MICM: Rethinking Unsupervised Pretraining for Enhanced Few-shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Zhimeng Huang 等ACM MM 2024 · 被引用 5 次
它引用的顶会 Paper6
- Background Suppression Network for Weakly-Supervised Temporal Action LocalizationPilhyeon Lee, Youngjung Uh, Hyeran ByunAAAI 2020 · 被引用 234 次
- Weakly-Supervised Action Localization With Background ModelingPhuc Xuan Nguyen, Deva Ramanan, Charless C. FowlkesICCV 2019 · 被引用 176 次
- Learning Compositional Representations for Few-Shot RecognitionPavel Tokmakov, Yu-Xiong Wang, Martial HebertICCV 2019 · 被引用 133 次
- Compositional Few-Shot Recognition with Primitive Discovery and EnhancingYixiong Zou, Shanghang Zhang, Ke Chen, Yonghong Tian 等ACM MM 2020 · 被引用 30 次
- METAL: Minimum Effort Temporal Activity Localization in Untrimmed VideosDa Zhang, Xiyang Dai, Yuan-Fang WangCVPR 2020
相关 Paper
- Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed VideosXiaoyu Zhang, Haichao Shi, Changsheng Li, Peng LiAAAI 2020 · 被引用 37 次
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu 等ACM MM 2020 · 被引用 97 次
- Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action LocalizationJingjing Li, Tianyu Yang, Wei Ji, Jue Wang 等CVPR 2022 · 被引用 57 次
- Relational Prototypical Network for Weakly Supervised Temporal Action LocalizationLinjiang Huang, Yan Huang, Wanli Ouyang, Liang WangAAAI 2020 · 被引用 71 次
- CDFSL-V: Cross-Domain Few-Shot Learning for VideosSarinda Samarasinghe, Mamshad Nayeem Rizve, Navid Kardan, Mubarak ShahICCV 2023 · 被引用 17 次
