End-to-End Semi-Supervised Learning for Video Action Detection
Akash Kumar, Yogesh Singh Rawat
Abstract
In this work, we focus on semi-supervised learning for video action detection which utilizes both labeled as well as unlabeled data. We propose a simple end-to-end consistency based approach which effectively utilizes the unlabeled data. Video action detection requires both, action class prediction as well as a spatio-temporal localization of actions. Therefore, we investigate two types of constraints, classification consistency, and spatio-temporal consistency. The presence of predominant background and static regions in a video makes it challenging to utilize spatio-temporal consistency for action detection. To address this, we propose two novel regularization constraints for spatio-temporal consistency; 1) temporal coherency, and 2) gradient smoothness. Both these aspects exploit the temporal continuity of action in videos and are found to be effective for utilizing unlabeled videos for action detection. We demonstrate the effectiveness of the proposed approach on two different action detection benchmark datasets, UCF101-24 and IHMDB-21. In addition, we also show the effectiveness of the proposed approach for video object segmentation on the Youtube-VOS which demonstrates its generalization capability The proposed approach achieves competitive performance by using merely 20% of annotations on UCF101-24 when compared with recent fully supervised methods. On UCF101-24, it improves the score by +8.9% and +11% at 0.5 f-mAP and v-mAP respectively, compared to supervised approach. The code and models will be made publicly available at: https://github.com/AKASH2907/End-to-End-Semi-Supervised-Learning-for-Video-Action-Detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc0edf45-6a6a-48b7-83db-99548237645cCited by top-tier papers9
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu et al.ICCV 2023 · 31 citations
- Semi-supervised Active Learning for Video Action DetectionAyush Singh, Aayush Jung Rana, Akash Kumar, Shruti Vyas et al.AAAI 2024 · 22 citations
- Are all Frames Equal? Active Sparse Labeling for Video Action DetectionAayush Jung Rana, Yogesh S. RawatNeurIPS 2022 · 16 citations
- Knowledge Guided Semi-supervised Learning for Quality Assessment of User Generated VideosShankhanil Mitra, Rajiv SoundararajanAAAI 2024 · 11 citations
- Mamba Only Glances Once (MOGO): A Lightweight Framework for Efficient Video Action DetectionYunqing Liu, Nan Zhang, Fangjun Wang, Kengo Murata et al.NeurIPS 2025 · 1 citation
Builds on12
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
- End-to-End Semi-Supervised Object Detection with Soft TeacherMengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang et al.ICCV 2021 · 622 citations
Related papers
- Stable Mean Teacher for Semi-supervised Video Action DetectionAkash Kumar, Sirshapan Mitra, Yogesh Singh RawatAAAI 2025 · 5 citations
- Semi-supervised Learning for Multi-label Video Action DetectionHongcheng Zhang, Xu Zhao, Dongqi WangACM MM 2022 · 10 citations
- SCT: Set Constrained Temporal Transformer for Set Supervised Action SegmentationMohsen Fayyaz, Jürgen GallCVPR 2020
- Learning from Temporal Gradient for Semi-supervised Action RecognitionJunfei Xiao, Longlong Jing, Lin Zhang, Ju He et al.CVPR 2022 · 77 citations
- Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action LocalizationJun-Tae Lee, Mihir Jain, Hyoungwoo Park, Sungrack YunICLR 2021 · 73 citations
