Semi-supervised Active Learning for Video Action Detection
Ayush Singh, Aayush Jung Rana, Akash Kumar, Shruti Vyas, Yogesh Singh Rawat
Abstract
In this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as unlabeled data along with informative sample selection for action detection. Video action detection requires spatio-temporal localization along with classification, which poses several challenges for both active learning (informative sample selection) as well as semi-supervised learning (pseudo label generation). First, we propose NoiseAug, a simple augmentation strategy which effectively selects informative samples for video action detection. Next, we propose fft-attention, a novel technique based on high-pass filtering which enables effective utilization of pseudo label for SSL in video action detection by emphasizing on relevant activity region within a video. We evaluate the proposed approach on three different benchmark datasets, UCF-101-24, JHMDB-21, and Youtube-VOS. First, we demonstrate its effectiveness on video action detection where the proposed approach outperforms prior works in semi-supervised and weakly-supervised learning along with several baseline approaches in both UCF101-24 and JHMDB-21. Next, we also show its effectiveness on Youtube-VOS for video object segmentation demonstrating its generalization capability for other dense prediction tasks in videos. The code and models is publicly available at: https://github.com/AKASH2907/semi-sup-active-learning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7927b8af-a43e-46f6-8aea-e813d8287199Cited by top-tier papers4
- Stable Mean Teacher for Semi-supervised Video Action DetectionAkash Kumar, Sirshapan Mitra, Yogesh Singh RawatAAAI 2025 · 5 citations
- STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingAaryan Garg, Akash Kumar, Yogesh S. RawatCVPR 2025
- Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video GroundingAkash Kumar, Zsolt Kira, Yogesh S. RawatICLR 2025
- Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang et al.SIGIR 2026
Builds on12
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Active Learning for Deep Detection Neural NetworksHamed H. Aghdam, Abel Gonzalez-Garcia, Antonio M. López, Joost van de WeijerICCV 2019 · 155 citations
- Deep Reinforcement Active Learning for Human-in-the-Loop Person Re-IdentificationZimo Liu, Jingya Wang, Shaogang Gong, Dacheng Tao et al.ICCV 2019 · 117 citations
- Rethinking Pseudo Labels for Semi-supervised Object DetectionHengduo Li, Zuxuan Wu, Abhinav Shrivastava, Larry S. DavisAAAI 2022 · 105 citations
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen et al.CVPR 2022 · 77 citations
Related papers
- End-to-End Semi-Supervised Learning for Video Action DetectionAkash Kumar, Yogesh Singh RawatCVPR 2022 · 31 citations
- Are all Frames Equal? Active Sparse Labeling for Video Action DetectionAayush Jung Rana, Yogesh S. RawatNeurIPS 2022 · 16 citations
- Hybrid Active Learning via Deep Clustering for Video Action DetectionAayush Jung Rana, Yogesh S. RawatCVPR 2023
- Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationKun Xia, Le Wang, Sanping Zhou, Gang Hua et al.ICCV 2023 · 16 citations
- Semi-supervised Learning for Multi-label Video Action DetectionHongcheng Zhang, Xu Zhao, Dongqi WangACM MM 2022 · 10 citations
