Semi-supervised Active Learning for Video Action Detection
Ayush Singh, Aayush Jung Rana, Akash Kumar, Shruti Vyas, Yogesh Singh Rawat
摘要
In this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as unlabeled data along with informative sample selection for action detection. Video action detection requires spatio-temporal localization along with classification, which poses several challenges for both active learning (informative sample selection) as well as semi-supervised learning (pseudo label generation). First, we propose NoiseAug, a simple augmentation strategy which effectively selects informative samples for video action detection. Next, we propose fft-attention, a novel technique based on high-pass filtering which enables effective utilization of pseudo label for SSL in video action detection by emphasizing on relevant activity region within a video. We evaluate the proposed approach on three different benchmark datasets, UCF-101-24, JHMDB-21, and Youtube-VOS. First, we demonstrate its effectiveness on video action detection where the proposed approach outperforms prior works in semi-supervised and weakly-supervised learning along with several baseline approaches in both UCF101-24 and JHMDB-21. Next, we also show its effectiveness on Youtube-VOS for video object segmentation demonstrating its generalization capability for other dense prediction tasks in videos. The code and models is publicly available at: https://github.com/AKASH2907/semi-sup-active-learning .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Stable Mean Teacher for Semi-supervised Video Action DetectionAkash Kumar, Sirshapan Mitra, Yogesh Singh RawatAAAI 2025 · 被引用 5 次
- STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingAaryan Garg, Akash Kumar, Yogesh S. RawatCVPR 2025
- Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video GroundingAkash Kumar, Zsolt Kira, Yogesh S. RawatICLR 2025
- Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang 等SIGIR 2026
它引用的顶会 Paper12
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Active Learning for Deep Detection Neural NetworksHamed H. Aghdam, Abel Gonzalez-Garcia, Antonio M. López, Joost van de WeijerICCV 2019 · 被引用 155 次
- Deep Reinforcement Active Learning for Human-in-the-Loop Person Re-IdentificationZimo Liu, Jingya Wang, Shaogang Gong, Dacheng Tao 等ICCV 2019 · 被引用 117 次
- Rethinking Pseudo Labels for Semi-supervised Object DetectionHengduo Li, Zuxuan Wu, Abhinav Shrivastava, Larry S. DavisAAAI 2022 · 被引用 105 次
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen 等CVPR 2022 · 被引用 77 次
相关 Paper
- End-to-End Semi-Supervised Learning for Video Action DetectionAkash Kumar, Yogesh Singh RawatCVPR 2022 · 被引用 31 次
- Are all Frames Equal? Active Sparse Labeling for Video Action DetectionAayush Jung Rana, Yogesh S. RawatNeurIPS 2022 · 被引用 16 次
- Hybrid Active Learning via Deep Clustering for Video Action DetectionAayush Jung Rana, Yogesh S. RawatCVPR 2023
- Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationKun Xia, Le Wang, Sanping Zhou, Gang Hua 等ICCV 2023 · 被引用 16 次
- Semi-supervised Learning for Multi-label Video Action DetectionHongcheng Zhang, Xu Zhao, Dongqi WangACM MM 2022 · 被引用 10 次
