Hybrid Active Learning via Deep Clustering for Video Action Detection
Aayush Jung Rana, Yogesh S. Rawat
Abstract
In this work, we focus on reducing the annotation cost for video action detection which requires costly frame-wise dense annotations. We study a novel hybrid active learning (AL) strategy which performs efficient labeling using both intra-sample and inter-sample selection. The intra-sample selection leads to labeling of fewer frames in a video as opposed to inter-sample selection which operates at video level. This hybrid strategy reduces the annotation cost from two different aspects leading to significant labeling cost reduction. The proposed approach utilize Clustering-Aware Uncertainty Scoring (CLAUS), a novel label acquisition strategy which relies on both informativeness and diversity for sample selection. We also propose a novel Spatio-Temporal Weighted (STeW) loss formulation, which helps in model training under limited annotations. The proposed approach is evaluated on UCF-101-24 and J-HMDB-21 datasets demonstrating its effectiveness in significantly reducing the annotation cost where it consistently outperforms other baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55979a95-3a8f-4351-b6fb-3d3ecef5bdaaCited by top-tier papers6
- ActiveDC: Distribution Calibration for Active FinetuningWenshuai Xu, Zhenghui Hu, Yu Lu, Jinzhou Meng et al.CVPR 2024 · 5 citations
- Learning Group Activity Features Through Person Attribute PredictionChihiro Nakatani, Hiroaki Kawashima, Norimichi UkitaCVPR 2024 · 4 citations
- A²LC: Active and Automated Label Correction for Semantic SegmentationYoujin Jeon, Kyusik Cho, Suhan Woo, Euntai KimAAAI 2026 · 1 citation
- STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingAaryan Garg, Akash Kumar, Yogesh S. RawatCVPR 2025
- Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video GroundingAkash Kumar, Zsolt Kira, Yogesh S. RawatICLR 2025
Builds on8
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
- Active Learning for Deep Detection Neural NetworksHamed H. Aghdam, Abel Gonzalez-Garcia, Antonio M. López, Joost van de WeijerICCV 2019 · 155 citations
- Deep Reinforcement Active Learning for Human-in-the-Loop Person Re-IdentificationZimo Liu, Jingya Wang, Shaogang Gong, Dacheng Tao et al.ICCV 2019 · 117 citations
Related papers
- Are all Frames Equal? Active Sparse Labeling for Video Action DetectionAayush Jung Rana, Yogesh S. RawatNeurIPS 2022 · 16 citations
- Semi-supervised Active Learning for Video Action DetectionAyush Singh, Aayush Jung Rana, Akash Kumar, Shruti Vyas et al.AAAI 2024 · 22 citations
- Active Domain Adaptation via Clustering Uncertainty-weighted EmbeddingsViraj Prabhu, Arjun Chandrasekaran, Kate Saenko, Judy HoffmanICCV 2021 · 160 citations
- The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly DetectionAnusha Achaya, Hitesh Sapkota, Qi Yu, Xumin LiuCVPR 2026
- Semi-supervised Learning for Multi-label Video Action DetectionHongcheng Zhang, Xu Zhao, Dongqi WangACM MM 2022 · 10 citations
