Are all Frames Equal? Active Sparse Labeling for Video Action Detection
Aayush Jung Rana, Yogesh S. Rawat
摘要
Video action detection requires annotations at every frame, which drastically increases the labeling cost. In this work, we focus on efficient labeling of videos for action detection to minimize this cost. We propose active sparse labeling (ASL) , a novel active learning strategy for video action detection. Sparse labeling will reduce the annotation cost but poses two main challenges; 1) how to estimate the utility of annotating a single frame for action detection as detection is performed at video level?, and 2) how these sparse labels can be used for action detection which require annotations on all the frames? This work attempts to address these challenges within a simple active learning framework. For the first challenge, we propose a novel frame-level scoring mechanism aimed at selecting most informative frames in a video. Next, we introduce a novel loss formulation which enables training of action detection model with these sparsely selected frames. We evaluate the proposed approach on two different action detection benchmark datasets, UCF-101-24 and J-HMDB-21, and observed that active sparse labeling can be very effective in saving annotation costs. We demonstrate that the proposed approach performs better than random selection, outperforming all other baselines, with performance comparable to supervised approach using merely 10% annotations. Project details available at https://www.crcv.ucf.edu/research/projects/active-sparse-labeling-for-video-action-detection/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Semi-supervised Active Learning for Video Action DetectionAyush Singh, Aayush Jung Rana, Akash Kumar, Shruti Vyas 等AAAI 2024 · 被引用 22 次
- Sample Less, Learn More: Efficient Action Recognition via Frame Feature RestorationHarry Cheng, Yangyang Guo, Liqiang Nie, Zhiyong Cheng 等ACM MM 2023 · 被引用 7 次
- STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingAaryan Garg, Akash Kumar, Yogesh S. RawatCVPR 2025
它引用的顶会 Paper9
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 被引用 630 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
- Active Learning for Deep Detection Neural NetworksHamed H. Aghdam, Abel Gonzalez-Garcia, Antonio M. López, Joost van de WeijerICCV 2019 · 被引用 155 次
- Active Learning for Deep Object Detection via Probabilistic ModelingJiwoong Choi, Ismail Elezi, Hyuk-Jae Lee, Clément Farabet 等ICCV 2021 · 被引用 144 次
- Deep Reinforcement Active Learning for Human-in-the-Loop Person Re-IdentificationZimo Liu, Jingya Wang, Shaogang Gong, Dacheng Tao 等ICCV 2019 · 被引用 117 次
相关 Paper
- Hybrid Active Learning via Deep Clustering for Video Action DetectionAayush Jung Rana, Yogesh S. RawatCVPR 2023
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See 等AAAI 2020 · 被引用 18 次
- End-to-End Semi-Supervised Learning for Video Action DetectionAkash Kumar, Yogesh Singh RawatCVPR 2022 · 被引用 31 次
- Weakly Supervised Action Selection Learning in VideoJunwei Ma, Satya Krishna Gorti, Maksims Volkovs, Guangwei YuCVPR 2021
- Action-Agnostic Point-Level Supervision for Temporal Action DetectionShuhei M. Yoshida, Takashi Shibata, Makoto Terao, Takayuki Okatani 等AAAI 2025 · 被引用 6 次
