Stable Mean Teacher for Semi-supervised Video Action Detection
Akash Kumar, Sirshapan Mitra, Yogesh Singh Rawat
摘要
In this work, we focus on semi-supervised learning for video action detection. Video action detection requires spatio-temporal localization in addition to classification, and a limited amount of labels makes the model prone to unreliable predictions. We present Stable Mean Teacher, a simple end-to-end student-teacher-based framework that benefits from improved and temporally consistent pseudo labels. It relies on a novel ErrOr Recovery (EoR) module, which learns from students' mistakes on labeled samples and transfers this to the teacher to improve pseudo labels for unlabeled samples. Moreover, existing spatio-temporal losses do not take temporal coherency into account and are prone to temporal inconsistencies. To overcome this, we present Difference of Pixels (DoP), a simple and novel constraint focused on temporal consistency, which leads to coherent temporal detections. We evaluate our approach on four different spatio-temporal detection benchmarks: UCF101-24, JHMDB21, AVA, and Youtube-VOS. Our approach outperforms the supervised baselines for action detection by an average margin of 23.5% on UCF101-24, 16% on JHMDB21, and 3.3% on AVA. Using merely 10% and 20% of data, it provides a competitive performance compared to the supervised baseline trained on 100% annotations on UCF101-24 and JHMDB21 respectively. We further evaluate its effectiveness on AVA for scaling to large-scale datasets and Youtube-VOS for video object segmentation, demonstrating its generalization capability to other tasks in the video domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Mamba Only Glances Once (MOGO): A Lightweight Framework for Efficient Video Action DetectionYunqing Liu, Nan Zhang, Fangjun Wang, Kengo Murata 等NeurIPS 2025 · 被引用 1 次
- STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingAaryan Garg, Akash Kumar, Yogesh S. RawatCVPR 2025
- Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video GroundingAkash Kumar, Zsolt Kira, Yogesh S. RawatICLR 2025
- Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang 等SIGIR 2026
- Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action DetectionSa Zhu, Wanqian Zhang, Lin Wang, Xiaohua Chen 等CVPR 2026
它引用的顶会 Paper16
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- End-to-End Semi-Supervised Object Detection with Soft TeacherMengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang 等ICCV 2021 · 被引用 622 次
- Dense Learning based Semi-Supervised Object DetectionBinghui Chen, Pengyu Li, Xiang Chen, Biao Wang 等CVPR 2022 · 被引用 80 次
- Learning from Temporal Gradient for Semi-supervised Action RecognitionJunfei Xiao, Longlong Jing, Lin Zhang, Ju He 等CVPR 2022 · 被引用 77 次
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen 等CVPR 2022 · 被引用 77 次
相关 Paper
- End-to-End Semi-Supervised Learning for Video Action DetectionAkash Kumar, Yogesh Singh RawatCVPR 2022 · 被引用 31 次
- TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action RecognitionIshan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak ShahCVPR 2023
- Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action LocalizationKun Xia, Le Wang, Sanping Zhou, Gang Hua 等ICCV 2023 · 被引用 16 次
- Semi-supervised Learning for Multi-label Video Action DetectionHongcheng Zhang, Xu Zhao, Dongqi WangACM MM 2022 · 被引用 10 次
- Semi-supervised Active Learning for Video Action DetectionAyush Singh, Aayush Jung Rana, Akash Kumar, Shruti Vyas 等AAAI 2024 · 被引用 22 次
