StartNet: Online Detection of Action Start in Untrimmed Videos
Mingfei Gao, Mingze Xu, Larry Davis, Richard Socher, Caiming Xiong
Abstract
We propose StartNet to address Online Detection of Action Start (ODAS) where action starts and their associated categories are detected in untrimmed, streaming videos. Previous methods aim to localize action starts by learning feature representations that can directly separate the start point from its preceding background. It is challenging due to the subtle appearance difference near the action starts and the lack of training data. Instead, StartNet decomposes ODAS into two stages: action classification (using ClsNet) and start point localization (using LocNet). ClsNet focuses on per-frame labeling and predicts action score distributions online. Based on the predicted action scores of the past and current frames, LocNet conducts class-agnostic start detection by optimizing long-term localization rewards using policy gradient methods. The proposed framework is validated on two large-scale datasets, THUMOS’14 and ActivityNet. The experimental results show that StartNet significantly outperforms the state-of-the-art by 15%-30% p-mAP under the offset tolerance of 1-10 seconds on THUMOS’14, and achieves comparable performance on ActivityNet with 10 times smaller time offset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li et al.NeurIPS 2021 · 196 citations
- OadTR: Online Action Detection with TransformersXiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao et al.ICCV 2021 · 159 citations
- Weakly-Supervised Online Action Segmentation in Multi-View Instructional VideosReza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Chiho Choi et al.CVPR 2022 · 22 citations
- CAG-QIL: Context-Aware Actionness Grouping via Q Imitation Learning for Online Temporal Action LocalizationHyolim Kang, Kyungmin Kim, Yumin Ko, Seon Joo KimICCV 2021 · 18 citations
- E2E-LOAD: End-to-End Long-form Online Action DetectionShuqiang Cao, Weixin Luo, Bairui Wang, Wei Zhang et al.ICCV 2023 · 12 citations
Related papers
- WOAD: Weakly Supervised Online Action Detection in Untrimmed VideosMingfei Gao, Yingbo Zhou, Ran Xu, Richard Socher et al.CVPR 2021
- ActionBytes: Learning From Trimmed Videos to Localize ActionsMihir Jain, Amir Ghodrati, Cees G. M. SnoekCVPR 2020
- Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksZiyi Liu, Le Wang, Qilin Zhang, Zhanning Gao et al.ICCV 2019 · 122 citations
- Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextZiyi Liu, Le Wang, Wei Tang, Junsong Yuan et al.AAAI 2021 · 28 citations
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis et al.ICCV 2019 · 201 citations
