E2E-LOAD: End-to-End Long-form Online Action Detection
Shuqiang Cao, Weixin Luo, Bairui Wang, Wei Zhang, Lin Ma
Abstract
Recently, feature-based methods for Online Action Detection (OAD) have been gaining traction. However, these methods are constrained by their fixed backbone design, which fails to leverage the potential benefits of a trainable backbone. This paper introduces an end-toend learning network that revises these approaches, incorporating a backbone network design that improves effectiveness and efficiency. Our proposed model utilizes a shared initial spatial model for all frames and maintains an extended sequence cache, which enables low-cost inference. We promote an asymmetric spatiotemporal model that caters to long-form and short-form modeling. Additionally, we propose an innovative and efficient inference mechanism that accelerates extensive spatiotemporal exploration. Through comprehensive ablation studies and experiments, we validate the performance and efficiency of our proposed method. Remarkably, we achieve an end-to-end learning OAD of 17.3 (+12.6) FPS with 72.4% (+1.2%), 90.3% (+0.7%), and 48.1% (+26.0%) mAP on THMOUS'14, TVSeries, and HDD, respectively. The source code is available at https://github. com/sqiangcao99/E2E-LOAD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8269fd46-9ef9-4736-81e8-c753cebd67c2Cited by top-tier papers4
- Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?Qingsong Zhao, Yi Wang, Jilan Xu, Yinan He et al.NeurIPS 2024 · 16 citations
- OnlineTAS: An Online Baseline for Temporal Action SegmentationQing Zhong, Guodong Ding, Angela YaoNeurIPS 2024 · 15 citations
- Hierarchical Event Memory for Accurate and Low-Latency Online Video Temporal GroundingMinghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang et al.ICCV 2025 · 3 citations
- Backtrace Mamba: Reviving Critical Temporal Contexts via Hierarchical Memory Compression for Online Action DetectionSu Yan, Jiahua Li, Kun Wei, Cheng DengAAAI 2026
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- An Empirical Study of End-to-End Temporal Action DetectionXiaolong Liu, Song Bai, Xiang BaiCVPR 2022 · 72 citations
- End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesShuming Liu, Chen-Lin Zhang, Chen Zhao, Bernard GhanemCVPR 2024 · 35 citations
- Watch Only Once: An End-to-End Video Action Detection FrameworkShoufa Chen, Peize Sun, Enze Xie, Chongjian Ge et al.ICCV 2021 · 68 citations
- OadTR: Online Action Detection with TransformersXiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao et al.ICCV 2021 · 159 citations
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu et al.ICCV 2023 · 31 citations
