MiniROAD: Minimal RNN Framework for Online Action Detection
Joungbin An, Hyolim Kang, Su Ho Han, Ming-Hsuan Yang, Seon Joo Kim
Abstract
Online Action Detection (OAD) is the task of identifying actions in streaming videos without access to future frames. Much effort has been devoted to effectively capturing longrange dependencies, with transformers receiving the spotlight for their ability to capture long-range temporal structures. In contrast, RNNs have received less attention lately, due to their lower performance compared to recent methods that utilize transformers. In this paper, we investigate the underlying reasons for the inferior performance of RNNs compared to transformer-based algorithms. Our findings indicate that the discrepancy between training and inference is the primary hindrance to the effective training of RNNs. To address this, we propose applying non-uniform weights to the loss computed at each time step, which allows the RNN model to learn from the predictions made in an environment that better resembles the inference stage. Extensive experiments on three benchmark datasets, THU-MOS, TVSeries, and FineAction demonstrate that a minimal RNN-based model trained with the proposed methodology performs equally or better than the existing best methods with a significant increase in efficiency. The code is available at https://github.com/jbistanbul/MiniROAD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1dd8006-55f6-4eda-9d13-a5430dad5d8eCited by top-tier papers7
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 36 citations
- Hierarchical Event Memory for Accurate and Low-Latency Online Video Temporal GroundingMinghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang et al.ICCV 2025 · 3 citations
- Open-Ended Hierarchical Streaming Video Understanding with Vision Language ModelsHyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi et al.ICCV 2025 · 2 citations
- Online Generic Event Boundary DetectionHyungrok Jung, Daneul Kim, Seunggyun Lim, Jeany Son et al.ICCV 2025
- Backtrace Mamba: Reviving Critical Temporal Contexts via Hierarchical Memory Compression for Online Action DetectionSu Yan, Jiahua Li, Kun Wei, Cheng DengAAAI 2026
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando et al.ICML 2023 · 474 citations
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis et al.ICCV 2019 · 201 citations
Related papers
- OadTR: Online Action Detection with TransformersXiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao et al.ICCV 2021 · 159 citations
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li et al.NeurIPS 2021 · 196 citations
- Context-Enhanced Memory-Refined Transformer for Online Action DetectionZhanzhong Pang, Fadime Sener, Angela YaoCVPR 2025
- Learning to Discriminate Information for Online Action DetectionHyunjun Eun, Jinyoung Moon, Jongyoul Park, Chanho Jung et al.CVPR 2020
- Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?Qingsong Zhao, Yi Wang, Jilan Xu, Yinan He et al.NeurIPS 2024 · 16 citations
