Memory-and-Anticipation Transformer for Online Action Understanding
Jiahao Wang, Guo Chen, Yifei Huang, Limin Wang, Tong Lu
Abstract
Most existing forecasting systems are memory-based methods, which attempt to mimic human forecasting ability by employing various memory mechanisms and have progressed in temporal modeling for memory dependency. Nevertheless, an obvious weakness of this paradigm is that it can only model limited historical dependence and can not transcend the past. In this paper, we rethink the temporal dependence of event evolution and propose a novel memory-anticipation-based paradigm to model an entire temporal structure, including the past, present, and future. Based on this idea, we present Memory-and-Anticipation Transformer (MAT), a memoryanticipation-based approach, to address the online action detection and anticipation tasks. In addition, owing to the inherent superiority of MAT, it can process online action detection and anticipation tasks in a unified manner. The proposed MAT model is tested on four challenging benchmarks TVSeries, THUMOS'14, HDD, and EPIC-Kitchens-100, for online action detection and anticipation tasks, and it significantly outperforms all existing methods. Code is available at https://github.com/Echo0125/ Memory-and-Anticipation-Transformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 816b0714-55cb-4207-8a2c-e5a3b565098cCited by top-tier papers28
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- FB-BEV: BEV Representation from Forward-Backward View TransformationsZhiqi Li, Zhiding Yu, Wenhai Wang, Anima Anandkumar et al.ICCV 2023 · 144 citations
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang et al.CVPR 2024 · 95 citations
- UniFormerV2: Unlocking the Potential of Image ViTs for Video UnderstandingKunchang Li, Yali Wang, Yinan He, Yizhuo Li et al.ICCV 2023 · 85 citations
Builds on25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- Context-Enhanced Memory-Refined Transformer for Online Action DetectionZhanzhong Pang, Fadime Sener, Angela YaoCVPR 2025
- OadTR: Online Action Detection with TransformersXiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao et al.ICCV 2021 · 159 citations
- BiOMamba: Mamba-based Forward-Then-Backward Temporal Modeling for Online Action Detection and AnticipationSensen Wang, Yuehu Liu, Chi ZhangACM MM 2025
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li et al.NeurIPS 2021 · 196 citations
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis et al.ICCV 2019 · 201 citations
