DirecFormer: A Directed Attention in Transformer Approach to Robust Action Recognition
Thanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo, Son Lam Phung, Xin Li, Khoa Luu
Abstract
Human action recognition has recently become one of the popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in the task of video action recognition with competitive results. However, these methods have suffered some fundamental limitations such as lack of robustness and generalization, e.g., how does the temporal ordering of video frames affect the recognition results? This work presents a novel end-to-end Transformer-based Directed Attention (Direc-Former) framework <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> The implementation of DirecFormer is available at https://github.com/uark-cviu/DirecFormer for robust action recognition. The method takes a simple but novel perspective of Transformer-based approach to understand the right order of sequence actions. Therefore, the contributions of this work are three-fold. Firstly, we introduce the problem of ordered temporal learning issues to the action recognition problem. Secondly, a new Directed Attention mechanism is introduced to understand and provide attentions to human actions in the right order. Thirdly, we introduce the conditional dependency in action sequence modeling that includes orders and classes. The proposed approach consistently achieves the state-of-the-art (SOTA) results compared with the recent action recognition methods [4, 18, 72, 74]. on three standard large-scale benchmarks, i.e. Jester, Kinetics-400 and Something-Something-V2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a309ec2b-7f19-4789-bbf6-a827d0139828Cited by top-tier papers13
- A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action RecognitionAndong Deng, Taojiannan Yang, Chen ChenICCV 2023 · 18 citations
- Insect-Foundation: A Foundation Model and Large-Scale 1M Dataset for Visual Insect UnderstandingHoang-Quan Nguyen, Thanh-Dat Truong, Xuan-Bac Nguyen, Ashley Dowling et al.CVPR 2024 · 17 citations
- Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed TransformerWenhan Wu, Ce Zheng, Zihao Yang, Chen Chen et al.ACM MM 2024 · 16 citations
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren et al.NeurIPS 2024 · 13 citations
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion LearningThanh-Dat Truong, Christophe Bobda, Nitin Agarwal, Khoa LuuNeurIPS 2025 · 6 citations
Builds on22
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
Related papers
- Recurring the Transformer for Video Action RecognitionJiewen Yang, Xingbo Dong, Liujun Liu, Chao Zhang et al.CVPR 2022 · 119 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- Spatio-Temporal Fusion for Human Action Recognition via Joint Trajectory GraphYaolin Zheng, Hongbo Huang, Xiuying Wang, Xiaoxu Yan et al.AAAI 2024 · 22 citations
- Deformable Video TransformerJue Wang, Lorenzo TorresaniCVPR 2022 · 40 citations
- Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionWei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du et al.ACM MM 2022 · 21 citations
