Hierarchical Self-Attention Network for Action Localization in Videos
Rizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien Fang
Abstract
This paper presents a novel Hierarchical Self-Attention Network (HISAN) to generate spatial-temporal tubes for action localization in videos. The essence of HISAN is to combine the two-stream convolutional neural network (CNN) with hierarchical bidirectional self-attention mechanism, which comprises of two levels of bidirectional self-attention to efficaciously capture both of the long-term temporal dependency information and spatial context information to render more precise action localization. Also, a sequence rescoring (SR) algorithm is employed to resolve the dilemma of inconsistent detection scores incurred by occlusion or background clutter. Moreover, a new fusion scheme is invoked, which integrates not only the appearance and motion information from the two-stream network, but also the motion saliency to mitigate the effect of camera motion. Simulations reveal that the new approach achieves competitive performance as the state-of-the-art works in terms of action localization and recognition accuracy on the widespread UCF101-24 and J-HMDB datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 835054f8-49ed-4a93-83f0-d03611c841e6Cited by top-tier papers4
- Probabilistic Attention for Interactive SegmentationPrasad Gabbur, Manjot Bilkhu, Javier R. MovellanNeurIPS 2021 · 19 citations
- Video Geo-Localization Employing Geo-Temporal Feature Learning and GPS Trajectory SmoothingKrishna Regmi, Mubarak ShahICCV 2021 · 16 citations
- Cefdet: Cognitive Effectiveness Network Based on Fuzzy Inference for Action DetectionZhe Luo, Weina Fu, Shuai Liu, Saeed Anwar et al.ACM MM 2024 · 4 citations
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
Related papers
- Two-Stream Action Recognition-Oriented Video Super-ResolutionHaochen Zhang, Dong Liu, Zhiwei XiongICCV 2019 · 58 citations
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See et al.AAAI 2020 · 18 citations
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 189 citations
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
