Hierarchical Self-Attention Network for Action Localization in Videos
Rizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien Fang
摘要
This paper presents a novel Hierarchical Self-Attention Network (HISAN) to generate spatial-temporal tubes for action localization in videos. The essence of HISAN is to combine the two-stream convolutional neural network (CNN) with hierarchical bidirectional self-attention mechanism, which comprises of two levels of bidirectional self-attention to efficaciously capture both of the long-term temporal dependency information and spatial context information to render more precise action localization. Also, a sequence rescoring (SR) algorithm is employed to resolve the dilemma of inconsistent detection scores incurred by occlusion or background clutter. Moreover, a new fusion scheme is invoked, which integrates not only the appearance and motion information from the two-stream network, but also the motion saliency to mitigate the effect of camera motion. Simulations reveal that the new approach achieves competitive performance as the state-of-the-art works in terms of action localization and recognition accuracy on the widespread UCF101-24 and J-HMDB datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Probabilistic Attention for Interactive SegmentationPrasad Gabbur, Manjot Bilkhu, Javier R. MovellanNeurIPS 2021 · 被引用 19 次
- Video Geo-Localization Employing Geo-Temporal Feature Learning and GPS Trajectory SmoothingKrishna Regmi, Mubarak ShahICCV 2021 · 被引用 16 次
- Cefdet: Cognitive Effectiveness Network Based on Fuzzy Inference for Action DetectionZhe Luo, Weina Fu, Shuai Liu, Saeed Anwar 等ACM MM 2024 · 被引用 4 次
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
相关 Paper
- Two-Stream Action Recognition-Oriented Video Super-ResolutionHaochen Zhang, Dong Liu, Zhiwei XiongICCV 2019 · 被引用 58 次
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See 等AAAI 2020 · 被引用 18 次
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 被引用 189 次
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu 等ICCV 2019 · 被引用 442 次
