STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action Recognition
Yuhan Zhang, Bo Wu, Wen Li, Lixin Duan, Chuang Gan
Abstract
Skeleton-based action recognition has been widely investigated considering their strong adaptability to dynamic circumstances and complicated backgrounds. To recognize different actions from skeleton sequences, it is essential and crucial to model the posture of the human represented by the skeleton and its changes in the temporal dimension. However, most of the existing works treat skeleton sequences in the temporal and spatial dimension in the same way, ignoring the difference between the temporal and spatial dimension in skeleton data which is not an optimal way to model skeleton sequences. The posture represented by the skeleton in each frame is proposed to be modeled individually. Meanwhile, capturing the movement of the entire skeleton in the temporal dimension is needed. So, we designed Spatial Transformer Block and Directional Temporal Transformer Block for modeling skeleton sequences in spatial and temporal dimensions respectively. Due to occlusion/sensor/raw video, etc., there are noises on both temporal and spatial dimensions in the extracted skeleton data reducing the recognition capabilities of models. To adapt to this imperfect information condition, we propose a multi-task self-supervised learning method by providing confusing samples in different situations to improve the robustness of our model. Combining the above design, we propose our Spatial-Temporal Specialized Transformer (STST) and conduct experiments with our model on the SHREC, NTU-RGB+D, and Kinetics-Skeleton. Extensive experimental results demonstrate the improved performances and analysis of the proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 355d8d73-01bb-4294-9f39-7f733b89cc23Cited by top-tier papers15
- RNTrajRec: Road Network Enhanced Trajectory Recovery with Spatial-Temporal TransformerYuqi Chen, Hanyuan Zhang, Weiwei Sun, Baihua ZhengICDE 2023 · 70 citations
- DirecFormer: A Directed Attention in Transformer Approach to Robust Action RecognitionThanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo et al.CVPR 2022 · 70 citations
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo et al.ICCV 2023 · 38 citations
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 37 citations
- Novel Motion Patterns Matter for Practical Skeleton-Based Action RecognitionMengyuan Liu, Fanyang Meng, Chen Chen, Songtao WuAAAI 2023 · 36 citations
Related papers
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 158 citations
- Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action RecognitionWentian Xin, Qiguang Miao, Yi Liu, Ruyi Liu et al.ACM MM 2023 · 66 citations
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li et al.ICCV 2023 · 77 citations
- Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton SequencesYujie Zhou, Haodong Duan, Anyi Rao, Bing Su et al.AAAI 2023 · 62 citations
- Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action RecognitionTailin Chen, Desen Zhou, Jian Wang, Shidong Wang et al.ACM MM 2021 · 79 citations
