Spatio-Temporal Fusion for Human Action Recognition via Joint Trajectory Graph
Yaolin Zheng, Hongbo Huang, Xiuying Wang, Xiaoxu Yan, Longfei Xu
摘要
Graph Convolutional Networks (GCNs) and Transformers have been widely applied to skeleton-based human action recognition, with each offering unique advantages in capturing spatial relationships and long-range dependencies. However, for most GCN methods, the construction of topological structures relies solely on the spatial information of human joints, limiting their ability to directly capture richer spatiotemporal dependencies. Additionally, the self-attention modules of many Transformer methods lack topological structure information, restricting the robustness and generalization of the models. To address these issues, we propose a Joint Trajectory Graph (JTG) that integrates spatio-temporal information into a uniform graph structure. We also present a Joint Trajectory GraphFormer (JT-GraphFormer), which directly captures the spatio-temporal relationships among all joint trajectories for human action recognition. To better integrate topological information into spatio-temporal relationships, we introduce a Spatio-Temporal Dijkstra Attention (STDA) mechanism to calculate relationship scores for all the joints in the JTG. Furthermore, we incorporate the Koopman operator into the classification stage to enhance the model's representation ability and classification performance. Experiments demonstrate that JT-GraphFormer achieves outstanding performance in human action recognition tasks, outperforming state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120, and N-UCLA datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Revealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action RecognitionHongda Liu, Yunfan Liu, Min Ren, Hao Wang 等CVPR 2025
- KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse SignalsShuting Zhao, Zeyu Xiao, Xinrong ChenAAAI 2026
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- InfoGCN: Representation Learning for Human Skeleton-based Action RecognitionHyung-Gun Chi, Myoung Hoon Ha, Seung-geun Chi, Sang Wan Lee 等CVPR 2022 · 被引用 383 次
- Dynamic GCN: Context-enriched Topology Learning for Skeleton-based Action RecognitionFanfan Ye, Shiliang Pu, Qiaoyong Zhong, Chao Li 等ACM MM 2020 · 被引用 348 次
- Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action RecognitionZhan Chen, Sicheng Li, Bing Yang, Qinghan Li 等AAAI 2021 · 被引用 341 次
相关 Paper
- Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action RecognitionWentian Xin, Qiguang Miao, Yi Liu, Ruyi Liu 等ACM MM 2023 · 被引用 66 次
- Disentangling and Unifying Graph Convolutions for Skeleton-Based Action RecognitionZiyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang 等CVPR 2020
- Skeleton-based Human Action Recognition via Large-kernel Attention Graph Convolutional NetworkYanan Liu, Hao Zhang, Yanqiu Li, Kangjian He 等IEEE VR 2023 · 被引用 123 次
- Towards To-a-T Spatio-Temporal Focus for Skeleton-Based Action RecognitionLipeng Ke, Kuan-Chuan Peng, Siwei LyuAAAI 2022 · 被引用 47 次
- Leveraging Spatio-Temporal Dependency for Skeleton-Based Action RecognitionJungho Lee, Minhyeok Lee, Suhwan Cho, Sungmin Woo 等ICCV 2023 · 被引用 28 次
