Towards To-a-T Spatio-Temporal Focus for Skeleton-Based Action Recognition
Lipeng Ke, Kuan-Chuan Peng, Siwei Lyu
Abstract
Graph Convolutional Networks (GCNs) have been widely used to model the high-order dynamic dependencies for skeleton-based action recognition. Most existing approaches do not explicitly embed the high-order spatio-temporal importance to joints’ spatial connection topology and intensity, and they do not have direct objectives on their attention module to jointly learn when and where to focus on in the action sequence. To address these problems, we propose the To-a-T Spatio-Temporal Focus (STF), a skeleton-based action recognition framework that utilizes the spatio-temporal gradient to focus on relevant spatio-temporal features. We first propose the STF modules with learnable gradient-enforced and instance-dependent adjacency matrices to model the high-order spatio-temporal dynamics. Second, we propose three loss terms defined on the gradient-based spatio-temporal focus to explicitly guide the classifier when and where to look at, distinguish confusing classes, and optimize the stacked STF modules. STF outperforms the state-of-the-art methods on the NTU RGB+D 60, NTU RGB+D 120, and Kinetics Skeleton 400 datasets in all 15 settings over different views, subjects, setups, and input modalities, and STF also shows better accuracy on scarce data and dataset shifting settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cf02ead-13e2-4eef-90ae-066c877ec43bCited by top-tier papers6
- Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action RecognitionJungho Lee, Minhyeok Lee, Dogyoon Lee, Sangyoun LeeICCV 2023 · 236 citations
- Leveraging Spatio-Temporal Dependency for Skeleton-Based Action RecognitionJungho Lee, Minhyeok Lee, Suhwan Cho, Sungmin Woo et al.ICCV 2023 · 28 citations
- Skeleton-based Action Recognition with Non-linear Dependency Modeling and Hilbert-Schmidt Independence CriterionHaipeng Chen, Yuheng Yang, Yingda LyuAAAI 2025 · 5 citations
- Revealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action RecognitionHongda Liu, Yunfan Liu, Min Ren, Hao Wang et al.CVPR 2025
- Neural Koopman Pooling: Control-Inspired Temporal Dynamics Encoding for Skeleton-Based Action RecognitionXinghan Wang, Xin Xu, Yadong MuCVPR 2023
Builds on7
- Learning Graph Convolutional Network for Skeleton-Based Human Action Recognition by Neural SearchingWei Peng, Xiaopeng Hong, Haoyu Chen, Guoying ZhaoAAAI 2020 · 362 citations
- Dynamic GCN: Context-enriched Topology Learning for Skeleton-based Action RecognitionFanfan Ye, Shiliang Pu, Qiaoyong Zhong, Chao Li et al.ACM MM 2020 · 348 citations
- Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action RecognitionZhen Huang, Xu Shen, Xinmei Tian, Houqiang Li et al.ACM MM 2020 · 80 citations
- AdaSGN: Adapting Joint Number and Model Size for Efficient Skeleton-Based Action RecognitionLei Shi, Yifan Zhang, Jian Cheng, Hanqing LuICCV 2021 · 60 citations
- Sharpen Focus: Learning With Attention Separability and ConsistencyLezi Wang, Ziyan Wu, Srikrishna Karanam, Kuan-Chuan Peng et al.ICCV 2019 · 37 citations
Related papers
- Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action RecognitionZhan Chen, Sicheng Li, Bing Yang, Qinghan Li et al.AAAI 2021 · 341 citations
- Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action RecognitionJianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen et al.AAAI 2024 · 59 citations
- Skeleton-based Human Action Recognition via Large-kernel Attention Graph Convolutional NetworkYanan Liu, Hao Zhang, Yanqiu Li, Kangjian He et al.IEEE VR 2023 · 123 citations
- Disentangling and Unifying Graph Convolutions for Skeleton-Based Action RecognitionZiyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang et al.CVPR 2020
- Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action RecognitionWentian Xin, Qiguang Miao, Yi Liu, Ruyi Liu et al.ACM MM 2023 · 66 citations
