Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action Recognition
Zhen Huang, Xu Shen, Xinmei Tian, Houqiang Li, Jianqiang Huang, Xian-Sheng Hua
Abstract
Skeleton-based human action recognition has attracted much attention with the prevalence of accessible depth sensors. Recently, graph convolutional networks (GCNs) have been widely used for this task due to their powerful capability to model graph data. The topology of the adjacency graph is a key factor for modeling the correlations of the input skeletons. Thus, previous methods mainly focus on the design/learning of the graph topology. But once the topology is learned, only a single-scale feature and one transformation exist in each layer of the networks. Many insights, such as multi-scale information and multiple sets of transformations, that have been proven to be very effective in convolutional neural networks (CNNs), have not been investigated in GCNs. The reason is that, due to the gap between graph-structured skeleton data and conventional image/video data, it is very challenging to embed these insights into GCNs. To overcome this gap, we reinvent the split-transform-merge strategy in GCNs for skeleton sequence processing. Specifically, we design a simple and highly modularized graph convolutional network architecture for skeleton-based action recognition. Our network is constructed by repeating a building block that aggregates multi-granularity information from both the spatial and temporal paths. Extensive experiments demonstrate that our network outperforms state-of-the-art methods by a significant margin with only 1/5 of the parameters and 1/10 of the FLOPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 459bb882-7064-4dee-96d7-43fcb97e1fd8Cited by top-tier papers6
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 158 citations
- Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action RecognitionTailin Chen, Desen Zhou, Jian Wang, Shidong Wang et al.ACM MM 2021 · 79 citations
- Towards To-a-T Spatio-Temporal Focus for Skeleton-Based Action RecognitionLipeng Ke, Kuan-Chuan Peng, Siwei LyuAAAI 2022 · 47 citations
- Novel Motion Patterns Matter for Practical Skeleton-Based Action RecognitionMengyuan Liu, Fanyang Meng, Chen Chen, Songtao WuAAAI 2023 · 36 citations
- Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal ApproachYuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen et al.NeurIPS 2025 · 6 citations
Builds on2
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Learning Graph Convolutional Network for Skeleton-Based Human Action Recognition by Neural SearchingWei Peng, Xiaopeng Hong, Haoyu Chen, Guoying ZhaoAAAI 2020 · 362 citations
Related papers
- Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action RecognitionJungho Lee, Minhyeok Lee, Dogyoon Lee, Sangyoun LeeICCV 2023 · 236 citations
- Skeleton-Based Action Recognition With Shift Graph Convolutional NetworkKe Cheng, Yifan Zhang, Xiangyu He, Weihan Chen et al.CVPR 2020
- Topology-Aware Convolutional Neural Network for Efficient Skeleton-Based Action RecognitionKailin Xu, Fanfan Ye, Qiaoyong Zhong, Di XieAAAI 2022 · 168 citations
- Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action RecognitionJianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen et al.AAAI 2024 · 59 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
