Ske2Grid: Skeleton-to-Grid Representation Learning for Action Recognition
Dongqi Cai, Yangyuxuan Kang, Anbang Yao, Yurong Chen
Abstract
This paper presents Ske2Grid, a new representation learning framework for improved skeletonbased action recognition. In Ske2Grid, we define a regular convolution operation upon a novel grid representation of human skeleton, which is a compact image-like grid patch constructed and learned through three novel designs. Specifically, we propose a graph-node index transform (GIT) to construct a regular grid patch through assigning the nodes in the skeleton graph one by one to the desired grid cells. To ensure that GIT is a bijection and enrich the expressiveness of the grid representation, an up-sampling transform (UPT) is learned to interpolate the skeleton graph nodes for filling the grid patch to the full. To resolve the problem when the one-step UPT is aggressive and further exploit the representation capability of the grid patch with increasing spatial size, a progressive learning strategy (PLS) is proposed which decouples the UPT into multiple steps and aligns them to multiple paired GITs through a compact cascaded design learned progressively. We construct networks upon prevailing graph convolution networks and conduct experiments on six mainstream skeleton-based action recognition datasets. Experiments show that our Ske2Grid significantly outperforms existing GCN-based solutions under different benchmark settings, without bells and whistles. Code and models are available at https: //github.com/OSVAI/Ske2Grid .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 813140e3-bb68-4bcb-b844-e460e12ae3b4Cited by top-tier papers3
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 37 citations
- Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-Shot Skeleton-Based Action RecognitionJeonghyeok Do, Munchurl KimICCV 2025 · 6 citations
- SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action RecognitionQilang Ye, Yu Zhou, Lian He, Jie Zhang et al.AAAI 2026
Builds on9
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- Learning Graph Convolutional Network for Skeleton-Based Human Action Recognition by Neural SearchingWei Peng, Xiaopeng Hong, Haoyu Chen, Guoying ZhaoAAAI 2020 · 362 citations
- STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action RecognitionYuhan Zhang, Bo Wu, Wen Li, Lixin Duan et al.ACM MM 2021 · 135 citations
- 3D Human Pose Lifting with Grid ConvolutionYangyuxuan Kang, Yuyang Liu, Anbang Yao, Shandong Wang et al.AAAI 2023 · 6 citations
Related papers
- Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action RecognitionZhen Huang, Xu Shen, Xinmei Tian, Houqiang Li et al.ACM MM 2020 · 80 citations
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li et al.ICCV 2023 · 77 citations
- Skeleton-Based Action Recognition With Shift Graph Convolutional NetworkKe Cheng, Yifan Zhang, Xiangyu He, Weihan Chen et al.CVPR 2020
- Part-Level Graph Convolutional Network for Skeleton-Based Action RecognitionLinjiang Huang, Yan Huang, Wanli Ouyang, Liang WangAAAI 2020 · 111 citations
- Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action RecognitionZhan Chen, Sicheng Li, Bing Yang, Qinghan Li et al.AAAI 2021 · 341 citations
